[PATCH] riscv: Avoid vector gathers in RVV memcmp
Zhang Jinhan
zhangjinhan@isrc.iscas.ac.cn
Tue Sep 15 08:40:30 GMT 2026
The mismatch path gathers each mismatching byte into a new vector
register, moves both bytes to scalar registers, and masks off the upper
bits. This is costly on implementations with slow vector gather and
vector-to-scalar operations.
vfirst.m already provides the byte offset of the first mismatch, and
the source pointers still address the start of the active chunk. Load
the two bytes directly instead. lbu provides the required zero
extension, so the masks are unnecessary.
Benchmark results on SpacemiT K3 (X100 cores, RV64, VLEN=256), using
bench-memcmp built with the following CFLAGS:
-O3 -g -fno-omit-frame-pointer -march=rv64gcv -mabi=lp64d
Each side was run three times, with 491 cases per run. Values are the
hp_timing values reported per iteration; lower is better. Improvement
is the percentage reduction in execution time.
Baseline Patched Improvement
Run 1 45.1949 32.4938 28.10%
Run 2 45.2616 27.5491 39.13%
Run 3 45.4126 32.5649 28.29%
Per-case median geomean 45.1266 32.3903 28.22%
Breakdown of the per-case medians:
Cases Baseline Patched Improvement
Zero length 9 44.6090 44.5972 0.03%
Equal, nonzero length 166 32.1503 32.1348 0.05%
Different, nonzero length 316 53.9419 32.2304 40.25%
Breakdown by length for differing inputs, where the final byte differs:
Length Cases Baseline Patched Improvement
1 6 48.1146 27.6128 42.61%
8 10 48.0704 27.5465 42.70%
16 12 48.0649 27.4185 42.96%
32 10 48.2266 27.5591 42.86%
64 10 48.7290 27.4802 43.61%
128 10 48.4850 27.4871 43.31%
256 10 48.4546 27.4934 43.26%
512 10 67.0307 46.5056 30.62%
1024 10 104.6051 83.6261 20.06%
2048 6 176.8793 156.2749 11.65%
4096 6 323.1028 302.0463 6.52%
All 316 nonempty differing-input cases improved. The two comparison
directions improved by 40.29% and 40.21%, respectively. Across all
491 cases, the largest improvement was 45.71%, the largest regression
was 0.84%, and no regression of 5% or more was repeatable.
Tested on riscv64 with string/test-memcmp and string/test-memcmpeq.
Co-authored-by: Ning Tian <tianning24@iscas.ac.cn>
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
Signed-off-by: Ning Tian <tianning24@iscas.ac.cn>
Signed-off-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
Signed-off-by: Zhang Jinhan <zhangjinhan@isrc.iscas.ac.cn>
---
sysdeps/riscv/rvv/memcmp.S | 15 +++++++--------
1 file changed, 7 insertions(+), 8 deletions(-)
diff --git a/sysdeps/riscv/rvv/memcmp.S b/sysdeps/riscv/rvv/memcmp.S
index 004bf538f0..3d81cc6eb0 100644
--- a/sysdeps/riscv/rvv/memcmp.S
+++ b/sysdeps/riscv/rvv/memcmp.S
@@ -31,13 +31,12 @@
#define ivl a3
#define temp a4
+#define addr a5
#define ELEM_LMUL_SETTING m8
#define vdata1 v0
#define vdata2 v8
#define vmask v16
-#define vtemp1 v16
-#define vtemp2 v24
ENTRY (MEMCMP)
.option push
@@ -62,12 +61,12 @@ L(loop):
li result, 0
ret
L(found):
- vrgather.vx vtemp1, vdata1, temp
- vrgather.vx vtemp2, vdata2, temp
- vmv.x.s result, vtemp1
- vmv.x.s temp, vtemp2
- andi result, result, 0xff
- andi temp, temp, 0xff
+ /* src1/src2 still point at the chunk bases when this branch is taken,
+ so the differing bytes can be loaded directly. */
+ add addr, src1, temp
+ lbu result, 0(addr)
+ add addr, src2, temp
+ lbu temp, 0(addr)
sub result, result, temp
ret
.option pop
--
2.43.0
More information about the Libc-alpha
mailing list