[PATCH] riscv: Avoid vector gathers in RVV memcmp

Zhang Jinhan zhangjinhan@isrc.iscas.ac.cn
Tue Sep 15 08:40:30 GMT 2026


The mismatch path gathers each mismatching byte into a new vector
register, moves both bytes to scalar registers, and masks off the upper
bits.  This is costly on implementations with slow vector gather and
vector-to-scalar operations.

vfirst.m already provides the byte offset of the first mismatch, and
the source pointers still address the start of the active chunk.  Load
the two bytes directly instead.  lbu provides the required zero
extension, so the masks are unnecessary.

Benchmark results on SpacemiT K3 (X100 cores, RV64, VLEN=256), using
bench-memcmp built with the following CFLAGS:

  -O3 -g -fno-omit-frame-pointer -march=rv64gcv -mabi=lp64d

Each side was run three times, with 491 cases per run.  Values are the
hp_timing values reported per iteration; lower is better.  Improvement
is the percentage reduction in execution time.

                              Baseline   Patched  Improvement
  Run 1                        45.1949   32.4938       28.10%
  Run 2                        45.2616   27.5491       39.13%
  Run 3                        45.4126   32.5649       28.29%
  Per-case median geomean       45.1266   32.3903       28.22%

Breakdown of the per-case medians:

                              Cases  Baseline  Patched  Improvement
  Zero length                     9   44.6090  44.5972        0.03%
  Equal, nonzero length         166   32.1503  32.1348        0.05%
  Different, nonzero length     316   53.9419  32.2304       40.25%

Breakdown by length for differing inputs, where the final byte differs:

  Length  Cases  Baseline  Patched  Improvement
       1      6   48.1146  27.6128       42.61%
       8     10   48.0704  27.5465       42.70%
      16     12   48.0649  27.4185       42.96%
      32     10   48.2266  27.5591       42.86%
      64     10   48.7290  27.4802       43.61%
     128     10   48.4850  27.4871       43.31%
     256     10   48.4546  27.4934       43.26%
     512     10   67.0307  46.5056       30.62%
    1024     10  104.6051  83.6261       20.06%
    2048      6  176.8793 156.2749       11.65%
    4096      6  323.1028 302.0463        6.52%

All 316 nonempty differing-input cases improved.  The two comparison
directions improved by 40.29% and 40.21%, respectively.  Across all
491 cases, the largest improvement was 45.71%, the largest regression
was 0.84%, and no regression of 5% or more was repeatable.

Tested on riscv64 with string/test-memcmp and string/test-memcmpeq.

Co-authored-by: Ning Tian <tianning24@iscas.ac.cn>
Co-authored-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
Signed-off-by: Ning Tian <tianning24@iscas.ac.cn>
Signed-off-by: Yuansheng <yuansheng@isrc.iscas.ac.cn>
Signed-off-by: Zhang Jinhan <zhangjinhan@isrc.iscas.ac.cn>
---
 sysdeps/riscv/rvv/memcmp.S | 15 +++++++--------
 1 file changed, 7 insertions(+), 8 deletions(-)

diff --git a/sysdeps/riscv/rvv/memcmp.S b/sysdeps/riscv/rvv/memcmp.S
index 004bf538f0..3d81cc6eb0 100644
--- a/sysdeps/riscv/rvv/memcmp.S
+++ b/sysdeps/riscv/rvv/memcmp.S
@@ -31,13 +31,12 @@
 
 #define ivl a3
 #define temp a4
+#define addr a5
 
 #define ELEM_LMUL_SETTING m8
 #define vdata1 v0
 #define vdata2 v8
 #define vmask v16
-#define vtemp1 v16
-#define vtemp2 v24
 
 ENTRY (MEMCMP)
 .option push
@@ -62,12 +61,12 @@ L(loop):
     li result, 0
     ret
 L(found):
-    vrgather.vx vtemp1, vdata1, temp
-    vrgather.vx vtemp2, vdata2, temp
-    vmv.x.s result, vtemp1
-    vmv.x.s temp, vtemp2
-    andi result, result, 0xff
-    andi temp, temp, 0xff
+    /* src1/src2 still point at the chunk bases when this branch is taken,
+       so the differing bytes can be loaded directly.  */
+    add addr, src1, temp
+    lbu result, 0(addr)
+    add addr, src2, temp
+    lbu temp, 0(addr)
     sub result, result, temp
     ret
 .option pop
-- 
2.43.0



More information about the Libc-alpha mailing list