[V2] riscv: Implement Zbb based strlen and prefer it over the RVV based strlen implementation when Zbb is available
Jeffrey Law
jeffrey.law@oss.qualcomm.com
Sat Jul 4 06:34:26 GMT 2026
So this is the V2 patch of a Zbb strlen implementation. As was
previously noted, this is 2-4X faster than the current RVV
implementation on the K3 and meaningfully faster on the K1 as well (I
don't remember that data offhand other than Zbb was the best choice
there too).
The most important difference between this and the first patch is
there's no longer a Zbb specific directory. Per the discussion from
last month there aren't any plans to make any Implies relationships and
such.
The ifunc resolver has been improved ever-so-slightly to avoid an extra
round trip through the hwprobe interface. We can get the state of Zbb
and RVV with a single round trip. A few comment typos spotted by an
LLM have been fixed as well.
With dropping the Zbb subdirectory, the bits left in the multiarch
directory have all the hidden symbol, alias and related stuff. I'm not
at all familiar with what needs to be done in this case. So I'd
appreciate a close look at that code.
I've built and tested glibc on the K3 with this patch. It shows no
regressions relative to the baseline build. I've also verified the
performance data is not meaningfully changed.
Obviously hoping we can get this included in the upcoming release...
--
So we've had Zbb variants for strlen, strcmp and a few other routines
sitting here in our local repositories for a long time. The original
implementations were done by the VRULL team, then adjusted for minor
bugs caught by the glibc testsuite and later wired into the hwprobe
mechanism.
Much like the RVV implementations that have been dropping into the tree,
I want to focus on one routine at a time to make sure we're happy with
the result, then move onto the next one. In this particular patch I'm
focused on strlen.
The implementation is largely derived from the bitmanip examples, just
cleaned up so that it ought to work for both rv32/rv64 and either big or
little endian (little endian is untested, I believe VRULL tested rv32 at
some point).
Neither the Zbb nor the RVV implementation seems at all sensitive to
data alignment concerns on the K3. So we can safely ignore that input
axis and focus on how many cycles it takes to handle a string of a
particular length.
I asked the LLM model to take the performance data, convert it to cycles
per byte, then get the average cycles per byte over a range of lengths
new buckets starting a power of 2 boundaries.
Bucket ZBB CPB Vector CBP Winner
1-1 4.227 17.312 ZBB is ~4.1x faster
2-3 1.714 7.232 ZBB is ~4.2x faster
4-7 0.870 3.287 ZBB is ~3.8x faster
8-15 0.563 1.950 ZBB is ~3.5x faster
16-31 0.446 0.954 ZBB is ~2.1x faster
32-63 0.299 0.477 ZBB is ~1.6x faster
64-127 0.190 0.414 ZBB is ~2.2x faster
And so-on with the cycles-per-byte dropping for both, but ZBB
consistently running ~2.1x faster than RVV up to a length of 8k.
We can see the Zbb is just better all around. There wasn't a single
case where RVV won. It's pretty obvious that the vector version has a
higher fixed overhead, but I really expected vector to overcome that
overhead as the strings got longer. As it stands the data says quite
clearly that we should be using Zbb on the K3 design and likely the K1
design (currently being tested).
Given the K1/K3 designs are what folks can get their hands on, I'd
recommend we make Zbb preferred over RVV. We'll likely have to adjust
that as newer designs come into the market, but the decision should be
data driven. I'm going to run this on our Veyron V2 design and Peter is
going to run on the Ascalon design, but neither of those are generally
available and probably shouldn't drive decisions, those are mostly for
informational purposes and to give a sense of whether or not higher
targeted designs are likely to benefit from the RVV variant when those
higher performance designs hit the market.
You could also legitimately ask what GCC should be doing here. Right now
GCC will inline the strlen call, generating RVV code that is nearly
identical to what's in glibc. So it's probably not a win for GCC to
inline an RVV strlen, though inlining does at least avoid the function
call overhead and allow for secondary optimization affects since there's
no call.
This has been built and regression tested on the c920 and K3, the K1 is
still running. The c920 is interesting because it has neither RVV nor
Zbb, so confirming I didn't do anything dumb in the resolver was useful.
--
OK for the trunk?
jeff
-------------- next part --------------
From c26fdb04b7b64e36768686407fe110fe30ee49a3 Mon Sep 17 00:00:00 2001
From: Jeff Law <jlaw@k3.lan>
Date: Tue, 9 Jun 2026 23:05:37 +0000
Subject: [PATCH] strlen zbb implementation
---
sysdeps/riscv/multiarch/strlen-zbb.S | 126 ++++++++++++++++++
.../unix/sysv/linux/riscv/multiarch/Makefile | 1 +
.../linux/riscv/multiarch/ifunc-impl-list.c | 6 +
.../unix/sysv/linux/riscv/multiarch/strlen.c | 17 ++-
4 files changed, 147 insertions(+), 3 deletions(-)
create mode 100644 sysdeps/riscv/multiarch/strlen-zbb.S
diff --git a/sysdeps/riscv/multiarch/strlen-zbb.S b/sysdeps/riscv/multiarch/strlen-zbb.S
new file mode 100644
index 0000000000..f503acfda7
--- /dev/null
+++ b/sysdeps/riscv/multiarch/strlen-zbb.S
@@ -0,0 +1,126 @@
+/* Re-include the RISC-V Zbb based strlen implementation.
+ Copyright (C) 2026 Free Software Foundation, Inc.
+ This file is part of the GNU C Library.
+
+ The GNU C Library is free software; you can redistribute it and/or
+ modify it under the terms of the GNU Lesser General Public
+ License as published by the Free Software Foundation; either
+ version 2.1 of the License, or (at your option) any later version.
+
+ The GNU C Library is distributed in the hope that it will be useful,
+ but WITHOUT ANY WARRANTY; without even the implied warranty of
+ MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU
+ Lesser General Public License for more details.
+
+ You should have received a copy of the GNU Lesser General Public
+ License along with the GNU C Library; if not, see
+ <https://www.gnu.org/licenses/>. */
+
+#if IS_IN(libc)
+# define STRLEN __strlen_zbb
+# undef libc_hidden_builtin_def
+# define libc_hidden_builtin_def(name)
+# undef weak_alias
+# define weak_alias(name, alias)
+
+#include <sysdep.h>
+#include <sys/asm.h>
+
+/* Assumptions: rvi_zbb. */
+/* Implementation from the Bitmanip specification. */
+
+#define src a0
+#define result a0
+#define addr a1
+#define data a2
+#define offset a3
+#define offset_bits a3
+#define valid_bytes a4
+#define m1 a4
+
+#if __riscv_xlen == 64
+# define REG_L ld
+# define SZREG 8
+# define PTRLOG 3
+#else
+# define REG_L lw
+# define SZREG 4
+# define PTRLOG 2
+#endif
+
+#if __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
+# define CZ clz
+# define SHIFT sll
+#else
+# define CZ ctz
+# define SHIFT srl
+#endif
+
+#ifndef STRLEN
+# define STRLEN __strlen_zbb
+#endif
+
+ENTRY (STRLEN)
+.option push
+.option arch,+zbb
+
+ /* Number of irrelevant bytes in the first word. */
+ andi offset, src, SZREG-1
+ /* Align pointer. */
+ andi addr, src, -SZREG
+
+ li valid_bytes, SZREG
+ sub valid_bytes, valid_bytes, offset
+ slli offset_bits, offset, PTRLOG
+
+ /* Get the first word. */
+ REG_L data, 0(addr)
+ /* Shift away the partial data we loaded to remove the irrelevant bytes
+ * preceding the string with the effect of adding NUL bytes at the
+ * end of the string. */
+ SHIFT data, data, offset_bits
+ /* Convert non-NUL into 0xff and NUL into 0x00. */
+ orc.b data, data
+ /* Convert non-NUL into 0x00 and NUL into 0xff. */
+ not data, data
+ /* Search for the first set bit (corresponding to a NUL byte in the
+ * original chunk). */
+ CZ data, data
+ /* The first chunk is special: compare against the number
+ * of valid bytes in this chunk. */
+ srli result, data, 3
+ bgtu valid_bytes, result, L(done)
+
+ /* Prepare for the word comparison loop. */
+ addi offset, addr, SZREG
+ li m1, -1
+
+ /* Our critical loop is 4 instructions and processes data in
+ * 4 byte or 8 byte chunks. */
+ .p2align 3
+L(loop):
+ REG_L data, SZREG(addr)
+ addi addr, addr, SZREG
+ orc.b data, data
+ beq data, m1, L(loop)
+
+L(epilogue):
+ not data, data
+ CZ data, data
+ /* Get number of processed words. */
+ sub offset, addr, offset
+ /* Add number of characters in the first word. */
+ add result, result, offset
+ srli data, data, 3
+ /* Add number of characters in the last word. */
+ add result, result, data
+L(done):
+ ret
+
+.option pop
+
+END (STRLEN)
+libc_hidden_builtin_def (STRLEN)
+weak_alias (STRLEN, strlen)
+
+#endif
diff --git a/sysdeps/unix/sysv/linux/riscv/multiarch/Makefile b/sysdeps/unix/sysv/linux/riscv/multiarch/Makefile
index 30d92c1ba9..0e42022fab 100644
--- a/sysdeps/unix/sysv/linux/riscv/multiarch/Makefile
+++ b/sysdeps/unix/sysv/linux/riscv/multiarch/Makefile
@@ -31,6 +31,7 @@ sysdep_routines += \
strlen \
strlen-generic \
strlen-vector \
+ strlen-zbb \
strncmp \
strncmp-generic \
strncmp-vector \
diff --git a/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c b/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c
index 8027578529..4b0ac4ea51 100644
--- a/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c
+++ b/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c
@@ -27,6 +27,7 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
size_t i = max;
bool fast_unaligned = false;
+ bool zbb_enabled = false;
bool rvv_enabled = false;
struct riscv_hwprobe pairs[2] = {
@@ -41,6 +42,9 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
if (pairs[1].value & RISCV_HWPROBE_IMA_V)
rvv_enabled = true;
+
+ if (pairs[1].value & RISCV_HWPROBE_EXT_ZBB)
+ zbb_enabled = true;
}
IFUNC_IMPL (i, name, memcpy,
@@ -66,6 +70,8 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
IFUNC_IMPL_ADD (array, i, strcpy, 1, __strcpy_generic))
IFUNC_IMPL (i, name, strlen,
+ IFUNC_IMPL_ADD (array, i, strlen, zbb_enabled,
+ __strlen_zbb)
IFUNC_IMPL_ADD (array, i, strlen, rvv_enabled,
__strlen_vector)
IFUNC_IMPL_ADD (array, i, strlen, 1, __strlen_generic))
diff --git a/sysdeps/unix/sysv/linux/riscv/multiarch/strlen.c b/sysdeps/unix/sysv/linux/riscv/multiarch/strlen.c
index 9975286b85..c00f2787ff 100644
--- a/sysdeps/unix/sysv/linux/riscv/multiarch/strlen.c
+++ b/sysdeps/unix/sysv/linux/riscv/multiarch/strlen.c
@@ -32,14 +32,25 @@ extern __typeof (__redirect_strlen) __libc_strlen;
extern __typeof (__redirect_strlen) __strlen_generic attribute_hidden;
extern __typeof (__redirect_strlen) __strlen_vector attribute_hidden;
+extern __typeof (__redirect_strlen) __strlen_zbb attribute_hidden;
static inline __typeof (__redirect_strlen) *
select_strlen_ifunc (uint64_t dl_hwcap, __riscv_hwprobe_t hwprobe_func)
{
unsigned long long int v;
- if (__riscv_hwprobe_one (hwprobe_func, RISCV_HWPROBE_KEY_IMA_EXT_0, &v) == 0
- && (v & RISCV_HWPROBE_IMA_V) == RISCV_HWPROBE_IMA_V)
- return __strlen_vector;
+
+ /* Testing has shown that on circa 2026 hardware a Zbb based strlen is
+ consistently faster than a V implementation. So we prefer Zbb for
+ now. As vector impementations mature this will likely need revisiting. */
+ if (__riscv_hwprobe_one (hwprobe_func, RISCV_HWPROBE_KEY_IMA_EXT_0, &v) == 0)
+ {
+ if ((v & RISCV_HWPROBE_EXT_ZBB) == RISCV_HWPROBE_EXT_ZBB)
+ return __strlen_zbb;
+
+ if ((v & RISCV_HWPROBE_IMA_V) == RISCV_HWPROBE_IMA_V)
+ return __strlen_vector;
+ }
+
return __strlen_generic;
}
--
2.47.3
More information about the Libc-alpha
mailing list