[PATCH] riscv: Optimize vectorized memory routines with ifunc dispatch
Peter Bergner
bergner@oss.tenstorrent.com
Wed Sep 9 03:10:12 GMT 2026
On 9/7/26 10:37 PM, Zheng Ziyang wrote:
> Performance results:
>
> vlen 128 (SG2044):
[snip]
> - memchr: performance degradation
> - memmove-large: performance degradation
How large of degradations?
> vlen 256 (Spacemit® A100):
Do you mean the Spacemit X100 core? IIRC, it's the X100 core that has
VLEM=256, while the A100 has VLEN=1024.
> - memcmp: 0.31
> - memcpy: 0.203
> - memcpy-large: 0.013
> - memcpy-random: 0.425
> - memmove: 0.26
> - memmove-large: -0.002
> - memrchr: -0.008
> - memset-zero: 0.06
> - memset: 0.39
> - memset-zero: 0.33
> - memset-large: 0.005
> - memset-large-zero: 0.012
Looking at your data for the SG2044 and Spacemit cores, it looks like your
larger improvements are due mainly to your special handling of smaller and
medium sized data, correct?
> diff --git a/sysdeps/riscv/multiarch/memrchr-generic.c b/sysdeps/riscv/multiarch/memrchr-generic.c
> new file mode 100644
> index 0000000000..fbe1598382
> --- /dev/null
> +++ b/sysdeps/riscv/multiarch/memrchr-generic.c
> @@ -0,0 +1,26 @@
> +/* Re-include the default memcpy implementation.
> + Copyright (C) 2024 Free Software Foundation, Inc.
> + This file is part of the GNU C Library.
There are lots of issues with you either adding new files with the wrong Copyright
date (it should be 2026) or you patching a 2026 Copyright date on an existing file
with the wrong date/dates. My guess is you have not refreshed your changes before
creating the patch. Please fix that. It happens in many files.
> diff --git a/sysdeps/riscv/multiarch/rawmemchr-generic.c b/sysdeps/riscv/multiarch/rawmemchr-generic.c
> new file mode 100644
> index 0000000000..d32e4c69a0
> --- /dev/null
> +++ b/sysdeps/riscv/multiarch/rawmemchr-generic.c
> @@ -0,0 +1,31 @@
> +/* Re-include the default memcpy implementation.
> + Copyright (C) 2024 Free Software Foundation, Inc.
Here's another example. You also mention reincluding the default
"memcpy" implementation when in actualality, you're reincluding the
rawmemchr implementation. You should check your other comments to
make sure they are correct too.
> diff --git a/sysdeps/riscv/rvv/memchr.S b/sysdeps/riscv/rvv/memchr.S
> index 609b2f94c1..f85e250819 100644
> --- a/sysdeps/riscv/rvv/memchr.S
> +++ b/sysdeps/riscv/rvv/memchr.S
[snip]
> + beqz a2, .Lnot_found
I'm not against zero checks if they come up in actual usage, but is
calling memchr with n == 0 actually a common thing? If not, then we're
just adding an extra untaken branch on the common path.
> + li t4, 64
> + bltu a2, t4, .Ltail_entry
> +
> + vsetvli t0, t4, e8, m4, ta, ma
I don't like forcing AVL=64 here. With your LMUL=4, that maps
perfectly to a VLEN=128 core, but cores with longer VLENs are
just going to have unused vector registers. Looking at your perf
data, it seems your new memchr code is slower than the existing
code on the SG2044 and you don't give any perf data for the Spacemit
core.
> diff --git a/sysdeps/riscv/rvv/memcmp.S b/sysdeps/riscv/rvv/memcmp.S
[snip]
> + /* main loop, vlenb*2 elts at a time */
> + vsetvli t1, a2, e8, m2, ta, ma
I'd like to know more about where your memcmp speedups are
coming from. Is it the LMUL=8 from the original code doesn't
perform as well as your LMUL=2 code? ...or is it your handling
of small # elts code or ??? Can you speak to that?
> diff --git a/sysdeps/riscv/rvv/memcpy.S b/sysdeps/riscv/rvv/memcpy.S
[snip]
> +L(copy128):
> + li t3, 64
> + vsetvli zero, t3, e8, m4, ta, ma # vl=64 for all VLEN
I have the same complaint of the forced AVL here, similar to memchr above.
> diff --git a/sysdeps/riscv/rvv/memmove.S b/sysdeps/riscv/rvv/memmove.S
...
This is a lot of changes, even to the large memove code and your perf
data only shows your new large memmove code is slightly slower. Do we
only need to add smaller sized memmove handling to the current code to
get improvements? If so, which smaller handling do we need?
> diff --git a/sysdeps/riscv/rvv/memset.S b/sysdeps/riscv/rvv/memset.S
[snip]
> +ENTRY(MEMSET)
> +
> + .option push
> + .option arch, +v, +zicboz
I see you're adding zicboz to the arch, but I don't see you actually using it.
Old code or ??? I believe I've heard that a zicboz memset is actually slower
on the Spacemit K3 than our current RVV based memset.
> diff --git a/sysdeps/riscv/rvv/strlen.S b/sysdeps/riscv/rvv/strlen.S
Have you seen the Zbb based strlen patch sent before the release?
It significantly out performed the RVV based routine and will likely be
our preferred strlen implementation when we move to a profile based design.
> diff --git a/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c b/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c
> index ac2b25de74..cf1f783139 100644
> --- a/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c
> +++ b/sysdeps/unix/sysv/linux/riscv/multiarch/ifunc-impl-list.c
> @@ -28,6 +28,7 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
>
> bool fast_unaligned = false;
> bool rvv_enabled = false;
> + bool zicboz_ext = false;
>
> struct riscv_hwprobe pairs[2] = {
> {.key = RISCV_HWPROBE_KEY_CPUPERF_0},
> @@ -41,8 +42,13 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
>
> if (pairs[1].value & RISCV_HWPROBE_IMA_V)
> rvv_enabled = true;
> +
> + if (pairs[1].value & RISCV_HWPROBE_EXT_ZICBOZ)
> + zicboz_ext = true;
> +
> }
As mentioned above, I see no actual usage of zicboz, so I think this code
isn't needed.
> IFUNC_IMPL (i, name, memset,
> - IFUNC_IMPL_ADD (array, i, memset, rvv_enabled,
> + IFUNC_IMPL_ADD (array, i, memset, (rvv_enabled && zicboz_ext),
> __memset_vector)
> IFUNC_IMPL_ADD (array, i, memset, 1, __memset_generic))
...nor this.
> diff --git a/sysdeps/unix/sysv/linux/riscv/multiarch/memset.c b/sysdeps/unix/sysv/linux/riscv/multiarch/memset.c
> index 1b0c2db711..e43a2d65a2 100644
> --- a/sysdeps/unix/sysv/linux/riscv/multiarch/memset.c
> +++ b/sysdeps/unix/sysv/linux/riscv/multiarch/memset.c
> @@ -39,7 +39,8 @@ select_memset_ifunc (uint64_t dl_hwcap, __riscv_hwprobe_t hwprobe_func)
> unsigned long long v;
>
> if (__riscv_hwprobe_one (hwprobe_func, RISCV_HWPROBE_KEY_IMA_EXT_0, &v) == 0
> - && (v & RISCV_HWPROBE_IMA_V) == RISCV_HWPROBE_IMA_V)
> + && (v & RISCV_HWPROBE_IMA_V) == RISCV_HWPROBE_IMA_V
> + && (v & RISCV_HWPROBE_EXT_ZICBOZ) == RISCV_HWPROBE_EXT_ZICBOZ)
> return __memset_vector;
...nor this.
> diff --git a/sysdeps/unix/sysv/linux/riscv/sys/hwprobe.h b/sysdeps/unix/sysv/linux/riscv/sys/hwprobe.h
> index ff200272b3..f39fc57fa9 100644
> --- a/sysdeps/unix/sysv/linux/riscv/sys/hwprobe.h
> +++ b/sysdeps/unix/sysv/linux/riscv/sys/hwprobe.h
> @@ -62,6 +62,11 @@ struct riscv_hwprobe {
>
> #endif /* RISCV_HWPROBE_KEY_MVENDORID */
>
> +/* Fallback definitions for kernel headers that predate these macros. */
> +#ifndef RISCV_HWPROBE_EXT_ZICBOZ
> +# define RISCV_HWPROBE_EXT_ZICBOZ (1 << 6)
> +#endif
...nor this.
Peter
More information about the Libc-alpha
mailing list