[Bug string/33091] New: Power 10 rawmemchr clobbers v20
fweimer at redhat dot com
sourceware-bugzilla@sourceware.org
Mon Jun 16 19:49:18 GMT 2025
https://sourceware.org/bugzilla/show_bug.cgi?id=33091
Bug ID: 33091
Summary: Power 10 rawmemchr clobbers v20
Product: glibc
Version: unspecified
Status: NEW
Severity: normal
Priority: P2
Component: string
Assignee: unassigned at sourceware dot org
Reporter: fweimer at redhat dot com
CC: carlos at redhat dot com, fweimer at redhat dot com,
Sachin.Monga at ibm dot com, siddhesh at sourceware dot org,
unassigned at sourceware dot org
Depends on: 33059
Target Milestone: ---
Target: powerpc
+++ This bug was initially created as a clone of Bug #33059 +++
The Power10 rawmemchr uses v20 as a zero register without saving/restoring it,
which is a violation of the ppc64le psABI.
This dates back to this commit in glibc 2.34:
commit 1a594aa986ffe28657a03baa5c53c0a0e7dc2ecd
Author: Matheus Castanho <msc@linux.ibm.com>
Date: Tue May 11 17:53:07 2021 -0300
powerpc: Add optimized rawmemchr for POWER10
Reuse code for optimized strlen to implement a faster version of rawmemchr.
This takes advantage of the same benefits provided by the strlen
implementation,
but needs some extra steps. __strlen_power10 code should be unchanged after
this
change.
rawmemchr returns a pointer to the char found, while strlen returns only
the
length, so we have to take that into account when preparing the return
value.
To quickly check 64B, the loop on __strlen_power10 merges the whole block
into
16B by using unsigned minimum vector operations (vminub) and checks if
there are
any \0 on the resulting vector. The same code is used by rawmemchr if the
char c
is 0. However, this approach does not work when c != 0. We first need to
subtract each byte by c, so that the value we are looking for is converted
to a
0, then taking the minimum and checking for nulls works again.
The new code branches after it has compared ~256 bytes and chooses which of
the
two strategies above will be used in the main loop, based on the char c.
This
extra branch adds some overhead (~5%) for length ~256, but is quickly
amortized
by the faster loop for larger sizes.
Compared to __rawmemchr_power9, this version is ~20% faster for length <
256.
Because of the optimized main loop, the improvement becomes ~35% for c != 0
and ~50% for c = 0 for strings longer than 256.
Reviewed-by: Lucas A. M. Magalhaes <lamm@linux.ibm.com>
Reviewed-by: Raphael M Zinsly <rzinsly@linux.ibm.com>
The strlen function is not impacted by this bug because it uses v18 as
VREG_ZERO.
Referenced Bugs:
https://sourceware.org/bugzilla/show_bug.cgi?id=33059
[Bug 33059] Power 10 memchr clobbers v20
--
You are receiving this mail because:
You are on the CC list for the bug.
More information about the Glibc-bugs
mailing list