Is it possible to get all occurrences of \< in a string using regexec?
Florian Weimer
fweimer@redhat.com
Mon Feb 10 08:20:58 GMT 2025
* Arkadiusz Drabczyk:
> \< GNU extension matches the beginning of a word, for example when
> used with GNU sed:
>
> $ echo abc def xyz | sed 's,\<,*,g'
> *abc *def *xyz
>
> The problem is when that \< is used with regexec() on all substrings
> of a string incremented by 1 from the beginning until the end matches
> with rm_so and rm_eo set to 0 are always returned so effectively all
> characters in a word match. The behavior can be demonstrated with
> busybox sed that uses regexec() that way:
>
> $ echo abc def xyz | busybox sed 's,\<,*,g'
> *a*b*c *d*e*f *x*y*z
>
> The GNU sed example works because it uses gnulib regex. The difference
> is that gnulib implementation can operate on the same string and
> return the matches from the given offset but regexc() takes a new
> string and returns all matches form its beginning.
>
> Is it possible to use regexc() in a way that would make it possible to
> correctly match all occurrences of \< in a string?
Have you tried calling egexec with REG_NOTBOL? It is documented as:
‘REG_NOTBOL’
Do not regard the beginning of the specified string as the
beginning of a line; more generally, don't make any assumptions
about what text might precede it.
Maybe it helps your use case as well?
Probably even better is REG_STARTEND (only documented in the manual
pages). This avoids one source of quadratic run-time behavior.
Thanks,
Florian
More information about the Libc-help
mailing list