Is it possible to get all occurrences of \< in a string using regexec?

Florian Weimer fweimer@redhat.com
Mon Feb 10 08:20:58 GMT 2025


* Arkadiusz Drabczyk:

> \< GNU extension matches the beginning of a word, for example when
> used with GNU sed:
>
> $ echo abc def xyz | sed 's,\<,*,g'
> *abc *def *xyz
>
> The problem is when that \< is used with regexec() on all substrings
> of a string incremented by 1 from the beginning until the end matches
> with rm_so and rm_eo set to 0 are always returned so effectively all
> characters in a word match. The behavior can be demonstrated with
> busybox sed that uses regexec() that way:
>
> $ echo abc def xyz | busybox sed 's,\<,*,g'
> *a*b*c *d*e*f *x*y*z
>
> The GNU sed example works because it uses gnulib regex. The difference
> is that gnulib implementation can operate on the same string and
> return the matches from the given offset but regexc() takes a new
> string and returns all matches form its beginning.
>
> Is it possible to use regexc() in a way that would make it possible to
> correctly match all occurrences of \< in a string?

Have you tried calling egexec with REG_NOTBOL?  It is documented as:

‘REG_NOTBOL’

     Do not regard the beginning of the specified string as the
     beginning of a line; more generally, don't make any assumptions
     about what text might precede it.

Maybe it helps your use case as well?

Probably even better is REG_STARTEND (only documented in the manual
pages).  This avoids one source of quadratic run-time behavior.

Thanks,
Florian



More information about the Libc-help mailing list