Is it possible to get all occurrences of \< in a string using regexec?

Arkadiusz Drabczyk arkadiusz@drabczyk.org
Wed Feb 12 22:02:36 GMT 2025


On Mon, Feb 10, 2025 at 10:20:09PM +0100, Arkadiusz Drabczyk wrote:
> On Mon, Feb 10, 2025 at 07:52:46PM +0100, Florian Weimer wrote:
> > This:
> > 
> >   size_t eo = strlen(text);
> >   while (offset < len) {
> >     match[0].rm_so = offset;
> >     match[0].rm_eo = eo;
> >     ret = regexec(&regex, text, 1, match, REG_STARTEND);
> > 
> >     if (ret == REG_NOMATCH) {
> >       break;
> >     }
> > 
> >     printf("match[0].rm_so == %d\n", match[0].rm_so);
> >     printf("match[0].rm_eo == %d\n", match[0].rm_eo);
> > 
> >     int start = match[0].rm_so;
> >     int end = match[0].rm_eo;
> > 
> >     printf("Match found at: %d to %d\n", start, end);
> > 
> >     offset = start + 1;
> >   }
> > 
> > Gives me with glibc:
> > 
> > match[0].rm_so == 0
> > match[0].rm_eo == 0
> > Match found at: 0 to 0
> > match[0].rm_so == 4
> > match[0].rm_eo == 4
> > Match found at: 4 to 4
> > match[0].rm_so == 8
> > match[0].rm_eo == 8
> > Match found at: 8 to 8
> > 
> > Is this what you had in mind?  We should perhaps add it as an example to
> > the manual.
> 
> ah, yes, thanks. I didn't set match[0].rm_so and match[0].rm_eo before
> running regexec() with REG_STARTEND. So the conclusion is that if you
> want to use GNU regular expression extensions you also have to use
> non-POSIX flags in some cases.

Hi again Florian,

I was surprised to see that on Android both built-in Toybox and
Busybox installed in Termux give the expected output, that is:

:/data/user/0/org.galexander.sshd/files $ printf 'abc def xyz\n' | sed  's,\<,*,g'
*abc *def *xyz

Android uses Bionic libc and the part that implements regular
expressions is copied from BSD libc.

This example reproduces the exact same method that busybox sed uses to
find matches in a string without using REG_STARTEND:

#include <regex.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

void find_matches(char *pattern, char *text) {
  regex_t regex;
  regmatch_t match[10];
  size_t offset = 0;
  size_t len = strlen(text);
  int ret;

  ret = regcomp(&regex, pattern, REG_EXTENDED);
  if (ret) {
    fprintf(stderr, "regcomp %d\n", ret);
    return;
  }

  if (REG_NOMATCH == regexec(&regex, text, 10, match, 0)) {
    puts("no match");
    return;
  }

  do {
    int start = match[0].rm_so;
    int end = match[0].rm_eo;
    printf("Match found at: %d to %d\n", start, end);
    printf("text now == %s\n", text);
    text += end;

    if (start == end && *text) {
      text++;
    }

    if (*text == '\0') {
      break;
    }

  } while (regexec(&regex, text, 10, match, REG_NOTBOL) != REG_NOMATCH);

  regfree(&regex);
}

int main(void) {
  char *pattern = "\\<";
  char *text = "abc def xyz";

  find_matches(pattern, text);
}

With glibc output is:

Match found at: 0 to 0
text now == abc def xyz
Match found at: 0 to 0
text now == bc def xyz
Match found at: 0 to 0
text now == c def xyz
Match found at: 1 to 1
text now ==  def xyz
Match found at: 0 to 0
text now == ef xyz
Match found at: 0 to 0
text now == f xyz
Match found at: 1 to 1
text now ==  xyz
Match found at: 0 to 0
text now == yz
Match found at: 0 to 0
text now == z

But on FreeBD 12.0:

Match found at: 0 to 0
text now == abc def xyz
Match found at: 3 to 3
text now == bc def xyz
Match found at: 3 to 3
text now == ef xyz

It must be why busybox and toybox sed implementations show correct
output with \< pattern on Android. Is the behavior in glibc a bug or
intended?

-- 
Arkadiusz Drabczyk <arkadiusz@drabczyk.org>


More information about the Libc-help mailing list