[PATCH 0/7] stdio-common: Verify printf format tests with Python rather than AWK

Adhemerval Zanella Netto adhemerval.zanella@linaro.org
Thu Aug 27 16:03:09 GMT 2026



On 24/08/26 19:16, Matt Turner wrote:
> The formatted printf output tests verify their records against GNU AWK, on
> the grounds that it supplies an implementation of format processing that is
> independent from ours.  It does not do that as cleanly as one would like.
> 
> AWK has to run in the bignum mode for the floating-point conversions,
> because otherwise it uses the system sprintf(3) internally and we end up
> verifying our code against itself.  That in turn makes gawk compiled with
> MPFR support a requirement for testing the library at all.  Beyond that
> gawk mishandles a number of cases which the AWK script then has to undo by
> hand: the extraneous leading 0 produced for the alternative form with the
> octal conversion, the 0 produced where no characters are expected for the
> hexadecimal conversions, the missing + and space characters for a zero
> value with the precision of zero, and a collection of sign, flag and field
> width anomalies for Inf and NaN values.  Each such workaround suppresses
> whatever we might get wrong in the same place.  The a, A, b, and B
> conversions cannot be verified at all, because gawk either does not handle
> them or produces different output.  And where gawk was built without the f
> and F conversions the tests reported unsupported rather than checking
> anything.

Hi Matt,

Thanks for working on this. It turned out that using awk as a printf oracle
was not the best approach, and the -M being on parole is not a good indication
to keep relying on it (even though there is no actual plan to remove it). 

The Python multi-precision is also builtin, so one less conditional knob to
kind in mind during testing.

But I think the main advantage of this approach is it increases coverage
(stdio-common goes from 1515 tests to 1611) and it is way faster (on my
box stdio-common tests time is cut in half).

> 
> The first patch replaces the AWK script with an equivalent one written in
> Python, which is already a requirement for building the library.  Rather
> than calling into any formatting routine it computes the reference output
> directly, using exact integer and rational arithmetic.  Working exactly
> means the result does not depend on the range or precision of any host
> floating-point type, so the wider types are handled without
> arbitrary-precision arithmetic having to be built into the interpreter, and
> none of the workarounds listed above are needed: the corner cases they
> cover are computed correctly.  The capability probes go away along with the
> AWK script, so f and F are now always verified, and the note on MPFR is
> dropped from the installation instructions.

I though we had Python as requirement only for testing, but we now do use
for gen-as-const-headers since 7e1d42400c1b8f03316fe14176133c8853cd3bbe.
> 
> The rest of the series spends the freedom that buys:
> 
>   2  stops iterating over the huge width for eleven of the twelve printf
>      family functions -- see below;

This can be pushed independently. 

>   3  adds the b and B conversions, which gawk does not handle;
>   4  adds the a and A conversions, which gawk renders differently;
>   5  adds subnormal values, which needed the minimum exponent for the type
>      to be reported to the verification;
>   6  fixes a printf_fp bug the added coverage exposed: with the alternative
>      form, a g or G conversion that rounds into a new decade dropped the
>      trailing zeros it is required to keep, so that printf ("%#.2g", 99.9)
>      gave "1.e+02" rather than "1.0e+02";

This can done independently and I think it would be useful to open a bug
report to track it.

>   7  marks a and A unsupported where long double has the IBM extended
>      format, which is a pair of doubles rather than a significand of a
>      single fixed width and so cannot be modelled the way the verification
>      does it.
> 
> This changes both what the tests cover and what they cost, so here are the
> numbers.  Measured on x86_64-linux-gnu with the tests run serially,
> processor time for the tst-printf-format-* tests, which varies by under 2%
> between runs:
> 
>                                  tests  results      records      CPU
>   before the series                288      576   32,588,304     935s
>   1  verify with Python            288      576   32,588,304     684s
>   2  huge width for printf only    288      576   27,563,064     375s
>   3  add b and B                   288      576   33,879,864     392s
>   4  add a and A                   336      672   37,258,224     402s
>   5  add subnormal values          336      672   38,609,568     411s
>   6  keep %#g trailing zeros       336      672   39,960,912     416s
>   7  skip a and A for IBM          336      672   39,960,912     416s
> 
> So the suite ends up checking 23% more records, spread over 48 more tests,
> in 2.25 times less processor time.  All 672 results pass, as do all 576
> before the series.
> 
> Two things account for the time.  Python turns out to be faster than gawk
> in the bignum mode to begin with, once rendered digits are memoized per
> value, which is what patch 1 alone buys; without the memoization the exact
> arithmetic makes the long double conversions slower than AWK instead.  The
> rest is patch 2.  HUGE_WIDTH is chosen so that none of the strings produced
> are truncated, which for the floating-point types means every record it
> takes part in carries hundreds or thousands of digits, and those records
> dominate the cost of this whole family of tests.  The digits are produced
> by the same conversion code whichever of the printf family of functions is
> used; what differs between the twelve of them is the sink the result is
> written to, which the smaller widths cover already.  Iterating over the
> huge width for printf alone therefore gives up no coverage worth having.

On my box the main difference of stdio-common tests time while running
in parallel was the 'stdio-common: Iterate over the huge width for one printf
function only'.

But at same time, with the new Python testing framework it does enable
a bunch a more tests.

> 
> Records are not equal in cost, which is why the two columns move
> independently: the 15% of records patch 2 drops are the most expensive in
> the suite and take 45% of the run time with them, while the b and B records
> patch 3 adds are integer ones, and the a and A records patch 4 adds are
> short as well, the hexadecimal significand having a fixed width.  What
> remains is mostly intrinsic -- the f and F conversions print the whole
> integer part regardless of the precision requested, so LDBL_MAX runs to
> some 4932 digits even at MID_WIDTH.
> 
> Patch 7 costs nothing here, as expected: powerpc64le is the one target with
> the IBM extended format and it builds everything with the IEEE format
> instead, so the conversions are only skipped where shared libraries have
> been turned off.
> 
> Tested on x86_64-linux-gnu and powerpc64le-linux-gnu, and with the long
> double conversions forced to the IBM extended format, where the a and A
> conversions report unsupported and the rest continue to pass.

So, as a side note, I started to check whether using mpfr/gmp directly would 
result in a simpler way to check this. The mpfr/gmp also had some corner cases
that we need to handle (like no standard %a/%A support, missing sign for
infinities and NaNs), and it required adding an additional build step in
build-many-glibcs.py to enable it.

It turned out to *not* be simpler than this approach. I will review the
rest of this patchset.


More information about the Libc-alpha mailing list