<div dir="ltr"><div dir="auto">Hi, Alejandro (or anyone else interested),<div dir="auto"><br></div><div dir="auto">There's a discrepancy in the wording of the mbrtowc(3) function (and similarly, mbsrtowcs(3) function) between in POSIX and ISO C. It could be reported as an issue to POSIX (the Austin Group), and I am not sure if you can do that.</div><div dir="auto"><br></div><div dir="auto">In ISO C (I checked in both C99 and C23, in particular the N3220 draft), there's a statement that if mbrtowc() returns a (size_t)(-1) as an encoding error occurs, "the conversion state is unspecified".</div><div dir="auto"><br></div><div dir="auto">POSIX (see <<a href="https://pubs.opengroup.org/onlinepubs/9799919799/functions/mbrtowc.html" target="_blank">https://pubs.opengroup.org/onlinepubs/9799919799/functions/mbrtowc.html</a>>), for the same part it says "the conversion state is undefined".</div><div dir="auto"><br></div><div dir="auto">This wording difference matters when the "unspecified behavior" and "undefined behavior" are technically different. An example is how the mbstate_t object can be reused after an invalid sequence is encountered. When the state is said to be "undefined" it's implied to be not usable again (unless it is reset, e.g., by an `mbrtowc(NULL, "", 1, ps)` call). When it's "unspecified" then implementations can allow the state to be reused for certain encodings (possible for UTF-8, for example).</div><div dir="auto"><br></div><div>This is something I discovered accidentally when researching the multibyte functions in the C standard library and how they work with an encoding like UTF-8.</div></div>
</div>