[PATCH v6 2/3] aarch64: clear ZA state of SME before clone and clone3 syscalls
Yury Khrustalev
yury.khrustalev@arm.com
Wed Oct 8 09:08:02 GMT 2025
On Tue, Oct 07, 2025 at 11:05:59AM -0300, Adhemerval Zanella Netto wrote:
>
> On 06/10/25 08:53, Yury Khrustalev wrote:
> > This change adds a call to the __arm_za_disable() function immediately
> > before the SVC instruction inside clone() and clone3() wrappers. It also
> > adds a macro for inline clone() used in fork() and adds the same call to
> > the vfork implementation. This sets the ZA state of SME to "off" on return
> > from these functions (for both the child and the parent).
> >
> > ...
> >
> > However, in certain cases, the kernel may ensure that a "child" doesn't
> > start in an incorrect state. This is what is done by the recent change
> > included in 6.16 kernel [5]. This is not enough to ensure that code that
> > uses clone() and clone3() function conforms to [1] when it runs on a
> > system that provides SME, hence this change.
>
> This makes old glibc without this fix non-compliant with the lazy saving scheme.
> I assume you will backport this from 2.39 (the first release with SME support)
> and forward,
Yes, this is the plan: to backport this patch to 2.39 and above as well
as the a7f6fd976c17b82dc198290b4ab7087f35855a0e change.
> but it does not help on older releases where HWCAP2_SME will still
> be advertised by the kernel.
HWCAP2_SME will only be advertised by 6.16 and above. If it happens to
be advertised for an earlier kernel version, it means it was forcibly
enabled during kernel build or the version of kernel is not up-to-date.
I don't think we can do anything with either of those cases.
> I am not sure if this would be a real problem; users of SME should rarely call
> clone/clone3/setjmp during a za operation (maybe some might hit this while running
> containers with old glibc on newer kernels). Rich has suggested masking off the
> hwcap SME bits in such a case [1], but it would require updating old release
> branches, which is also not optimal.
I agree, doing so would be optimising for a rare and unlikely case at
the expense of most users.
> I assume this would not be a problem for Android, as they typically sync their
> libc support with the kernel version.
>
> And I think this should not be a blocker for this change.
Agreed.
> > +/* Clear ZA state of SME (C version). */
> > +/* The __libc_arm_za_disable function has special calling convention
> > + that allows to call it without stack manipulation and preserving
> > + most of the registers. */
> > +#define CALL_LIBC_ARM_ZA_DISABLE() \
> > +({ \
> > + asm volatile ( \
> > + " mov x13, x30\n" \
> > + " .cfi_register x30, x13\n" \
> > + " bl __libc_arm_za_disable\n" \
> > + " mov x30, x13\n" \
> > + " .cfi_register x13, x30\n" \
> > + : : : "x13", "x14", "x15", "x16", "x17", "x18", "memory" ); \
> > +})
>
> I think there is no need to constrait the temporary to x13, we can let
> compiler pick it instead:
>
> #define CALL_LIBC_ARM_ZA_DISABLE() \
> ({ \
> unsigned long int __tmp; \
> asm volatile ( \
> " mov %0, x30\n" \
> " .cfi_register x30, %0\n" \
> " bl __libc_arm_za_disable\n" \
> " mov x30, %0\n" \
> " .cfi_register %0, x30\n" \
> : "=r" (__tmp) \
> : \
> : "x14", "x15", "x16", "x17", "x18", "memory" ); \
> })
Compiler might not know the calling convention of__libc_arm_za_disable().
It's a private Glibc function and a compiler might treat it as such
applying common rules. I think it's safer to force use of x13 that we
know will stay unchanged.
Yury
More information about the Libc-alpha
mailing list