[Valgrind-developers] [PATCH 1/2] x86: optimize XCHG to MOV for same-register forms

H.J. Lu hjl.tools@gmail.com
Tue Jul 7 21:18:23 GMT 2026


On Tue, Jul 7, 2026 at 10:35 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 07.07.2026 16:21, Sam James wrote:
> > Jan Beulich <jbeulich@suse.com> writes:
> >
> >> On 06.07.2026 10:22, Sam James wrote:
> >>> Jan Beulich <jbeulich@suse.com> writes:
> >>>
> >>>> On 03.07.2026 15:32, H.J. Lu wrote:
> >>>>> On Fri, Jul 3, 2026 at 8:28 PM Jan Beulich <jbeulich@suse.com> wrote:
> >>>>>>
> >>>>>> On 03.07.2026 13:39, H.J. Lu wrote:
> >>>>>>> On Fri, Jul 3, 2026 at 6:04 PM Sam James <sam@gentoo.org> wrote:
> >>>>>>>>
> >>>>>>>> Jan Beulich <jbeulich@suse.com> writes:
> >>>>>>>>
> >>>>>>>>> On 03.07.2026 08:09, Paul Floyd wrote:
> >>>>>>>>>> On 2026-07-03 08:00, Jan Beulich wrote:
> >>>>>>>>>>> On 03.07.2026 06:55, Paul Floyd wrote:
> >>>>>>>>>>>> Would it be possible for gas to only do this transformation if the source and destination registers are different?
> >>>>>>>>>>> When the registers are different, this transformation is invalid to do.
> >>>>>>>>>>
> >>>>>>>>>> OK so you are optimising a no-op. Does GCC use it as a no-op?
> >>>>>>>>>
> >>>>>>>>> I don't expect so. In fact, my take is that -O... should not be used on
> >>>>>>>>> compiler generated code. The compiler should do whatever optimizations
> >>>>>>>>> are possible / sensible, and it should not emit code which can (easily)
> >>>>>>>>> further be optimized. (Easily because the assembler really only does
> >>>>>>>>> very simple and pretty obvious transformations.)
> >>>>>>>>
> >>>>>>>> Yes, that's reasonable. We should document it though.
> >>>>>>>
> >>>>>>> -O should be safe for compiler generated codes.
> >>>>>>
> >>>>>> The question isn't about "being safe". -O should be safe on whatever input.
> >>>>>> If it's not, it's a bug.
> >>>>>>
> >>>>>> The question is whether it is plausible to use -O... at all for compiler
> >>>>>> generated code. I causes extra overhead in the assembler, after all. If the
> >>>>>> compiler did a decent job, all of that extra overhead is going to be in
> >>>>>> vein. (As said elsewhere, the situation is different for code coming from
> >>>>>> asm() - that's not really compiler generated code.)
> >>>>>>
> >>>>>
> >>>>> From what we have learned so far, "XCHG REG,REG" has been done
> >>>>> on purpose and compilers never generate them automatically.   Assembler
> >>>>> should leave them alone even with encoding optimization.
> >>>>
> >>>> No, why? Optimization is specifically for hand-coded assembly, so what a
> >>>> compiler emits doesn't matter here. Following this argumentation of yours,
> >>>> we should remove all optimization again from gas. People can use any
> >>>> particular encoding "on purpose", after all. As said elsewhere, if you're
> >>>> after particular encodings, don't engage optimization in the first place.
> >>>
> >>> Let's please document that though.
> >>
> >> I'm having a hard time seeing what exactly you want documented. Likely not
> >> "optimization can change what you've written", as that entirely obvious. So
> >> please could you propose something that fits your desires?
> >
> > Just what you keep saying?
> >
> > "Assembler optimization is intended for hand-written code, not to be
> > used on compiler output." which then gives us licence to dispose with
> > bugs reports like this, as opposed to the ambiguity right now.
>
> I fear I wouldn't be happy with making such a statement in doc. For one,
> "intended" is weak enough that people may still think the options are
> worthwhile to use on compiler output. Plus there's the issue with inline
> assembly, which imo can plausibly be subject to optimization. Yet at
> this time we have no way to have optimization "kick in" only on those
> portions.
>
> I could perhaps live with a yet weaker version of what you suggest:
>
> "Assembler optimization is intended primarily for hand-written code.  If

I disagree.  I added -O to assembler for

On x86, some instructions have alternate shorter encodings:

1. When the upper 32 bits of destination registers of

andq $imm31, %r64
testq $imm31, %r64
xorq %r64, %r64
subq %r64, %r64

known to be zero, we can encode them without the REX_W bit:

andl $imm31, %r32
testl $imm31, %r32
xorl %r32, %r32
subl %r32, %r32

This optimization is enabled with -O, -O2 and -Os.
2. Since 0xb0 mov with 32-bit destination registers zero-extends 32-bit
immediate to 64-bit destination register, we can use it to encode 64-bit
mov with 32-bit immediates.  This optimization is enabled with -O, -O2
and -Os.
3. Since the upper bits of destination registers of VEX128 and EVEX128
instructions are extended to zero, if all bits of destination registers
of AVX256 or AVX512 instructions are zero, we can use VEX128 or EVEX128
encoding to encode AVX256 or AVX512 instructions.  When 2 source
registers are identical, AVX256 and AVX512 andn and xor instructions:

VOP %reg, %reg, %dest_reg

can be encoded with

VOP128 %reg, %reg, %dest_reg

This optimization is enabled with -O2 and -Os.
4. 16-bit, 32-bit and 64-bit register tests with immediate may be
encoded as 8-bit register test with immediate.  This optimization is
enabled with -Os.

These optimizations were intended for compiler generated assembly
codes.  But if we discovered such optimization has a significant
drawback, we should disable it unless its benefits outweigh its
drawback.

>  you notice any effects on compiler generated code, please consider
>  raising a bug against the compiler.  Note, however, that this does not
>  extend to there possibly being effects on code originating from inline
>  assembly."
>
> Jan



-- 
H.J.


More information about the Binutils mailing list