[PATCH 0/5] x86/Intel: AVX512 syntax enhancements

Cui, Lili lili.cui@intel.com
Wed May 18 03:15:27 GMT 2022


> -----Original Message-----
> From: Jan Beulich <jbeulich@suse.com>
> Sent: Tuesday, May 17, 2022 8:00 PM
> To: Cui, Lili <lili.cui@intel.com>
> Cc: H.J. Lu <hjl.tools@gmail.com>; Binutils <binutils@sourceware.org>
> Subject: Re: [PATCH 0/5] x86/Intel: AVX512 syntax enhancements
> 
> > 1. If we use BCST instead {1to*}, it cannot directly reflect the broadcast
> number. When the register size is zmm, but broadcast number is not the
> same.
> >
> > -[      ]*[a-f0-9]+:[   ]*62 f5 54 58 58 31[     ]*vaddph zmm6,zmm5,WORD PTR
> \[ecx\]\{1to32\}
> > +[      ]*[a-f0-9]+:[   ]*62 f5 54 58 58 31[     ]*vaddph zmm6,zmm5,WORD
> BCST \[ecx\]
> >
> > -[      ]*[a-f0-9]+:[   ]*62 65 7d df 5b 72 80[          ]*vcvtph2dq
> zmm30\{k7\}\{z\},WORD PTR \[rdx-0x100\]\{1to16\}
> > +[      ]*[a-f0-9]+:[   ]*62 65 7d df 5b 72 80[          ]*vcvtph2dq
> zmm30\{k7\}\{z\},WORD BCST \[rdx-0x100\]
> 
> This case is clearly disambiguated by the destination register.
> What I think you're worried about are conversions where the field size
> shrinks (e.g. from 32 bits to 16 bits, like in vcvtdq2ph). In this case you will
> note that for the purpose of keeping things unambiguous the disassembler
> will continue to emit {1to<N>}, and the assembler will continue to require
> that extra bit of information.
> 

The format of appending {1to<N>} for vcvtdq2ph special case is great. 
There is no ambiguity for the format of vcvtph2dq zmm30{k7}{z},WORD BCST [rdx-0x100], but we cannot direct know the N ({1to<N>}) for this BCST format, although we can confirm it with the SDM. I just trying to say for the first impression, BAST format has this disadvantage.

> > 2. Just remove the last comma, it's ok for me, I remember FP16 has an
> instruction with {sae} on the middle position for the ATT format. But the intel
> format is placed at the end, I don't know if there is any problem.
> >
> > -[      ]*[a-f0-9]+:[   ]*62 f5 54 18 58 f4[     ]*vaddph zmm6,zmm5,zmm4,\{rn-
> sae\}
> > +[      ]*[a-f0-9]+:[   ]*62 f5 54 18 58 f4[     ]*vaddph zmm6,zmm5,zmm4\{rn-
> sae\}
> >
> > FP16:
> > vcvtusi2sh %edx, {rn-sae}, %xmm29, %xmm30 vcvtusi2sh
> > xmm6,xmm5,edx\{rn-sae\}
> 
> Well, yes, this is not only not a problem, but intended. See how the SDM
> places the rounding/SAE modifiers. It's also not FP16-specific in any way.
> 

Yes, SDM put the rounding/SAE behind the last register operand, if the last operand is immediate, it will put rounding/SAE before the immediate. But I don't quite understand why ATT format put it after %edx instead of before.

> > 3. This can reduce the  templates size, it is good to me.
> +Modrm|EVexLIG|Masking=3|EVexMap5|VexVVVV|VexW0|Disp8MemShift
> =1|No_bSu
> > +f|No_wSuf|No_lSuf|No_sSuf|No_qSuf|No_ldSuf|StaticRounding|SAE, {
> > +RegXMM|Word|Unspecified|BaseIndex, RegXMM, RegXMM }
> 
> I'm afraid I don't understand what you're trying to tell me here.
> Are you asking for some kind of change to be made to the patch(es)?

Please ignore it, no changes are required here, it's ok.

Thanks, 
Lili.


More information about the Binutils mailing list