[PATCH 1/2] [PATCH 1/2] Enable Intel AVX512_FP16 instructions
Cui, Lili
lili.cui@intel.com
Fri Jul 9 11:47:25 GMT 2021
> -----Original Message-----
> From: Jan Beulich <jbeulich@suse.com>
> Sent: Monday, July 5, 2021 2:30 PM
> To: Cui, Lili <lili.cui@intel.com>; hjl.tools@gmail.com
> Cc: binutils@sourceware.org
> Subject: Re: [PATCH 1/2] [PATCH 1/2] Enable Intel AVX512_FP16 instructions
>
> On 01.07.2021 09:47, Cui,Lili wrote:
> > opcodes/i386-opc.tbl | 376 +++++++++++++++++++++
>
> A few more observations:
>
> While personally I think what you have is the best way of encoding it
> (allowing 64-bit register use to control EVEX.W in 64-bit bit mode), VMOVW
> is neither consistent with VPEXTRW (using EvexWIG) nor with VMOVD (only
> permitting Reg32). H.J., what are your thoughts here?
>
Removed Reg64 and add VexWIG on VMOVW.
> The VMOVW template needs splitting afaict: By OR-ing Word with
> Reg32 and/or Reg64, git s have separate register and
> memory operand templates, which is for this reason, iirc. (You may recall my
> other remark regarding combining e.g. Reg32 and Dword - there it is merely
> redundant, but having such is liable to suggest to people that combinations
> like
> Reg32 and Word are also okay. I intend to have i386-gen warn about such
> down the road, but obviously only once all prese/vnt redundancies have been
> eliminated.)
>
Done.
> VCVT{,T}SH2{,U}SI should have EvexWIG for their non-64bit encodings.
> But really it's unclear why each of them has three templates when the
> corresponding pre-existing SD and SS insns get away with two. I would have
> expected new templates to have been cloned from similar existing ones,
> rather than introducing new ones (with new inconsistencies). Of course
> there's (again) the possibility that you've spotted a bug with pre-existing
> templates, but then - if you don't want to fix those right away - I'd expect you
> to at least point out why you deviate from what we've got.
>
vcvtss2si has two templates because there is a special judgment in check_long_reg function, then the instruction can encode as EVEX.W = 1 without explicit VexW1.
if (intel_syntax
&& i.tm.opcode_modifier.toqword
&& i.types[0].bitfield.class != RegSIMD)
{
/* Convert to QWORD. We want REX byte. */
i.suffix = QWORD_MNEM_SUFFIX;
}
I add a special judgment in check_word_reg function, then VCVT{,T}SH2{,U}SI can also have two templates. I think that in order to reduce the number of templates and make the code less readable, this is a trade-off. I changed it anyway. Jan, what are your thoughts here?
else if (i.types[op].bitfield.qword
&& (i.tm.operand_types[op].bitfield.class == Reg
|| i.tm.operand_types[op].bitfield.instance == Accum)
&& i.tm.operand_types[op].bitfield.qword)
{
if (intel_syntax
&& i.tm.opcode_modifier.toqword
&& i.types[0].bitfield.class != RegSIMD)
{
/* Convert to QWORD. We want REX byte. */
i.suffix = QWORD_MNEM_SUFFIX;
}
}
> I suppose VCMP{P,S}H should have a large set of pseudos just like
> VCMP{P,S}{S,D} do, even if (for now) the spec doesn't spell those out.
>
> Jan
More information about the Binutils
mailing list