[PATCH 1/2] [PATCH 1/2] Enable Intel AVX512_FP16 instructions

Cui, Lili lili.cui@intel.com
Fri Jul 9 11:47:25 GMT 2021


> -----Original Message-----
> From: Jan Beulich <jbeulich@suse.com>
> Sent: Monday, July 5, 2021 2:30 PM
> To: Cui, Lili <lili.cui@intel.com>; hjl.tools@gmail.com
> Cc: binutils@sourceware.org
> Subject: Re: [PATCH 1/2] [PATCH 1/2] Enable Intel AVX512_FP16 instructions
> 
> On 01.07.2021 09:47, Cui,Lili wrote:
> >  opcodes/i386-opc.tbl           | 376 +++++++++++++++++++++
> 
> A few more observations:
> 
> While personally I think what you have is the best way of encoding it
> (allowing 64-bit register use to control EVEX.W in 64-bit bit mode), VMOVW
> is neither consistent with VPEXTRW (using EvexWIG) nor with VMOVD (only
> permitting Reg32). H.J., what are your thoughts here?
> 
Removed Reg64 and add VexWIG on VMOVW.

> The VMOVW template needs splitting afaict: By OR-ing Word with
> Reg32 and/or Reg64, git s have separate register and
> memory operand templates, which is for this reason, iirc. (You may recall my
> other remark regarding combining e.g. Reg32 and Dword - there it is merely
> redundant, but having such is liable to suggest to people that combinations
> like
> Reg32 and Word are also okay. I intend to have i386-gen warn about such
> down the road, but obviously only once all prese/vnt redundancies have been
> eliminated.)
> 
Done.

> VCVT{,T}SH2{,U}SI should have EvexWIG for their non-64bit encodings.
> But really it's unclear why each of them has three templates when the
> corresponding pre-existing SD and SS insns get away with two. I would have
> expected new templates to have been cloned from similar existing ones,
> rather than introducing new ones (with new inconsistencies). Of course
> there's (again) the possibility that you've spotted a bug with pre-existing
> templates, but then - if you don't want to fix those right away - I'd expect you
> to at least point out why you deviate from what we've got.
> 
vcvtss2si has two templates because there is a special judgment in check_long_reg function, then the instruction can encode as EVEX.W = 1 without explicit VexW1.

if (intel_syntax
    && i.tm.opcode_modifier.toqword
    && i.types[0].bitfield.class != RegSIMD)
          {
            /* Convert to QWORD.  We want REX byte. */
            i.suffix = QWORD_MNEM_SUFFIX;
          }

I add a special judgment in check_word_reg function, then VCVT{,T}SH2{,U}SI can also have two templates. I think that in order to reduce the number of templates and make the code less readable, this is a trade-off. I changed it anyway. Jan, what are your thoughts here?

    else if (i.types[op].bitfield.qword
             && (i.tm.operand_types[op].bitfield.class == Reg
                 || i.tm.operand_types[op].bitfield.instance == Accum)
             && i.tm.operand_types[op].bitfield.qword)
      {
        if (intel_syntax
            && i.tm.opcode_modifier.toqword
            && i.types[0].bitfield.class != RegSIMD)
          {
            /* Convert to QWORD.  We want REX byte. */
            i.suffix = QWORD_MNEM_SUFFIX;
          }
      }


> I suppose VCMP{P,S}H should have a large set of pseudos just like
> VCMP{P,S}{S,D} do, even if (for now) the spec doesn't spell those out.
> 
> Jan



More information about the Binutils mailing list