<div dir="auto">-Os should optimize for code size. Other optimizations should<div dir="auto">take performance into account.</div></div><br><div class="gmail_quote"><div dir="ltr" class="gmail_attr">On Tue, Jun 18, 2024, 2:23 PM Jiang, Haochen <<a href="mailto:haochen.jiang@intel.com" target="_blank" rel="noreferrer">haochen.jiang@intel.com</a>> wrote:<br></div><blockquote class="gmail_quote" style="margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">> >> Wait. While the compiler may use PSRLDQ here, based on knowing<br>
> >> assumptions<br>
> >> made elsewhere, the assembler can't: The replacement insn must generate the<br>
> >> exact same result in the destination register. PSRLDQ with an immediate of<br>
> >> 0 (which effectively you're suggesting to use here) doesn't alter the<br>
> >> destination register at all, though. When really we want the upper bits of<br>
> >> the register cleared.<br>
> ><br>
> > pextrd/q also doesn't clear them at all. For vpextrd/q and vpsrldq, they will<br>
> > both clear higher bits. So they will be the same.<br>
> <br>
> Wait - your suggestion is even more confusing: The destination of PSRLDQ is<br>
> an XMM register, whereas the destination of PEXTR* is a GPR or memory. This<br>
> is properly expressed in the constraints in the compiler, but clearly we<br>
> can't replace insns like this in the assembler.<br>
<br>
Yes, I realized that I am wrong here, there are no constraints. vmovd/q would be<br>
definitely better and doable here if we would like to do something.<br>
<br>
> >>> Also, I suppose the optimization related to latency should not be done in<br>
> >>> assembler.<br>
> >><br>
> >> Why? We have -O, -O1, and -O2 alongside -Os for a reason.<br>
> ><br>
> > I am quite conservative on the optimization in assembler. If we are also going to<br>
> > optimize those hand-written code, the optimization could work.<br>
> ><br>
> > However, when they hand write some code, are we supposed to change them?<br>
> <br>
> Well, if we aren't to, people simply don't pass -O.<br>
> <br>
> > For -Os, we could give them all the optimizations we have, but for -O, I am not<br>
> > that sure.<br>
> ><br>
> > And I suppose we might add too much burden for the assembler if we are going<br>
> > to add too much optimizations related to latency. It will become another compiler.<br>
> > Are we supposed to copy all the optimizations from compiler?<br>
> <br>
> Probably not all (and many aren't the the insn level anyway, nor do we - so<br>
> far at least - optimize for latency/throughput at the expense of code size).<br>
> But yes - this specific aspect is why I keep raising questions on what<br>
> optimizations are worth it vs where we'd better leave code alone.<br>
<br>
H.J., what is your opinion on that?<br>
<br>
Thx,<br>
Haochen<br>
<br>
> <br>
> Jan<br>
> <br>
> > IMO, optimization to<br>
> > codesize is ok, but for latency, I am a little concerned.<br>
> ><br>
> > Thx,<br>
> > Haochen<br>
> ><br>
> >><br>
> >> Jan<br>
<br>
</blockquote></div>