[PATCH] aarch64: optimized memcpy implementation for thunderx2
Anton Youdkevitch
anton.youdkevitch@bell-sw.com
Tue Oct 2 14:46:00 GMT 2018
Szabolcs,
On 02.10.2018 14:21, Szabolcs Nagy wrote:
> On 01/10/18 23:42, Steve Ellcey wrote:
>> On Mon, 2018-10-01 at 19:22 +0300, Anton Youdkevitch wrote:
>>> +L(dst_unaligned):
>>> +Â Â Â Â Â Â Â /* For the unaligned store case the code loads two
>>> +Â Â Â Â Â Â Â Â Â Â aligned chunks and then merges them using ext
>>> +Â Â Â Â Â Â Â Â Â Â instrunction. This can be up to 30% faster than
>>> +Â Â Â Â Â Â Â Â Â Â the the simple unaligned store access.
>>> +
>>> +Â Â Â Â Â Â Â Â Â Â Current state: tmp1 = dst % 16; C_q, D_q, E_q
>>> +Â Â Â Â Â Â Â Â Â Â contains data yet to be stored. src and dst points
>>> +Â Â Â Â Â Â Â Â Â Â to next-to-be-processed data. A_q, B_q contains
>>> +Â Â Â Â Â Â Â Â Â Â data already stored before, count = bytes left to
>>> +Â Â Â Â Â Â Â Â Â Â be load decremented by 64.
>>> +
>>> +Â Â Â Â Â Â Â Â Â Â The control is passed here if at least 64 bytes left
>>> +Â Â Â Â Â Â Â Â Â Â to be loaded. The code does two aligned loads and then
>>> +Â Â Â Â Â Â Â Â Â Â extracts (16-tmp1) bytes from the first register and
>>> +Â Â Â Â Â Â Â Â Â Â tmp1 bytes from the next register forming the value
>>> +Â Â Â Â Â Â Â Â Â Â for the aligned store.
>>> +
>>> +Â Â Â Â Â Â Â Â Â Â As ext instruction can only have it's index encoded
>>> +Â Â Â Â Â Â Â Â Â Â as immediate. 15 code chunks process each possible
>>> +Â Â Â Â Â Â Â Â Â Â index value. Computed goto is used to reach the
>>> +Â Â Â Â Â Â Â Â Â Â required code. */
>>> +
>>> +Â Â Â Â Â Â Â /* Store the 16 bytes to dst and align dst for further
>>> +Â Â Â Â Â Â Â Â Â Â operations, several bytes will be stored at this
>>> +Â Â Â Â Â Â Â Â Â Â address once more */
>>> +       str     C_q, [dst], #16
>>> +       ldp     F_q, G_q, [src], #32
>>> +       bic     dst, dst, 15
>>> +       adr     tmp2, L(load_and_merge)
>>> +       add     tmp2, tmp2, tmp1, LSL 7
>>> +       sub     tmp2, tmp2, 128
>>> +       br      tmp2
>>
>> Anton,
>>
>> As far as the actual code, I think my only concern is this use of a
>> 'computed goto' to jump to one of the extract sections. Â It seems very
>> brittle since a change in the alignment of the various sections or a
>> change in the size of those sections could mess up this jump. Â Would
>> the code be any slower if you used a jump table instead of a computed
>> goto?
>
> is the 16byte alignment really needed (i.e. 8byte is not enough)?
> the code is fairly big with 16 alignment cases.
Unfortunately, yes. As the code deals with 16-bytes chunks the
optimal results are with the memory addresses that are 16 bytes
aligned.
> the indirect jump may be difficult to predict in real workloads.
> otherwise the computed jump is acceptable, just document how
> many instructions one entry can have at most (32?) so it's less
> brittle in case somebody tries to modify the code.
Like I answered Steve this is probably more or less the same
performance-wise. I will change the code to use jump table
and rerun the bencharks (I don't expect them to be different,
though).
More information about the Libc-alpha
mailing list