[PATCH 3/6] newlib: mem[p]cpy/memmove improve performance for optimized versions
Richard Earnshaw (lists)
Richard.Earnshaw@arm.com
Tue Jan 28 16:33:21 GMT 2025
On 28/01/2025 16:11, Corinna Vinschen wrote:
> On Jan 27 10:45, Alexey Lapshin wrote:
>> This change improves performance on memory blocks with sizes in range
>> [4..15]. Performance measurements made for RISCV machine (memset):
>>
>> size 4, CPU cycles change: 50 -> 37
>> size 5, CPU cycles change: 57 -> 40
>> size 6, CPU cycles change: 64 -> 47
>> size 7, CPU cycles change: 71 -> 54
>> size 8, CPU cycles change: 78 -> 44
>> size 9, CPU cycles change: 85 -> 47
>> size 10, CPU cycles change: 92 -> 54
>> size 11, CPU cycles change: 99 -> 61
>> size 12, CPU cycles change: 106 -> 51
>> size 13, CPU cycles change: 113 -> 54
>> size 14, CPU cycles change: 120 -> 61
>> size 15, CPU cycles change: 127 -> 68
>
> But is that generally true for other architectures as well?
>
No, it can be very dependent on the microarchitecture. I know of Arm implementations where it would be better and implementations where it would be (much) worse. The other variable is that for misaligned copies there's a choice of bringing the source data to alignment or the target data (you really don't want to do a large copy with both misaligned). That can also vary by micro-architecture.
But we have custom assembler versions for Arm, so it probably doesn't matter for us, except at -Os and there I wouldn't expect us to want large expanded chunks of code for all the cases that misaligned copies might involve.
R.
More information about the Newlib
mailing list