[PATCH][AArch64] Optimized memcpy/memmove
Wilco Dijkstra
wdijkstr@arm.com
Fri Sep 25 13:16:00 GMT 2015
Further optimize memcpy/memmove for AArch64. Copies are split into 3 main cases: small copies of up
to 16 bytes, medium copies of 17..96 bytes which are fully unrolled. Large copies of more than 96
bytes align the destination and use an unrolled loop processing 64 bytes per iteration. In order to
share code with memmove, small and medium copies read all data before writing, allowing any kind of
overlap. All memmoves except for the large backwards case fall into memcpy for optimal performance.
On a random copy test memcpy/memmove are 40% faster on A57 and 28% on A53.
OK for commit?
ChangeLog:
2015-09-25 Wilco Dijkstra <wdijkstr@arm.com>
* sysdeps/aarch64/memcpy.S (memcpy):
Rewrite of optimized memcpy and memmove.
* sysdeps/aarch64/memmove.S (memmove): Remove
memmove code (merged into memcpy.S).
-------------- next part --------------
An embedded and charset-unspecified text was scrubbed...
Name: 0001-Optimized-memcpy.txt
URL: <http://sourceware.org/pipermail/libc-alpha/attachments/20150925/651df66a/attachment.txt>
More information about the Libc-alpha
mailing list