SIMD optimization x64/x86: Better code for *64x2.splat
Categories
(Core :: JavaScript: WebAssembly, enhancement, P3)
Tracking
()
People
(Reporter: lth, Unassigned)
References
(Blocks 1 open bug)
Details
I don't know how much this matters, but it's come up (https://github.com/WebAssembly/simd/issues/191):
For i64x2.splat we use pinsrq from int registers to float registers twice, while optimally we might use movd (to fill the low lane) + pinsrq (to replicate the low lane to the high lane), thus avoiding crossing the alu / fpu boundary twice.
For f64x2.splat we use shufpd, while optimally we might use pinsrq to replicate the low lane into the high lane, the point would be to avoid the shuffle, which may be contended or just plain slow.
For these we might want to scan the optimization manual for advice.
Comment 1•4 years ago
|
||
Since bug 1749788:
For i64x2.splat, it is platform dependent. For x64, we use vmovq w/reg64, then vbroadcastq or vpunpcklqdq. For x86, we have to vmovd low dword, and vpinsrd for high dword, and replicate using punpcklqdq. Staying in alu boundaries.
For f64x2.splat, we use vmovddup -- the data stays in fpu
Description
•