Closed Bug 1690533 Opened 5 years ago Closed 4 years ago

SIMD optimization x64/x86: Better code for *64x2.splat

Categories

(Core :: JavaScript: WebAssembly, enhancement, P3)

x86_64
All
enhancement

Tracking

()

RESOLVED DUPLICATE of bug 1749788

People

(Reporter: lth, Unassigned)

References

(Blocks 1 open bug)

Details

I don't know how much this matters, but it's come up (https://github.com/WebAssembly/simd/issues/191):

For i64x2.splat we use pinsrq from int registers to float registers twice, while optimally we might use movd (to fill the low lane) + pinsrq (to replicate the low lane to the high lane), thus avoiding crossing the alu / fpu boundary twice.

For f64x2.splat we use shufpd, while optimally we might use pinsrq to replicate the low lane into the high lane, the point would be to avoid the shuffle, which may be contended or just plain slow.

For these we might want to scan the optimization manual for advice.

Since bug 1749788:

For i64x2.splat, it is platform dependent. For x64, we use vmovq w/reg64, then vbroadcastq or vpunpcklqdq. For x86, we have to vmovd low dword, and vpinsrd for high dword, and replicate using punpcklqdq. Staying in alu boundaries.

For f64x2.splat, we use vmovddup -- the data stays in fpu

Status: NEW → RESOLVED
Closed: 4 years ago
Resolution: --- → DUPLICATE
You need to log in before you can comment on or make changes to this bug.