Improve performance of wasm wide-arithmetic on x86_64
Categories
(Core :: JavaScript: WebAssembly, enhancement, P3)
Tracking
()
| Tracking | Status | |
|---|---|---|
| firefox152 | --- | fixed |
People
(Reporter: jseward, Assigned: jseward)
References
(Blocks 2 open bugs)
Details
Attachments
(2 files)
|
12.53 KB,
patch
|
Details | Diff | Splinter Review | |
|
48 bytes,
text/x-phabricator-request
|
Details | Review |
Our initial implementation of i64.{add,sub}128 and i64.mul_wide_{u,s} can be
slow, especially on x86_64, per numbers at [1]. We should investigate and
improve.
[1] https://github.com/WebAssembly/wide-arithmetic/issues/7#issuecomment-4292562703
| Assignee | ||
Comment 1•5 months ago
|
||
Reworks the x64 sequence for {M,L}MulI64WideHI64. Improves
performance for https://github.com/alexcrichton/wasm-benchmark-i128
on Intel Tiger Lake by 3.5%.
| Assignee | ||
Updated•5 months ago
|
| Assignee | ||
Comment 2•5 months ago
|
||
This patch:
-
Reworks the x64 sequence for MacroAssembler::wasmMulI64WideHI64 and hence for
{M,L}MulI64WideHI64. Specifically, MacroAssembler::wasmMulI64WideHI64 no
longer tries to preserve RAX and RDX; those registers are trashed and instead
the problem of preserving them is given to the Ion register allocator / to
the baseline "manual" allocator.This makes the sequences shorter and allows removal of the "xchg"
instruction, which (judging from Agner Fog's latency tables) is a limiting
factor on ILP.No change for any non-x86_64 targets.
-
Adds a
foldsTomethod for MWasmAddSubI128HI64. This doesn't see much
action, but can fold out0 + xandx + 0cases.
Performance for https://github.com/alexcrichton/wasm-benchmark-i128 on Intel
Tiger Lake is improved by 3.5%, mostly via improved IPC.
Updated•4 months ago
|
Updated•1 month ago
|
Description
•