Open
Bug 2037383
Opened 5 months ago
Updated 2 months ago
wasm wide-arithmetic: add a folding rule `addHI 0:x 0:y = carry bit out of x +64 y`
Categories
(Core :: JavaScript: WebAssembly, enhancement, P3)
Core
JavaScript: WebAssembly
Tracking
()
NEW
People
(Reporter: jseward, Unassigned)
References
(Blocks 1 open bug)
Details
There is probably some mileage to be had by adding a rule in
MWasmAddSubI128HI64::foldsTo to do addHI 0:x 0:y = carry bit out of x +64 y,
and then implement the carry-bit part using a new MIR/LIR for that purpose.
Clang on x86_64 generates
xorl %eax, %eax
addq %rsi, %rdi // assumes rdi trashable
setb %al
and gcc on arm64 generates
cmn x0, x1
cset x0, cs
This would reduce register pressure on both targets.
The above pattern appears 339 times in Alex Crichton's benchmark suite [1].
[1] https://github.com/alexcrichton/wasm-benchmark-i128
The hottest (and, really only, modulo loop unrolling) block for the blind-sig
test (where we do relatively worst in comparison with wasmtime) looks like this
for the two targets. The !s mark pointless zeroing of a register.
=-=-=-=-=-=-=-=-=-=-=-=-=-= begin SB rank 0 =-=-=-=-=-=-=-=-=-=-=-=-=-=
0: (197146489 14.27%) 197146489 14.27% 0x2648caa84400
==== SB 103386 (evchecks 0) [tid 0] 0x2648caa84400 UNKNOWN_FUNCTION UNKNOWN_OBJECT+0x0
(arm64) 0x2648CAA84400: ldr x6, [x21, x20]
(arm64) 0x2648CAA84404: umulh x25, x6, x2
(arm64) 0x2648CAA84408: madd x27, x6, x2, xzr
(arm64) 0x2648CAA8440C: ldr x6, [x21, x19]
(arm64) 0x2648CAA84410: add x24, x27, x6, lsl #0
(arm64) 0x2648CAA84414:! movz x11, 0x0
(arm64) 0x2648CAA84418: adds xzr, x27, x6, lsl #0
(arm64) 0x2648CAA8441C: adc x26, x25, x11
(arm64) 0x2648CAA84420: add x25, x15, x24, lsl #0
(arm64) 0x2648CAA84424:! movz x6, 0x0
(arm64) 0x2648CAA84428:! movz x11, 0x0
(arm64) 0x2648CAA8442C: adds xzr, x15, x24, lsl #0
(arm64) 0x2648CAA84430: adc x27, x6, x11
(arm64) 0x2648CAA84434: add x15, x27, x26, lsl #0
(arm64) 0x2648CAA84438: str x25, [x21, x19]
(arm64) 0x2648CAA8443C: add w19, w19, 0x8
(arm64) 0x2648CAA84440: add w20, w20, 0x8
(arm64) 0x2648CAA84444: sub w22, w22, 0x1
(arm64) 0x2648CAA84448: cbz w22, 0x2648CAA844EC
(19 insns total)
=-=-=-=-=-=-=-=-=-=-=-=-=-= end SB rank 0 =-=-=-=-=-=-=-=-=-=-=-=-=-=
=-=-=-=-=-=-=-=-=-=-=-=-=-= begin SB rank 0 =-=-=-=-=-=-=-=-=-=-=-=-=-=
0: (284889400 13.51%) 284889400 13.51% 0x7f92bf9f50b wasmION:fI=554:wasm_benchmark_i128-70205fabfd0db6a3.wasm._ZN14num_bigint_dig7biguint5monty10montgomery17h8df79c32e788b3b4E+1163
==== SB 104225 (evchecks 0) [tid 0] 0x7f92bf9f50b wasmION:fI=554:wasm_benchmark_i128-70205fabfd0db6a3.wasm._ZN14num_bigint_dig7biguint5monty10montgomery17h8df79c32e788b3b4E+1163 UNKNOWN_OBJECT+0x0
0x7F92BF9F50B: movq (%r15,%r10),%rax
0x7F92BF9F50F: movq %rax,%r14
0x7F92BF9F512: mulq %rcx
0x7F92BF9F515: movq %rdx,%rax
0x7F92BF9F518: movq %rax,-104(%rbp)
0x7F92BF9F51C: movq %r14,%rax
0x7F92BF9F51F: imulq %rcx, %rax
0x7F92BF9F523: movq %rax,%r14
0x7F92BF9F526: movq (%r15,%r12),%r9
0x7F92BF9F52A: movq %r14,%r13
0x7F92BF9F52D: addq %r9,%r13
0x7F92BF9F530:! xorl %edi,%edi
0x7F92BF9F532: movq %r14,%rax
0x7F92BF9F535: movq -104(%rbp),%rdx
0x7F92BF9F539: movq %rax,%rbx
0x7F92BF9F53C: addq %r9,%rbx
0x7F92BF9F53F: movq %rdx,%rbx
0x7F92BF9F542: adcq %rdi,%rbx
0x7F92BF9F545: movq %r8,%rax
0x7F92BF9F548: addq %r13,%rax
0x7F92BF9F54B:! xorl %edi,%edi
0x7F92BF9F54D:! xorl %r9d,%r9d
0x7F92BF9F550: movq %r8,%r14
0x7F92BF9F553: addq %r13,%r14
0x7F92BF9F556: movq %rdi,%r14
0x7F92BF9F559: adcq %r9,%r14
0x7F92BF9F55C: addq %rbx,%r14
0x7F92BF9F55F: movq %r14,-104(%rbp)
0x7F92BF9F563: movq %rax,(%r15,%r12)
0x7F92BF9F567: leal 8(%r12), %r8d
0x7F92BF9F56C: leal 8(%r10), %r13d
0x7F92BF9F570: leal -1(%rsi), %r12d
0x7F92BF9F574: testl %r12d,%r12d
0x7F92BF9F577: je-32 0x7F92BF9F58F
(34 insns total)
=-=-=-=-=-=-=-=-=-=-=-=-=-= end SB rank 0 =-=-=-=-=-=-=-=-=-=-=-=-=-=
Updated•2 months ago
|
Blocks: wasm-wide-arithmetic
You need to log in
before you can comment on or make changes to this bug.
Description
•