Closed Bug 1911591 Opened 2 years ago Closed 2 years ago

wasm lazy tiering: improvements to hotness-counter decrementation

Categories

(Core :: JavaScript: WebAssembly, enhancement, P3)

enhancement

Tracking

()

RESOLVED FIXED
132 Branch
Tracking Status
firefox132 --- fixed

People

(Reporter: jseward, Assigned: jseward)

References

Details

Attachments

(2 files)

Currently, a function's hotness counter is decremented by 1 at each entry to
the function and for each loop iteration within the function. This means large
branchy functions without loops only arrive slowly at their hotness threshold,
when it might have been better to tier them up earlier. The same problem
occurs more generally for any function that has long basic blocks, or sequences
thereof, and few loops.

This could be ameliorated by making the entry-check decrement amount be
proportional to the function length, and the loop-check amount proportional to
the size of the body. Not by any means a perfect solution, but more
representative than what we have at present.

For the entry-check, it's easy to find the function (bytecode) length. For
loop bodies, probably not so easy, since the check is at the start of the loop
and we're doing single-pass compilation.

Generalising the title a bit. There are two improvments to the decrementing of
hotness-counter that make sense to do together:

(1) A code generation change. Currently we have

   Reg32 tmp = load( instancePtr + counter_offset )
   tmp = tmp - 1, and set flags
   jmp-if-negative OOL
   store( instancePtr + counterOffset, tmp)
resume:
   ..
------
OOLcode:
   call the request-tier-up stub
   goto resume

The hot path turns into 4 insns on {x86,x64,arm32,arm64}. This is not bad, but
a refactoring like this

   Reg32 tmp = load( instancePtr + counter_offset )
   tmp = tmp - 1, and set flags
   store( instancePtr + counterOffset, tmp)
   jmp-if-negative OOL
   // everything else unchanged

would facilitate doing the hot path in 2 insns on {x86,x64} whilst not making
the arm{32,64} cases worse:

   subl $1, $offset(%r14)
   js OOL
resume:

Under the hood (of the cpu), this would likely turn into 4 micro-ops, so no
perf win there. The point is to reduce the code size, and to avoid trashing a
register. It does change the entry invariant that the request-tier-up stub
would expect -- previously, at entry, the counter will be <= 0, with this
change, it will be < 0. But that's OK.


(2) The step-size change as described in comment 0.

The entry-check decrement can easily scaled by the function body size since
we know that beforehand.

Scaling the loop-head checks is more tricky, because we don't know the loop
size until we encounter the closing end. Hence (a) we'd need a patchable
version of (1) and (b) we can find the loop body size (and hence patch the
constant) by noting the iterator offset when pushing the control item for
the loop and subtracting it from the iterator offset when we encounter
the closing end.

Limiting the step-size to 127 might allow a shorter-form encoding on x86/x64,
whilst giving adequate flexibility in step-down rates for very long blocks. A
bit of initial experimentation suggests that a step-down rate of one step per
20 bytecode bytes is plausible, which would allow linear behaviour out to a
block size of 2540 bytes.

Assignee: nobody → jseward
Summary: wasm lazy tiering: scale hotness-counter decrements by block size → wasm lazy tiering: improvements to hotness-counter decrementation

For x86/x64, we have to assume that the OOLcode is more than 127 bytes
forwards, hence no short form encoding for the js. That produces:

41 83 AE 88 13 00 00 7F     subl $127, 5000(%r14)  // 0x00001388 == 5000
0F 88 80 00 00 00           js 1f
= 14 bytes

or, if the step is not limited to 127

41 81 AE 88 13 00 00 80 00 00 00     subl $128, 5000(%r14)
0F 88 80 00 00 00                    js 1f
= 17 bytes

This patch adds a new MacroAssembler method,
sub32FromMemAndBranchIfNegativeWithPatch. It subtracts a constant value in the
range 1 .. 127 from a 32-bit memory location, writes the new value back to the
location, then jumps to a label if the updated value is negative.

The constant (1 .. 127) is not specified. It must be patched in afterwards
using new method patchSub32FromMemAndBranchIfNegative.

There are no uses of these methods in this patch. They are used in the next
patch in the series to decrement hotness counters for wasm baseline code. The
constant is patchable because we create a decrement at the start of each loop
body, and its constant depends on the loop body size, but we don't know what
that is until we get to the end of the loop.

The limitation to 1 .. 127 allows generating the shortest possible instructions
on Intel, and is good enough for hotness-counting.

For wasm lazy tiering, we need to discover when a (baseline-compiled) function
is hot enough to tier up. Each function has an int32_t counter associated with
it, which lives in the Instance (hence is fast to access).

At the start of the function's body, and also at the start of each loop body,
the counter is reduced by 1. If it goes negative then we call a stub method
which requests tier-up. These hotness checks are created by
BaseCompiler::addHotnessCheck.

This works, kind-of. But it puts functions with long basic blocks at a
disadvantage relative to those with short blocks -- they will appear to "heat
up" more slowly than those with short blocks; hence they will tier-up
relatively later than those with short blocks.

This patch "fixes" this, by reducing the counters not by 1 but instead by a
number between 1 and 127, which depends on the size of the function / loop
body, in wasm bytecode bytes. This is easy for the function-entry hotness
check, because we know the function size at that point. But for loops that is
more difficult because we don't know the size of the loop until its end
instruction, but the hotness count is at the start of the loop body. Hence the
hotness checks are now patchable.

Changes:

  • WasmBCClass.h, struct Control: new fields loopBytecodeStart and
    offsetOfCtrDec so we can know, at the end of a loop, how big its body is,
    and where we should patch the hotness check.

  • BlockSizeToDownwardsStep: given a function/loop size, decide on the downwards
    step size.

  • BaseCompiler::beginFunction: create variable downwards step instead of 1

  • BaseCompiler::addHotnessCheck: update to use new
    sub32FromMemAndBranchIfNegativeWithPatch macro from previous patch.

  • BaseCompiler::patchHotnessCheck: new method

  • BaseCompiler::emitLoop: collect loop start-point info

  • BaseCompiler::emitEnd: for loop ends, use info collected by
    BaseCompiler::emitLoop to patch the loop-head hotness check

  • WasmHandleRequestTierUp: add a release assertion about the value of the
    hotness counter at the point where tier-up is requested.

  • WasmBaselineCompiler.cpp: a new SMDOC describing the tier-up mechanism.

Pushed by jseward@mozilla.com: https://hg.mozilla.org/integration/autoland/rev/b64564985f2d part 1: add MacroAssembler::sub32FromMemAndBranchIfNegativeWithPatch. r=yury. https://hg.mozilla.org/integration/autoland/rev/e9af9e7c0cef part 2: variable-sized downwards steps in wasm lazy tiering hotness counting. r=yury.
Status: NEW → RESOLVED
Closed: 2 years ago
Resolution: --- → FIXED
Target Milestone: --- → 132 Branch
Blocks: 1920617
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: