FunctionCompiler::closeLoop and fixupRedundantPhis dominates compile time on Wllama
Categories
(Core :: JavaScript: WebAssembly, task, P3)
Tracking
()
People
(Reporter: mayankleoboy1, Assigned: rhunt)
References
(Blocks 3 open bugs, )
Details
Go to https://github.ngxson.com/wllama/examples/main/dist/
Download and load the second model : "llama-3.2-1b-instruct-q4_k_m.gguf"
Profile: https://share.firefox.dev/4gZ9tmF
With the "Phi-3.1-mini-128k-instruct-Q3_K_M-(shards).gguf" model : https://share.firefox.dev/3Ch7uLm
Maybe something to improve?
On loading larger models, the proportion of time spent by that thread in hasFlags increases.
Comment 1•1 year ago
|
||
hasFlag is a constant-time operation, so spending that much time on it implies that we're doing it a lot. I suspect that this code is accidentally quadratic along some unexpected axis.
| Assignee | ||
Updated•1 year ago
|
| Assignee | ||
Updated•1 year ago
|
Comment 2•1 year ago
|
||
From lazy tiering logs, I see some very large functions being Ion-compiled, so
I'm not surprised if we've fall into yet another quadratic hole. sz is size in wasm
bytecode bytes.
MG::startPartialTier fI=1342 sz=24979 wasm-function[1342]
MG::startPartialTier fI=631 sz=150896 wasm-function[631]
MG::startPartialTier fI=462 sz=19218 wasm-function[462]
| Reporter | ||
Comment 3•1 year ago
|
||
Profile with the simple allocator: https://share.firefox.dev/4m0lXxk
| Reporter | ||
Updated•1 year ago
|
| Reporter | ||
Comment 4•1 year ago
•
|
||
This is what i get today: https://share.firefox.dev/42b2T6Y (1.1s on TC thread)
gemma-2-2b-it-abliterated-Q4_K_M-(shards).gguf: https://share.firefox.dev/4pfVWvA (1.4s on TC thread)
Phi-3.1-mini-128k-instruct-Q3_K_M-(shards).gguf : https://share.firefox.dev/4m6tdGL (1.2 s on TC thread)
Do we need to keep this open, considering the profiles highlight slowness in known areas?
Description
•