Running the summary task from about:inference using GPU takes 17s+9s on the Canvasthread (FXC?) and no result is generated
Categories
(Core :: Machine Learning: General, defect)
Tracking
()
People
(Reporter: mayankleoboy1, Unassigned, NeedInfo)
References
Details
run the summarizer task from about:inference. Select gpu as the device.
Click run.
Profile: https://share.firefox.dev/4fIFRIM
No result was generated. No error that i can see.
cc: egubler from webgpu perspective.
| Reporter | ||
Updated•1 year ago
|
| Reporter | ||
Comment 1•1 year ago
•
|
||
Still doesnt run to completion, but looks like the shader compilation time is reduced to less than a second.
https://share.firefox.dev/4alHnQ3
Comment 2•1 year ago
|
||
The severity field is not set for this bug.
:tarek, could you have a look please?
For more information, please visit BugBot documentation.
Comment 3•1 year ago
|
||
Thanks Mayank, can you confirm you have set dom.webgpu.enabled and dom.webgpu.workers.enabled to true.
| Reporter | ||
Comment 4•1 year ago
|
||
I retested right now and it didnt work.
Both dom.webgpu.enabled and dom.webgpu.workers.enabled are set to true
Profile: https://share.firefox.dev/40JZX0L
about:inference console output :
Creating engine if needed
Loading tokenizer.json from cache
Loading tokenizer.json from cache
Loading config.json from cache
Loading config.json from cache
Loading tokenizer_config.json from cache
Loading tokenizer_config.json from cache
Loading generation_config.json from cache
Loading generation_config.json from cache
Loading onnx/decoder_model_merged_quantized.onnx from cache
Loading onnx/decoder_model_merged_quantized.onnx from cache
Loading onnx/encoder_model_quantized.onnx from cache
Loading onnx/encoder_model_quantized.onnx from cache
Running inference request
Error: Internal error: worker terminated
Aside: consider adding some ML preset to about:logging
Comment 5•1 year ago
|
||
Thanks for your testing. Erich, do you have a bit of time to dig on this one with me?
Comment 6•1 year ago
|
||
Tarek: Sure! Hit me up in Matrix, and we can coordinate something. 🙂
Comment 7•1 year ago
|
||
I gave debugging a go myself, and I'm running into walls with debugging here, because DevTools doesn't have visibility into the mjs modules we're using to implement the inference engine. Just based on rg'ing into toolkit/components/ml/ in my Gecko checkout, it seems that we're using Transformers to implement this. The diagnostic states worker terminated— I wonder if the Transformers engine is trying (and failing) to use a module service worker (see bug 1360870)?
Otherwise, I haven't had good transparency into how WebGPU is being used here. Hoping that somebody from the ML side will be able to debug this more effectively.
Updated•1 year ago
|
Comment 8•1 year ago
|
||
When I activate tracing, and remove the 30s timeout, I am able to have the run finishing with no GPU errors.
The onnx backend produces warnings on some operations falling back to using CPU
console.error: "\x1B[0;93m2025-01-30 12:42:47.911366 [W:onnxruntime:, session_state.cc:1168 VerifyEachNodeIsAssignedToAnEp] Some nodes were not assigned to the preferred execution providers which may or may not have an negative impact on performance. e.g. ORT explicitly assigns shape related ops to CPU to improve perf.\x1B[m"
console.error: "\x1B[0;93m2025-01-30 12:42:47.911789 [W:onnxruntime:, session_state.cc:1170 VerifyEachNodeIsAssignedToAnEp] Rerunning with verbose output on a non-minimal build will show node assignments.\x1B[m"
We are seeing a lot of of calls to jsepCopyGpuToCpu and back and force of data between the CPU memory and the GPU memory because of this. That slows down execution and the problem worsen on int8 (probably because some dequantize operations are not on GPU)
I tried different quantization levels with GPU, using Xenova/long-t5-tglobal-base-16384-book-summary (long-t5 arch) with the default example, and this is the current state (on my M1 )
- gpu + fp32 produces correct results, and takes 140 seconds for the run
- gpu + q8 threads produces garbage results, and takes over 15mn (!) for the run
as a comparison:
- cpu + fp32 + 1 thread: 1741 ms
- cpu + fp32 + 4 threads: 818 ms
- cpu + q8 + 1 threads: 2924 ms
- cpu + q8 + 4 threads: 1127 ms
The list of supported kernel operations are listed here https://github.com/microsoft/onnxruntime/blob/main/js/web/docs/webgpu-operators.md
I need to dig but some operations are picking CPU and the memory copies make things extremely slow
On a the smaller models like Xenova/all-MiniLM-L6-v2 for feature detection,
on my M1:
- gpu + fp32 + 1 thread takes 86 ms
- cpu + fp32 + 1 thread takes 32 ms
- cpu + q8 + 1 thread takes 24 ms
- gpu + q8 + 1 thread takes 2536 ms
GPU gets faster for the model if we send a lot of input. for instance a batch of 200 tokens:
- gpu + fp32 takes 120ms
- cpu + fp32 + single threaded takes 512 ms
- cpu + q8 + single threaded takes 302 ms
- cpu + q8 + 4 threads takes 106 ms
- cpu + fp32 + 4 threads takes 152 ms
So the bottom line so far is that:
- WebGPU can be used only for fp32, because is extremely slow in int8 and won't work with fp16
- WebGPU gets faster than WASM on non generative models as th einput gets bigger. But when using several threads, WASM is comparable.
So for our current use cases, I don't think WebGPU is useful yet
Comment 9•1 year ago
|
||
for reference, jsepCopyGpuToCpu is here:
https://searchfox.org/mozilla-central/source/toolkit/components/ml/vendor/ort.webgpu-dev.mjs#15297
which downloads GPU data see
Updated•1 year ago
|
Updated•8 months ago
|
Comment 10•8 months ago
|
||
The severity field is not set for this bug.
:tburrell, could you have a look please?
For more information, please visit BugBot documentation.
Description
•