Create reproducible smoke tests for running llama.cpp locally and in CI
Categories
(Core :: Machine Learning: On Device, task, P1)
Tracking
()
| Tracking | Status | |
|---|---|---|
| firefox154 | --- | fixed |
People
(Reporter: jbowser, Assigned: jbowser)
References
(Blocks 1 open bug)
Details
(Whiteboard: [aiplatform])
Attachments
(2 files)
Make tests reproducible so we can do smoke tests on llama.cpp to test the follwing criteria
- KL divergence
- Semantic similarity?
- Determine whether they are Machine dependent.
| Assignee | ||
Updated•3 months ago
|
Updated•3 months ago
|
Updated•3 months ago
|
Comment 1•3 months ago
|
||
There is some trickiness here on artifact download. For the first test I think we should just re-use the status quo with the perftest setup. I looked up a bunch of files and constraints I'm aware of in a chat session, and I'm including them here as an FYI. I read it and it all seams reasonable.
AI Summary – Where to put llama.cpp smoke tests (within the current status quo)
Test file location
The existing llama.cpp / wllama browser tests live in toolkit/components/ml/tests/browser/:
browser_ml_llama_summarizer_perf.js— wllama summarization perf test; usessmollm2-360m-instruct-q8_0.ggufandQwen3-0.6B-Q8_0.gguf(lines 15-48)browser_ml_security_perf.js— wllama-based security perf test
A new smoke test would slot into the same directory. Reuse head.js (perfTest, initializeEngine, perfSetup) for the engine setup; the head already handles model-hub redirection and Remote-Settings stubbing as described in status-quo.md.
CI wiring
Add a new task to taskcluster/kinds/perftest/linux.yml (and mirror it across macosx.yml, windows11-24h2.yml, windows11-24h2-ref.yml, android.yml). Model the new task after ml-llama-summarizer-perf (linux.yml:500-523) or ml-security-perf (linux.yml:526-549) — both already fetch:
wllama.wasm— defined intaskcluster/kinds/fetch/onnxruntime-web-fetch.yml:13-19smollm2-360-instruct-gguf—onnxruntime-web-fetch.yml:93-100Qwen3-0.6B-GGUF—onnxruntime-web-fetch.yml:102-109
The task command must include --hooks toolkit/components/ml/tests/tools/hooks_local_hub.py; this is what populates MOZ_MODELS_HUB and lets head.js (initializeEngine at lines 342-377) point browser.ml.modelHubRootUrl at the on-disk model files. Without it head.js errors out (line 350-352).
Adding new GGUF models
If smoke-test criteria (KL divergence, semantic similarity, machine-dependence checks) require additional weights beyond the two GGUFs above, add new type: git entries to taskcluster/kinds/fetch/onnxruntime-web-fetch.yml following the existing pattern (e.g. onnxruntime-web-fetch.yml:102-109):
<task-name>:
description: ...
fetch:
type: git
repo: https://huggingface.co/<Org>/<Repo>
revision: <pinned-sha>
path-prefix: "onnx-models/<Org>/<Repo>/main/"
artifact-name: <name>.tar.zst
The path-prefix must start with onnx-models/ because hooks_local_hub.py:83 serves <MOZ_FETCHES_DIR>/onnx-models as the hub root.
Updated•3 months ago
|
Updated•3 months ago
|
Updated•3 months ago
|
| Assignee | ||
Comment 2•3 months ago
|
||
These tests exercise LlamaRunner via Mozilla/test-llama (TinyStories)
under mochitest, plus the real Link Preview SmolLM2 model behind an
opt-in MOZ_ML_LOCAL_DIR env var (served over httpd.js because file://
URLs aren't yet supported by createFileUrl in Utils.sys.mjs).
The TinyStories assertions catch the common ways a broken llama.cpp
build fails: crashes, stalled decoder, garbage tokenizer, degenerate
loops, prompt-insensitivity, non-determinism at argmax, and silent
changes in token selection (pinned golden hash).
| Assignee | ||
Comment 3•3 months ago
|
||
So, using greedy sampling with llama.cpp has produced deterministic enough results on the test that we can keep more complex testing to the models team. One thing that needs to be noted is that for the stories model, it wildly diverges based on architecture because the model is a toy and doesn't have enough parameters. SmolLM2 doesn't suffer from this but has to be run on the perftest harness. Therefore, a simple string comparison is a good enough test in this case, since we're testing if llama.cpp broke between updates and not the model itself.
It might be noting that if we see sub 1B models diverge based on architecture, that maybe should be an evaluation metric that we watch out for as far as which model architectures are actually suitable for OnDevice deployment.
| Assignee | ||
Comment 4•3 months ago
|
||
Phase B of Bug 2047025. Mirrors the five structural assertions from
browser_ml_llama_native_smoke.js (Phase A) against the real Link Preview
model (SmolLM2-360M Q8). The test lives under the perftest harness so
hooks_local_hub.py can serve the GGUF from MOZ_FETCHES_DIR/onnx-models/;
mochitest can't reach the model because its network sandbox blocks
external hosts and the GGUF is too large to vendor.
Key finding for the bug's "machine-dependence" criterion: SmolLM2's
greedy output is byte-identical across every architecture and OS we
measured — Linux x86_64, Windows x86_64 (standard and ref hardware),
macOS Intel x86_64 (verified once on try push d34f4ca…), and macOS
aarch64 (verified locally on Apple Silicon). Unlike TinyStories
(Phase A), where NEON-vs-AVX rounding flips argmax on a handful of
tokens, SmolLM2's top-1 logit margins are wide enough to absorb the
FP noise from any of the SIMD widths or math libraries (Apple
Accelerate included) currently in our CI fleet.
Test structure (parallels Phase A):
- Survives + metrics populated.
- Output is printable text with diverse tokens (catches NaN/garbage
tokenizer, degenerate loops). - Greedy is deterministic across two runs on the same engine.
- Engine is prompt-sensitive (catches "model returns a constant"
failure mode). - Greedy output matches a pinned EXPECTED_TEXT + SHA-256 hash
(catches silent quantization drift and upstream llama.cpp rolls).
Each task emits a heartbeat info() line before and after engine init so
the mochitest harness's "no output for 370s" outer timeout can't fire
during the slow GGUF mmap path.
Wiring:
- ml-llama-smollm2-smoke task across the three desktop platforms
that are not end-of-life:
taskcluster/kinds/perftest/{linux,windows11-24h2,
windows11-24h2-ref}.yml. Each fetches only smollm2-360-instruct-gguf
(native llama.cpp path; no wllama.wasm). - perftest.toml entry with the standard "disabled = ..." line so the
test only runs via mach perftest. - python/mozperftest/perfdocs/config.yml entry + regenerated
testing/perfdocs/generated/mozperftest.rst so the perfdocs lint job
stays green.
Deliberately not on macOS Intel 10.15:
Two reasons. (1) Native llama.cpp on macosx1015-64-qr was unstable
for TinyStories (third hash bucket + within-engine determinism
failures), which is why browser_ml_llama_native_smoke.js excludes
it via apple_catalina skip-if; macOS 10.15 is EOL. (2) The
releng-hardware/gecko-1-t-osx-1015 worker pool has multi-hour queue
depth on every push, so even a passing run there isn't useful
signal for everyday CI. SmolLM2 was verified once on this platform
in try push d34f4ca226327593d954ff706ae602d8ac8642af (passed with
the same hash as Linux/Windows x86_64), captured for the record.
macosx.yml carries a comment-marker where the task entry would go;
revisit if a macOS aarch64 perftest worker pool materialises.
Deferred (not in this bug):
- Android perftest task. Link Preview is desktop-only today; native
llama.cpp on Android needs its own validation. - Semantic similarity / KL divergence. The pinned text + hash catches
the regression class this bug cares about. KL divergence requires
JS-exposed logits, which LlamaRunner doesn't surface today and
warrants its own bug.
Updated•3 months ago
|
https://hg.mozilla.org/mozilla-central/rev/8397e1582f2e
https://hg.mozilla.org/mozilla-central/rev/16e218324a4e
Updated•2 months ago
|
Description
•