Canonical document. Mirrored across four sibling WebGPU/WGSL research projects:
webgpu-q— quantum chemistrywebgpu-dna— radiation track-structure / radiobiologyzero-tvm— Phi-3 LLM inference (hand-written WGSL, head-to-head vs WebLLM)neuropulse— live 1:1 LLM forward-pass visualization (Phi-3, 3.8B params)
Edit any one and propagate. Project-specific examples in §§ 1, 6, 7, 8, 10 diverge per repo; sections 2–5, 9, 11–15 are universal.
This is the discipline that makes the work publishable in JOSS, citable
years later, and reproducible by reviewers on different hardware. The
patterns matured in different repos and back-port / forward-port between
them (research-grade artifact discipline first in webgpu-dna, the
"falsify before shipping" CPU pre-screen first in zero-tvm —
which validates every kernel against a CPU reference before the WGSL
even reaches the GPU — automated doc-vs-code drift detection in
neuropulse, full porting framework in webgpu-q). Future siblings
inherit the union.
Umbrella thesis: every advanced physics simulation in the world should ship as a URL. The browser/WebGPU layer is what's novel; the chemistry/physics/model architecture is textbook. Hand-write only the novel layer; port everything with a peer-reviewed reference.
For zero-tvm specifically, the novel contribution is the
hand-written WGSL stack — proving you can match TVM's autotuned
kernels in 37 files of WGSL and 2k lines of TypeScript, by
out-fusing TVM's default pipeline. The Phi-3 architecture, MLC weight
format, and BPE tokenizer are textbook and ported.
All measured numbers for zero-tvm live in one canonical place:
BENCH.md— every head-to-head vs WebLLM (dispatches/token, tokens/s, JS bundle, total kernels, WGSL LOC, KV-cache mode, identical hardware/weights). Every cell links to the underlying artifact intests/results/.README.md§ What's actually in the box — derived fromBENCH.md, cross-linked back to source files in the repo.
Anywhere else (CLAUDE.md, RESEARCH.md, index.html, badges,
landing page) may summarize numbers but never introduce new ones.
If a number isn't in BENCH.md, it isn't measured.
Before stating a measurement anywhere:
protocol → run head-to-head against WebLLM on identical hardware → commit JSON artifact → add BENCH.md row → quote
Not the other way around. Every claim is meant to be head-to-head
falsifiable by anyone who runs the same hardware against
zerotvm.com vs webllm.mlc.ai.
Path: tests/results/YYYY-MM-DD/<id>.json.
Shape (locked; don't add top-level keys without updating the harness):
{
"meta": { "protocol": "...", "hypothesis": "...", "passBar": "...",
"seed": "named-seed-id", "warmup": 5, "trials": 20 },
"env": { "gitSha": "...", "userAgent": "...", "adapter": {...},
"limits": {...}, "timestamp": "2026-05-14T...",
"shaderHashes": {"matmul_wgsl": "...", "qkv_fused_wgsl": "...",
"attention_wgsl": "...", "fused_ffn_wgsl": "...",
"add_norm_wgsl": "..."} },
"rows": [ { /* per-cell measurements, including head-to-head WebLLM cells */ } ],
"status": "pass" | "fail" | "noisy" | "partial",
"diagnosis": "first-failing-cell + smoking-gun explanation"
}Re-runnable deterministically given fixed seed + identical GPU + same
shader hash. fp16 GEMM and fp16 reductions are NOT order-deterministic
across GPU vendors — same WGSL on different hardware (Apple Metal vs
Nvidia Vulkan vs Intel iGPU) yields statistically equivalent logits
(top-k mass within ε of HF reference) but not bit-exact;
shaderHashes lets reviewers group rows correctly.
pass— meets the protocol's pass bar (top-k match vs HF reference logits, throughput within band of WebLLM, etc.).fail— doesn't. Commit anyway with adiagnosisfield naming the first failing cell and the smoking gun. Never silently rerun until pass.noisy—std/median > 0.1on any cell. Informational, not pass/fail.partial— some cells pass, others don't; explicitN of Mcount in the diagnosis.honest negative— failures that are evidence.CLAUDE.md § Known gapscites the artifact and the rejected hypothesis. The 2026-06 "22% behind WebLLM" headline was an honest negative encoded as a pass-bar miss with documented root cause; it stood until the 2026-07-25 M2 Max re-measurement (correctness fixes + vec4-load defaults) flipped it ahead, and the 2026-07-30 corrected protocol restated it as +16% total / +31% decode — the negative's full arc is preserved in BENCH.md, not retconned out.
Honest negatives become the project's evidence base. They are not bugs to fix; they are findings.
Math.random()is banned in any test/experiment path. Sampling uses argmax (deterministic) or seeded top-p with a named seed. WGSL random draws (none in the current forward pass) would use a uniform-routed seed channel.- Every JSON artifact records: git SHA (when available), full
navigator.userAgent,adapter.info, WebGPUlimits, UTC ISO8601 timestamp, shader-file SHA-256 / git-rev-parse hashes for each of the 10 kernel roles (37 WGSL files counting A/B/int8 variants). - 5 warmup samples are discarded; 20 trials retained.
- Report median + p10/p90/p99 + std + IQR for throughput measurements — never single-shot.
- If
std/median > 0.1on any cell → label the artifact"noisy".
performance.now() deltas around queue.submit alone are fiction —
WebGPU is asynchronous. Mandatory pattern: a mapped readback of a
tiny buffer (a single f32) before AND after the work. The throughput
counter shown live in the chat UI uses this pattern; the offline
benchmark harness in tests/perf/ likewise. WebLLM's reported
throughput uses the same pattern, so head-to-head comparisons are
apples-to-apples.
Match against more than one reference frame. Listed in increasing sophistication / decreasing strength:
- Closed-form invariants: norm preservation across RMSNorm, softmax mass = 1 across attention, KV-cache append idempotence, tokenizer round-trip identity on UTF-8 corpora.
- CPU pre-screen: every WGSL kernel has a TypeScript CPU
reference (plain-JS reimplementations inside
tests/kernels/run.mjs; thesrc/cpu/directory this used to name never existed) that runs on the same inputs. The CPU path catches algorithmic bugs before the WGSL even reaches the GPU — this iszero-tvm's back-port-worthy contribution to the sibling discipline. - Peer-reviewed reference packages:
- HuggingFace
microsoft/Phi-3-mini-4k-instructin PyTorch fp16 as the bit-comparable reference for logits (top-k mass match within ε of fp16 noise floor). - MLC q4f16_1 quantized weights as the bit-comparable reference
for weights. Both
zero-tvmand WebLLM read these from the same OPFS cache. - WebLLM (
mlc-ai/web-llm) as the head-to-head reference — same hardware, same weights, same prompt. The head-to-head headline (+16.0% total / +31.4% decode as of the 2026-07-30 M2 Max run; ~28% ahead on 2026-07-25; 22% behind before that) is the falsifiable claim. Both metrics are always quoted: total wall-clock and decode-only. So is the counter-result — WebLLM wins time-to-first-token on short prompts.
- HuggingFace
- Experiment: human-readable output equivalence on a fixed
prompt corpus, scored against WebLLM's output for "did this
answer the question equally well."
UNIMPLEMENTED as of 2026-08-10. Nothing in the repo does this.
The nearest thing that exists is the five-prompt lexical battery in
tests/e2e/zero-tvm.test.ts:160-167, which checks for "paris", "42", "len", "yes" and two colour words on 4 of 9 shipped models, and which CI cannot run. Seedocs/QUALITY.md.
Multiple independent reference frames > one. Each artifact should
state which it's checking against in meta.hypothesis.
This is the architectural rule. The differentiator of zero-tvm is
the hand-written, fused WGSL kernel stack — proving you can match
WebLLM's TVM-autotuned emission with 10 hand-written kernel roles by
out-fusing the default pipeline. (State that carefully: 85 is what a whole
WebLLM session dumps; RESEARCH.md:275 counts ~11 active on the decode path,
which is the path the comparison is about. Quoting 85 against 10 overstates by
~8x, in our favour, against our own writeup.) So:
- Hand-written and owned (the contribution):
- All 10 WGSL kernel roles / 37 files:
qkv_fused.wgsl(Q/K/V proj- RoPE + paged-KV append in one dispatch),
attention.wgsl(paged attention + page-table read),fused_ffn.wgsl(gate + up - SiLU),
add_norm.wgsl(residual + RMSNorm),matmul.wgsl(int4-dequant GEMM + subgroup/tiled variants),rmsnorm.wgsl,embed.wgsl,sample.wgsl, etc.
- RoPE + paged-KV append in one dispatch),
- The 228-dispatch (f16 KV) / 260-dispatch (int8 KV) per-token schedule that beats TVM's 342.
- All ~2k lines of TypeScript engine, tokenizer, and weight loader.
- The OPFS cache layer over HuggingFace's CDN.
- All 10 WGSL kernel roles / 37 files:
- Ported from peer-reviewed source with attribution:
- Phi-3 architecture spec (32 layers / 32 attention heads /
32 KV heads / 96 head dim / SwiGLU / RoPE / RMSNorm /
grouped-query attention) from Microsoft's released
Phi-3-mini-4k-instructmodel card andconfig.json. - MLC q4f16_1 quantization scheme and
ndarray-cache.jsonweight format frommlc-ai/mlc-llm/mlc-ai/web-llm, sozero-tvmand WebLLM read the same on-disk weights. - BPE tokenizer patterns from
huggingface/tokenizers/sentencepiece. - RoPE / GQA / SwiGLU reference numerics from the original papers (Su et al. 2021 RoFormer, Ainslie et al. 2023 GQA, Shazeer 2020 GLU).
- Apache-TVM kernel patterns referenced (not copied) as the head-to-head baseline; the comparison is precisely against that pipeline's emission.
- Phi-3 architecture spec (32 layers / 32 attention heads /
32 KV heads / 96 head dim / SwiGLU / RoPE / RMSNorm /
grouped-query attention) from Microsoft's released
Per-file header for ported code:
// Ported from <upstream> (<upstream-url>), <license> license.
// Source: <relative-path> at commit <SHA>
// Original authors: <upstream/AUTHORS>
// Adaptations for zero-tvm:
// - <substantive change 1>
// - ...
// See LICENSE-<UPSTREAM> at repo root for the <license> notice.
Repo-level: LICENSE-MLC and LICENSE-PHI3 at root (verbatim
from upstream). Per-module status table belongs in a MIGRATION.md
table:
| module | reference | license | status |
|---|---|---|---|
tokenizer.ts BPE |
huggingface/tokenizers |
Apache 2.0 | 🟢 |
| weight loader | MLC ndarray-cache.json format |
Apache 2.0 | 🟢 |
| Phi-3 architecture | Microsoft Phi-3 model card / config | MIT | 🟢 |
| WGSL kernels | hand-written (this repo) | MIT | n/a |
License compatibility: MIT + Apache 2.0 work together — the ported portion keeps its upstream license obligations (notice + state changes); the rest of the repo (including all hand-written WGSL) stays MIT.
Any tunable scalar in production code that isn't backed by a peer-reviewed source is:
- Labeled empirical in the code comment at point of use.
- Documented in
CLAUDE.md§ Known gaps with the magnitude of the empirical correction and what observable it was tuned against. - Queued for removal once the structural fix lands.
- Tracked in
CHANGELOG.md/ commit messages when added and when removed.
zero-tvm aims for zero fudge factors in the forward pass —
every numeric scale (head-dim scale, RMS epsilon, RoPE base, softmax
temperature, int4 dequant zero-point convention) is sourced from
Phi-3's released config and the MLC quantization spec. The
empirical knobs are kernel-tile sizes (subgroup vs scalar, A/B
variants of matmul.wgsl), which are auto-selected at init based
on adapter.info — not hidden physics constants.
Tested-and-rejected hypotheses (e.g., "fusing all of qkv_fused
will exceed register pressure on Apple M2" — falsified by the
228-dispatch result) go into the same documents so future sessions
don't re-test them.
Every artifact records the SHA-256 (or git rev-parse <gitSha>:<path>
short hash) of each of the 37 WGSL shader files the experiment
depended on. This lets reviewers group rows by shader version when a
kernel implementation changes (subgroup tile size, int4 dequant
strategy, fused-vs-unfused FFN, A/B variant selection, …).
The env block carries shaderHashes: { matmul_wgsl: "...", qkv_fused_wgsl: "...", attention_wgsl: "...", fused_ffn_wgsl: "...", add_norm_wgsl: "...", ... }.
CLAUDE.md § Known gaps at repo root lists each open issue as:
## N. The <observable> deficit vs <reference> (<artifact>, <date>)
Observed. <quantitative gap with σ-significance>
Hypothesis A — <candidate root cause>
Hypothesis B — <alternative>
Falsification experiment: <what would distinguish them>
Standing entry (closed 2026-07-25): the 22% throughput gap behind WebLLM on M2 Pro, identical-weights identical-prompt. Closed by correctness fixes (fused_ffn f32 accumulation, attention barrier, decode off-by-one) plus promoting the measured vec4-load kernels to default; the M2 Max head-to-head now reads +16.0% total / +31.4% decode ahead (2026-07-30, BENCH.md). Hardware changed too (M2 Pro → M2 Max), so per-hypothesis attribution is not fully resolved; the surviving open sub-question — split-K attention at long context — is queued in BENCH.md.
Entries are removed when the underlying gap closes; the artifact
references stay in CHANGELOG.md. Tested-and-rejected hypotheses
get a strikethrough entry with the refutation artifact link, so
the same hypothesis isn't tried twice.
When a prior claim turns out wrong, revise it in the same commit that surfaces the data, with the full arc preserved. Examples:
- "Zero-TVM matches WebLLM throughput" — false on initial measurement.
README and BENCH.md were updated in the same commit as the
head-to-head artifact to read "22% behind WebLLM" instead. The
earlier overclaim is preserved in
CHANGELOG.md. (The arc continued: the 2026-07-25 M2 Max re-measurement flipped the headline ahead, recorded the same way — measurement first, claim second.) - "+73% on Qwen3-4B" / "+93% on Qwen3.5-4B" (2026-07-29) — withdrawn 2026-07-30. The bench harness never reset the engine's absorbed-token record between runs, so once cross-turn prefix reuse shipped, our half of the A/B prefilled one token per run while WebLLM's half prefilled the whole prompt. The comparison was not like-for-like. Both pairs are marked withdrawn in place in BENCH.md with the reason, the harness was fixed to reset per run and to split TTFT from decode, and all three models were re-measured. Every pair now carries two numbers — total wall-clock and decode-only — because reporting one alone is cherry-picking. A finding that shrank the headline (+73% → +31.7% on Qwen3-4B) is still a finding.
- "228 dispatches/token" was the result after fusing QKV+RoPE+KV-append into one dispatch; the earlier 3-dispatch-per-layer count is archived rather than retconned out.
- "Hand-written kernels can't beat TVM on dispatch count" — refuted
by 228 vs 342. The fusion contribution is the answer; documented
in
README.md§ "What's actually in the box".
This is publication-grade transparency. Wrong hypotheses become part of the public scientific record, not an embarrassment to hide.
Each minor release ships:
- Git tag (
v0.X.Y) - GitHub Release with notes drawn from
CHANGELOG.md - Zenodo DOI minted via the GitHub-Zenodo integration
CITATION.cffpreferred-citationblock updated with the real DOI
Patch releases (doc-only, refactor, etc.) skip the Zenodo step.
initGPU()MUST passrequiredLimitsformaxStorageBufferBindingSizeandmaxBufferSize. The default 128 MiB cap silently truncates large dispatches; Phi-3-mini's full weight set exceeds this without explicit limits.atomicAddworks only onu32— not f32. The forward pass avoids atomic reductions entirely (tree reductions in shared memory instead).- No recursion in WGSL. All shaders are single-pass.
- Uniform buffers must be 16-byte aligned.
- No subgroup intrinsics in WebGPU 1.0 spec —
zero-tvmships A/B variants ofmatmul.wgsl(subgroup + scalar) and selects at init based onadapter.info. Once WebGPU 1.1 ships subgroup intrinsics, the A/B split collapses.
- TypeScript
strict+noUncheckedIndexedAccess. No exceptions. - ESLint clean — 0 errors. Warnings tracked, ideally 0.
- CI green. Every PR runs unit + Playwright e2e + typecheck + lint.
- Each kernel has paired test coverage by intent, not by metric:
- Closed-form invariant (norm preservation, softmax mass = 1, tokenizer round-trip) where it exists.
- CPU pre-screen (plain-JS references inside
tests/kernels/run.mjs; there is nosrc/cpu/directory) on every kernel — falsify on CPU before shipping the WGSL. - Peer-package (HuggingFace Phi-3-mini fp16 logits, MLC quantized weights, WebLLM throughput) on a fixed prompt.
- Honest negatives (status: "fail" tests) live alongside passes; they don't break CI but they're surfaced in the suite output.
- Minor releases (
v0.X.0) for substantive features or scientific findings. Tag + GitHub Release + Zenodo DOI. - Patch releases (
v0.X.Y) for doc-only, refactor, SVG refresh, narrative updates. Tag + GitHub Release, no DOI. - CHANGELOG follows Keep a Changelog
format:
### Added / Changed / Fixed / Documented / Honest negatives. - CITATION.cff version matches
package.jsonversion matches Git tag matches GitHub Release tag, all pinned per release.
Inherit these 15 principles from day one. Copy this file verbatim into the new repo. Replace project-specific references in sections 1, 6, 7, 8, 10 with the new project's analogs. Cross-link sibling projects in the header.
The discipline is the product.
Last revised: 2026-05-14. Canonical mirror of
webgpu-q/RESEARCH_STANDARDS.md.
Edit either and propagate.