Stage 2 was explicitly authorized on 2026-08-13. The approved baseline is frozen and measured optimization is active.
Next available ID: OPT-0092.
| ID | Subsystem | Hypothesis | Evidence | Disposition | Result | Record | Implementation |
|---|---|---|---|---|---|---|---|
| OPT-0001 | profiler | Exact-shape, low-overhead measurement can localize M3 compute, scheduling, and transfer bottlenecks without repeated full-song runs | positive | integrated | Current exact GEMM is materially below the product need; uniform VAE quanta create 62,622 CBs per 10.24 s window | record, result | bbc961121b379b314c8929b16ae37eb292cde3cb |
| OPT-0002 | VAE scheduling | Work-aware output/MAC budgets can remove most scalar-VAE command and idle overhead without changing arithmetic | positive | integrated | Decoder quanta fell 62,622 → 3,942 (15.89×) with zero bit mismatches and materially lower paired wall time | record, result | 80f3c8bf550bd16fe18c64992627027972be18a7 |
| OPT-0003 | dense GEMM | A Parakeet-inspired fixed-32 subgroup kernel with tile-major packed-BF16 weights can materially accelerate the exact M3 shapes while retaining target-browser numerical identity | positive | integrated | Original thermal microbenchmark: exact, 4/4 wins, 1.894× faster. Package-native 24-layer × 8-evaluation gate: all 8,256 final-latent U32 bits exact, identical memory, and complete 634/826 portable/subgroup scheduling reconciliation. The accepted 12-second product WAV was then reproduced exactly in 64.8 s; its ~11.115× change is a nonthermal combined-stack checkpoint, not solely OPT-0003 | record, result | runtime a4e4ce4d2a2a74b7d9d0b1a05e7fd25343e9d404; package identity 68d7795c616c1520b1d97ddef9f9d3147ab3973e; harness bf6da81647814737e88c8e881da72e04887cba07 |
| OPT-0004 | VAE Conv1D | A workgroup-tiled FP32 K7 Conv1D can reuse overlapping input windows while preserving target-browser scalar output bits | positive | integrated | Exact on the pinned Chrome/M3 representative with 4/4 paired wins and a 2.988x median active-wall speedup; the production selector also passed an exact mixed tiled/portable decoder final-output gate | record, result | 48148ad1c791d26653d95418f55720240bafddff |
| OPT-0005 | VAE Conv1D | K-outer, input-channel-chunked FP32 K7 tiling can remove the whole-input-channel workgroup-storage ceiling while preserving source reduction order across dilation 1/3/9 | positive | integrated | Exact on the pinned Chrome/M3 checks; complete block-0 dilation one was 1.7235x faster active and 1.5907x faster including cooperative idle, and the production hybrid selector passed an exact complete decoder gate | record, result | 31e8ef7f385b4c3b21180b356ca2d89ec00a7099 |
| OPT-0006 | VAE scheduling | Bounded multi-quantum command buffers can remove the dominant queue-drain/idle floor while preserving every dispatch, pass boundary, and arithmetic operation | positive | integrated | Batch 8 preserved 524,288 production-range output bits and reduced median bounded-screen wall from 110.95 ms to 70.00 ms (1.585x); production integration then preserved all 12 mixed-decoder U32 outputs while reducing its decoder command buffers from 109 to 14 | record, result | 66f69709f7e30a9309e2315de0ea2ad7bae71604 |
| OPT-0007 | VAE K1 Conv1D | A pointwise FP32 Conv1D specialization can reuse input and weight tiles while preserving each output's increasing-input-channel arithmetic order | negative | abandoned | Exact throughout; C128 was positive but noisy for one range, while the production batch-8 16-range sequence lost 0/4 pairs at 0.82993x, so this fixed geometry is not integrated | record, result | benchmark candidate 1fa164b91a03cf6838fd94812e21949a8e621664; not integrated |
| OPT-0008 | VAE profiler | One authenticated 256-latent-frame shipped production window can attribute current VAE wall time to operations, families, selected kernels, and mixed-boundary batches before another kernel is chosen | positive | benchmark-only | Completed measurement: exact thermally valid 11,427.100000023842 ms shipped window; pure channel-chunked K7 dominates at 6,028 ms submit-through-drain and the measurement-time follow-up was fixed32 implicit-GEMM, while ConvTranspose remained second; no boundary flush was decision-relevant and no production integration occurred | record; result | production 9dbd6e9cb85da211aa9e8224edfc08a2eef3f706; harness 511a6696e5229894cf4db70554e6bfb4a6d11486 |
| OPT-0009 | FP16 GEMM calibration | Source-authenticated Parakeet subgroup kernels and ACE production-shape variants can establish the local M3's FP16/FP32 dense-math behavior and the accumulation choices worth carrying into an FP16-first profile | positive | benchmark-only | Parakeet native FP16 reached 2.797 TFLOP/s; on all four exact ACE shapes FP16 operands with FP32 accumulation was 1.274–1.390x faster than the accepted oracle and reduced the weighted 24-layer × 8-evaluation dense diagnostic from 39.1392 s to 28.896 s. Native FP16 accumulation was rejected after adversarial zero collapse, overflow, cancellation, and long-K drift; no production integration occurred | record, result | allocation 303ab8df036df71768a56774c59c75c4cfe30aa9; native harness 30b3b76c8114d2fa55bdc020d21d57ae53be70f3; final accumulation harness b41108dc1be75da9ba7a72ea64faa98c9dc81ecd; Parakeet 7ee112738262a6f5a0efd2f150748a4087432fbb; benchmark-only |
| OPT-0010 | Planner token profiler | Package-native M1/M2 planner token execution can be attributed across dense layers, tied head, readback/sampling, and per-quantum scheduling before choosing GEMV, vocabulary restriction, or command batching work | positive | benchmark-only | Six unchanged raw-FP16 tokens were full-logit bit-exact at 335.5–498.4 ms/token. Full-vocabulary CPU sampling cost 85.9–279.2 ms, layers cost 133.3–165.5 ms submit-through-drain, and 34 per-token idles cost 39.5–43.7 ms; compact semantic head/sampling ranks first, low-row FP16/FP32 GEMV second, and batching third. The 132–140 ms semantic saving is explicitly an unmeasured projection; no production integration occurred | record, result | production/core 00dfd4732aa019bbbb238ae40265fe86cb38f27b; frozen harness 81e84df955e7cb812a60d9c6decadff3791234e3; manifest c5b547cd08aa5e6d2971b2c9c84940b8af193f2e230ce689258ca81fcd292a3b; benchmark-only |
| OPT-0011 | FP16 VAE storage/window | Authenticated FP16 VAE weights and internal activation storage, with explicit FP32 islands, can unlock FP16 heavy kernels and a 512-latent-frame/64-overlap window without increasing the current largest workspace binding | inconclusive | superseded | Candidate package authentication plus the actual-Chrome FP16 K7, selected synthetic production-local Conv1D, pointwise, ConvTranspose1D, and Snake correctness gates are positive. Production-local Conv1D passed 5,374,258 first/rerun raw-U16 comparisons over 2,687,129 selected outputs per execution with zero mismatches, but its raw 55.119 s rAF / 24.751 s timer gaps do not support responsiveness. Pointwise passed 723,551,236 first/rerun raw-bit comparisons over both ingress fixtures and all 15 B256 Add operations. ConvTranspose1D authenticated the exact five-operation, 322-quantum B256 topology and passed 11,564 first/rerun raw-U16 comparisons over 5,782 selected outputs. Snake authenticated the exact 36-operation, 813-quantum, 844,627,968-element topology and passed 13,632,070 first/rerun raw-U16 comparisons over 6,816,035 selected outputs with zero mismatches; its heartbeat is liveness only and supports no responsiveness claim. This remains correctness-only partial evidence; complete FP16-256/512 windows, timing, waveforms, selector integration, and listening remain pending, and responsiveness is unestablished. The literal A/B/C window gate never completed; narrower exact-packed and revision-7 work superseded this umbrella without an OPT-0011 production selection. | record, partial result | allocation f07afbeb425157ca9c1eb4a9bc2102365ab7f616; package bcb74ed6df6d4b77296f07fe059e8f23e4474830; K7 core/harness 82f0fa4b3d5e676ec9dc967c3563dc9650cc59bd/8648a4390b2decdb5bbbdc0c119d6562dfc8181a; production-local Conv1D core/harness 75f70f12bdb43ae33b9bd37391b7d49be5aa1704/1320051a2413e1f187143ac0f79958df9218b54f; pointwise core/harness dd36a04960f846e53c2fd948d67b9aa9ddced4f2/1ab637aa3b174dcf3593beaa56fba6ce8ab4cd44; ConvTranspose1D core/harness d2bf0819d0460f6bd60ebe0457eb091b45e7bf6a/356f49b20841d0f051fcae3825a87c645c88c386; Snake core/harness ae2106c9d5834a3cd5cb836cad484665752230e3/59d94643def58a96c1cc081c177c91bd308fd83b; partial benchmark receipts retained; superseded by OPT-0028/0054/0066/0072 |
| OPT-0012 | Planner M2 head/sampling | Semantic M2 can score and read only the exact 64,000 audio-code rows, or the separate forced-EOS row, then sample over an ascending compact global-ID domain without changing retained logits, browser-v1 sampling bits, emitted tokens, or the Philox cursor relative to authenticated raw-FP16 arm A | positive | benchmark-only | Under the owner-approved nominal-start rule, two balanced corrective runs made compact arm C faster than A in all six medians and 32/36 same-round pairs, and faster than same-head B in all six medians and 27/36 pairs. Exact same-immutable-byte replay won 36/36 with 41.8–45.1% median savings; the allocation-free converter won 12/12 with >98.9% lower medians. Final A/B/C trajectories self-repeated twice per arm with the same 150 codes plus EOS, zero mismatches, and zero NaNs across 207,991,224 actual-package FP16 words; cancellation and lifecycle passed. Both screens started after 45 nominal seconds but later became non-nominal, so they are decision-useful with a disclosed caveat, not continuously all-nominal or release-quality. A-to-B alone was marginal/noisy. No production integration occurred; park planner integration until the direct-target FP16 VAE/DiT priority allows it | record, result | allocation f5e8e5db0b88a9a44dc96b73319183114daf136a; final core f73380bceebdd5568d93908a67ff33cea2b7d8f0; final harness eb6d8cc4c1d1db8f4fc75f0cd836ceed8a12daa8; benchmark-only |
| OPT-0013 | DiT full attention | Eight fixed-32 subgroups can share each 128-dimensional K/V tile across two GQA heads and four adjacent queries at the exact M2250 full-attention shape, materially reducing K/V rereads and barriers while retaining FP32 online-softmax state | positive | benchmark-only | The immutable exact-shape gate compared all 4,608,000 F32 outputs: max abs 2.6542693376541138e-8, NRMSE 3.3672934508159816e-7, and zero non-finite values. Query8 reduced compute wall 1108.8 → 136.2 ms (8.140969x) and full drain/readback wall 1117.3 → 140.2 ms (7.969330x), with 7.992895x fewer K/V scalar loads. Immediate production-integration candidate; no layer, trajectory, listening, or product-speed claim |
record, result | core a88e1b41c7b127d20c3fc4dbdee63acb77612a8c; gate 5b9cc1852e5178c8bbb907dbe4c90eae6004d6b8; benchmark-only |
| OPT-0014 | VAE K7 Conv1D | GPU-repacked bit-preserving KIO weights and a fixed32 16-row × 64-Cout subgroup tile can make K7 weight loads contiguous and remove inter-subgroup duplication while preserving K→Cin FP32 order | negative | abandoned | Exact across 13,854,720 output words and 61,017,600 repack U16 comparisons, but the weighted C300 K7 projection improved only 9,506.50 → 8,562.85 ms (1.1102x) for +61,017,600 B, with 15 regressing strata; no production integration or product-speed claim | record, result | core 12e128ab323c0024ed683313b4d06c07041213e7; gate 3904d212148cf2ecf93f317f8dcce3d59ef232a8; not integrated |
| OPT-0015 | VAE ConvTranspose1D | An exact-order direct-congruent or packed FP16 kernel can avoid looping, staging, and barriering over all 2*stride taps when production outputs have only about two congruent valid taps |
positive | integrated | The primitive gate had zero mismatches across 8,404,992 raw-U16 comparisons and projected a 3.64348x weighted speedup. Production integration reproduced the prior FP16 WAV exactly; homogeneous transpose fell 8,016.4 → 2,002.0 ms and decoder wall 14,118.3 → 7,265.8 ms in the 12-second profile. No 180-second or under-60-second claim | record, result | core 075ecc0b34b7541cffc0a83412c17ee31bbadab6; gate 65603ade17b9f3b9ca92cc0c29be83fe51a6e885; integration 36608b857827b2b1d31ac91bf5cca9639fb0b9ed |
| OPT-0016 | VAE K7 Conv1D | Reducing each packed-KIO subgroup lane from 32 live FP32 accumulators to 16 or 8, while replacing u32 unpack with direct typed-FP16 KIO loads, can improve M3 occupancy enough to beat the shipped fixed32 K7 path without changing K→Cin arithmetic | negative | abandoned | All three candidates were raw-bit exact, but weighted speedups versus packed 16x64 were only 0.97796x, 0.98153x, and 1.00771x; none reached 1.15x, so the declared early stop skipped the full phase and production integration. The nearby exact-order microtile space is closed; any rounding-changing K7 follow-up needs a new reordered-rounding, quality-gated experiment |
record, result | core 997891de0fe449c9b6551e80abc55604256969ad; gate 085669d5aec0fc02f3268c8b462385b59fb72ab7; not integrated |
| OPT-0017 | VAE K7 Conv1D | A WG256 adaptive cooperative implicit-GEMM tile with shared FP16 R32 panels and FP32 dot4 reduction groups can cut C300 K7 global operand traffic more than 4x versus shipped fixed32 and produce a multi-x wall-time gain within a declared reordered-rounding quality envelope | negative | abandoned | Repack and the frozen numerical envelope passed, but weighted cooperative dot4 regressed 13,761.60 → 26,183.70 ms (0.52558x), missing the 1.75x threshold; the full phase was skipped and no production integration occurred | record, result | core b83f4fe94d56787ddb980629ea6f41804543ca69; gate c34efbb67017c679c7932eed1df783254af17631; not integrated |
| OPT-0018 | DiT production profiler | Capture-only attribution of every current M2250 production DiT graph drain can identify whether any remaining kernel family has a credible path to at least 10 seconds of absolute saving before another DiT optimization is allocated | inconclusive | benchmark-only | The actual C98 DiT-only run remains decision-useful: 73.0726 s generation wall and 62.1482 s graph drains, led by 26.8603 s pure feed-forward and 21.4465 s mixed commands. The compact result and thermal trace were persisted, but the full authenticated browser receipt was not, so the literal evidence closes inconclusive. This was observational capture only; later mechanism-specific DiT experiments supersede the open profiling investigation | record | profiler/C98 authority f92de5a209ebb5f05ba9b37e5f3b7bfc88633d82; benchmark-only, closed |
| OPT-0019 | DiT dense GEMM | A WG256 M64xN128xK16 cooperative FP16 panel can remove duplicated packed-weight traffic and halve per-thread FP32 accumulator pressure while preserving increasing-K arithmetic across every production dense projection | negative | abandoned | Raw-U32 exact with every shape faster, but the complete 4/2/2/1 score improved only 223.2000 -> 169.7000 ms (1.31526x; 53.5000 ms saved), missing the frozen 1.55x ratio despite clearing the absolute-saving threshold. The stop rule forbids package/native or M2250 escalation; no integration or performance retry occurred | record, result | benchmark core/harness 8900ce670271cbd227142c584fb4a917e5d9cfb9; not integrated |
| OPT-0020 | DiT dense GEMM | Transposing OPT-0019's cooperative weight panel and reducing consecutive K4 groups with explicit FP32 dot can shorten the dependency/instruction chain while retaining its traffic, geometry, and bounded storage | negative | abandoned | Correctness and every frozen numerical gate passed, but dot4 C was slower than exact B on all four M2250 shapes: complete 4/2/2/1 B/C 0.5947464078891549x versus 1.25x, A/C 0.7203039805137254x versus 1.55x, and feed-forward B/C 0.5524747348537575x versus 1.25x; stop with no package, C98, M2250 product, listening, or integration escalation |
record, result | allocation baseline 7f8db6c5cfb17b2c508669e1b136d1bb02bab239; registration fce77739841572942eca4e96cc6a9f48eb02a971; benchmark core/harness 1d825a1399fdcc23c4ef3ce18151a2efa2415626; not integrated |
| OPT-0021 | DiT dense GEMM | Reorienting OPT-0019's two cooperative panels into padded K-major vec4<f16> storage can replace scalar workgroup assembly and reads with aligned vector operations while retaining its M64xN128xK16 ownership and exact increasing-K FP32 arithmetic |
negative | abandoned | Raw-U32 exact across 76,032,000 comparisons, but C beat exact B on only 2/4 shapes; complete B/C was 0.991159923845376x versus 1.075x, and A-C saving was 40.94999998807907 ms versus 52.0834 ms, so all frozen timing conditions failed and no package, C98, M2250 product, or integration escalation occurred |
record, result | allocation baseline bc2484096ff5622bc21d03aad80d08f12719989e; benchmark core/harness fa446366fa404e5ce00cf7350c206f7a63ba791b; not integrated |
| OPT-0022 | VAE ConvTranspose1D | Converter-native polyphase K-I-O weights plus a barrier-free fixed32 subgroup kernel can preserve exact tap/Cin arithmetic while removing OPT-0015's shared panels, barriers, and duplicated input loads | negative | abandoned | Raw-U16 exact across all 8,404,992 output comparisons and all 49,610,752 inverse-layout words, but block 2 regressed and weighted C300 A/B was only 1.2605735458036402x versus 1.3398349037268882x; the primitive stop rule forbids converter, package, C300 production, long-window, waveform, product, or integration escalation |
record, result | allocation baseline 73a8a85334226a2f2eb888960796ade8875ea6ad; benchmark core/harness 8345fce46c92afa109c1b38fa45d15df31b94de5; not integrated |
| OPT-0023 | VAE profiler | One authenticated package-native C4500 production VAE sequence can replace the stale C300-linear extrapolation with exact current-family/mixed-batch walls, scheduling/readback counts, and full-sequence wall attribution before another kernel is selected | positive | benchmark-only | Sole authoritative C448 + 10xC512 + C340 capture completed in 161,392.39999997616 ms under a fully nominal trace: K7 59,993.59999811649 ms, ConvTranspose 42,401.00000369549 ms, K1 25,772.300002217293 ms, mixed 8,943.499999284744 ms, with 17,658.19999921322 ms unsplit within-decode residual. All 90,687 dispatches, 11,350 submissions/drains, 69,120,000 raw bytes, steady memory, and cleanup counts reconciled. One prior stale-launch page setup was rejected before worker run/timed dispatch and was not a timing sample. No utilization, speedup, integration, quality, product, or under-60-second claim |
record, result | allocation baseline dc08f76ce44a6a46edbd4b60c9b74e6a7b019363; profiler/harness 02230725e460323de7e82ebed00177ec2103ea55; benchmark-only |
| OPT-0024 | VAE K7 Conv1D | Keeping the shipped barrier-free 8-row x 128-Cout subgroup ownership while reducing native contiguous-Cin operands in fixed groups of four with FP16 dot partials and one FP32 accumulator add can remove three quarters of scalar load/broadcast and FP32-add instructions from the 16 biased K7 operations without a new weight layout | positive | benchmark-only | Primitive weighted speedup was 1.715789x. The authenticated revision-6 C512 subsystem then reduced pure K7 4805.70 -> 2232.85 ms (2.15227x), decoder 6999.55 -> 4327.55 ms, and outer wall 8300.95 -> 5613.75 ms; whole-window NRMSE 0.00184056, SNR 54.701 dB, Pearson 0.999998306, deterministic and clean. Trajectory, product waveform/listening, and long gates still precede integration |
record, primitive result, C512 result | primitive bb0a152; benchmark-only |
| OPT-0025 | VAE K1 Conv1D | Recasting every pointwise Conv1D as the existing fixed32 FP16-operands/FP32-accumulation subgroup GEMM with tile-major weights can replace the low-throughput shared-panel K1 kernel while preserving source-order reduction and explicit FP16 output rounding | positive | benchmark-only | Raw-FP16 exact across 241,172,480 outputs. The five-shape C512 score improved 2,613.60 -> 255.45 ms (10.231356x), projecting the authoritative K1 family from 25.7723 -> 2.5190 s, about 23.253 s saved |
record, result | primitive 9bfd359; mechanism integrated through OPT-0028 |
| OPT-0026 | VAE ConvTranspose1D | Polyphase weights plus K7-style subgroup ownership of multiple rows and adjacent output channels can reuse each input across more outputs than OPT-0022 while preserving phase/tap/Cin FP32 term order | positive | benchmark-only | Exact pack/inverse over 49,610,752 U16 words and exact output over 141,312,000 U16 words. Every block improved; the five-block score fell 1,705.70 -> 527.65 ms (3.232635x), projecting the authoritative family from 42.4010 -> 13.1165 s, about 29.284 s saved |
record, result | primitive 72de722; mechanism integrated through OPT-0028 |
| OPT-0027 | VAE submission batching | Increasing the exact FIFO VAE batch from 8 to 64 quanta per command buffer can remove roughly seven eighths of 11,338 queue drains and 1 ms empty intervals while retaining one outstanding command buffer and bounded cancellation latency | positive | benchmark-only | Exact on every warm/timed C512 run. Batch64 reduced decoder command buffers 982 -> 123, requested idle 982 -> 123 ms, mean decoder 6983.90 -> 6014.10 ms, and mean outer wall 8479.30 -> 6204.25 ms (1.366692x, 2275.05 ms saved), with clean lifecycle |
record, result | benchmark candidate present; production selection pending |
| OPT-0028 | VAE exact kernel integration | Converter-native K1 tile-major and ConvTranspose polyphase weights can realize OPT-0025/0026 in production without browser repacking or duplicate weights | positive | integrated | Converter-native K1 tile-major and ConvTranspose polyphase layouts, fixed32 subgroup owners, portable counterparts, authenticated revision-6 package/profile, and fail-closed selection are integrated. Subsequent product and diagnostic gates retain OPT-0028 as the exact revision-6 VAE oracle | record | integrated in 0b27dde; exact revision-6 oracle retained |
| OPT-0029 | VAE ConvTranspose1D | A dense-style 32-row x 256-channel fixed32 subgroup tile can lift exact polyphase transpose far above OPT-0026's still-low 0.26 TFLOP/s by reusing each input and weight across eight rows/channels per lane | negative | abandoned | Raw-U16 exact over 141,312,000 words, but dense-tile weighted medians regressed 455.0 -> 480.3 ms (0.947325x) versus OPT-0026; no integration |
record, result | benchmark-only; not integrated |
| OPT-0030 | DiT full attention | Requesting the adapter's stock 512/1024-thread limit can extend query8 to query16/query32 and share every K/V row across 2-4x more streams | negative | abandoned | Stock Chrome compiled WG512/WG1024 and all 4,608,000 outputs were bit-exact, but Q8/Q16/Q32 medians were 139.6/219.2/173.0 ms; Q16 0.636861x and Q32 0.806936x versus Q8. Wider occupancy/barriers outweighed K/V reuse |
record, result | benchmark-only; not integrated |
| OPT-0031 | DiT dense GEMM | A stock-WG512 M128xN128xK32 cooperative tile can use the M3's wider limits to quarter OPT-0019's panel/barrier iterations while retaining exact increasing-K FP32 arithmetic | negative | abandoned | Raw-U32 exact across 101,376,000 comparisons, but every shape regressed and weighted 4/2/2/1 was 166.80 -> 181.45 ms (0.91926x); wider shared-panel workgroups are not the missing mechanism |
record, result | benchmark-only; not integrated |
| OPT-0032 | DiT dense GEMM | Direct FP16 K4 dot partials widened into FP32 running accumulators can shorten the dense dependency chain without the rejected full-FP16 accumulator or OPT-0020's FP32 horizontal dot | positive | benchmark-only | Full/adversarial numerics passed (NRMSE 0.000311, SNR 70.13 dB, max error 0.01443); every shape won and weighted 4/2/2/1 improved 199.65 -> 142.10 ms (1.404996x). Layer/trajectory/listening and converter-native integration remain |
record, result | benchmark-only; follow-up authorized |
| OPT-0033 | DiT full attention | Keeping query8 ownership while staging 8-16 K/V rows per workgroup can preserve ascending-key online softmax exactly and replace two barriers per key with two barriers per key block | negative | abandoned | Both candidates were bit-exact over 9,216,000 comparisons, but query8/key-block8/key-block16 medians were 108.5/125.0/119.7 ms; even 16-key blocking was only 0.90643x |
record, result | benchmark-only; not integrated |
| OPT-0034 | DiT submission batching | Encoding 8-16 consecutive FIFO graph quanta per command buffer can remove most of the 2,553 drains and 1 ms empty intervals responsible for the 62.15 s graph / 73.07 s stage gap without changing dispatch order or arithmetic | inconclusive | benchmark-only | The receipt's literal decision is negative-below-complete-stage-speed-gate, but performance attribution is inconclusive: all 576,000 candidate raw-U32 comparisons were exact with identical final-latent SHA-256, and batch1/8/16 reduced graph command buffers/drains 2,553 -> 320 -> 160 and requested idle 2,552 -> 319 -> 159 ms while maximum batch drain grew 329.2 -> 578.4 -> 1,301.6 ms. Fixed-order graph walls rose 123,094.6 -> 128,405.3 -> 134,193.9 ms, yet unchanged batch1 was 1.981x the prior OPT-0018 authority and condition/load stages also drifted, so no arm is integrated. Reject the unchanged seven-minute fixed-order rerun; revisit with short balanced/interleaved repeated graph segments or layer/evaluation slices |
record, result | core/harness ab88972a561f05d1cc2a3a11658f49795091f680; receipt 6c87e7f07f24b3eb2b1f515c153c873ab41c6b1a6df3c8f00bb57ef99ffa8a38; benchmark-only, not integrated |
| OPT-0035 | VAE window geometry | A 2,378-frame chunk is the minimum-memory two-window cover of C4500 at overlap 64 and can remove 1,280 of 5,908 decoded latent frames while remaining within the M3 adapter's stock buffer limits | inconclusive | benchmark-only | Exact and deterministic across all 17.28M waveform samples and every seam; C2378 reduced decoded frames by 21.6655%, peaked at 3,758,347,792 live bytes, and all 111 buffers were destroyed exactly once with zero live and one owner. Aggregate C512/C2378 wall was 109350.10/120413.25 ms (0.9081x), but C512 drifted 81213.9 -> 137486.3 ms (+69.3%) and forward/reverse pairings flip from 0.6190x to 1.2542x; timing is not a stable negative. Revisit with shorter paired subsystem or interleaved chunk-level timing, not an unchanged full four-arm rerun |
record, result, thermal | result 1833087730427396a3a198f20a3d6a159ca54cf1d5e9286f30539a3aac5aafa9; thermal 20e37497f21819f41e9faa3d8792a342cdac6969a4c95216e963ee2fd2b8f564; benchmark-only, C512 default unchanged |
| OPT-0036 | VAE ConvTranspose1D | Splitting OPT-0029's 64-accumulator dense tile into isolated 8-row×4-channel and 4-row×8-channel variants can capture one reuse axis at 32 accumulators/lane without its register-pressure regression | positive | benchmark-only | Raw-U16 exact throughout. Row and channel reuse passed every shape and reduced the summed five-shape medians 534.0 -> 373.0/372.5 ms (1.4316x/1.4336x). Channel reuse wins blocks 0-2 and row reuse wins 3-4, making a per-shape exact selector the follow-up |
record, result | benchmark-only; production selector not changed |
| OPT-0037 | DiT dense integration | Replacing the resident repeated-layer dense layout with OPT-0032's K4-native layout can realize its 1.405x primitive win across the actual 192 layer evaluations without duplicating multi-gigabyte weights | negative | abandoned | Sequential authenticated M2250 comparison stayed finite and passed NRMSE 0.01658, SNR 35.61 dB, and Pearson 0.999862, but maximum final-latent error was 0.99558 versus the frozen 0.25 cap, with 802 finite sign changes. Stop before audio/listening and remove the all-dense K4 profile from the product default; investigate the unique K6144 down projection under a new ID |
record, result | gate/runtime 9323425; receipt 4bb35dcbca8356c92be1bbe9eec72aab03e2646275bec42377e5266c61316bc0; not production-authorized |
| OPT-0038 | DiT dense GEMM | Accumulating two or four adjacent FP16 dot4 results locally before each FP32 running-state update can shorten the remaining dependency chain while bounding FP16 reduction length to K8/K16 | negative | abandoned | Numerics passed, but weighted K4/K8/K16 was 157.35/161.45/241.30 ms; K8 0.9746x and K16 0.6521x versus K4. K8 helps only the two smaller shapes, so retain OPT-0032 K4 as the general owner |
record, result | benchmark-only; not integrated |
| OPT-0039 | DiT full attention | A fixed-WG256 dual-query-per-subgroup tile can share K/V across 16 streams without OPT-0030's WG512 occupancy loss | positive | benchmark-only | Bit-exact over 13,824,000 comparisons. Dual-query16 halved workgroups and reduced K/V loads and barrier events 1.99645x; under the nominal one-check protocol, median wall improved 132.2 -> 97.4 ms (1.3572895x). Cleanup was complete. Production remains unauthorized; graph/layer/trajectory integration requires a new-ID follow-up |
record, result | receipt ffc0a27929e903c3366f5b1c8da4c945cf2f0f8a17134d73a96c315227b41ac9; benchmark-only, not integrated |
| OPT-0040 | VAE ConvTranspose1D | A static per-operation selector can use OPT-0036 channel reuse for blocks 0-2 and row reuse for blocks 3-4, capturing their complementary exact wins | inconclusive | superseded | OPT-0036 projected 534.0 -> 326.75 ms (1.6343x) versus OPT-0026, but OPT-0040 produced no standalone frozen C512 receipt. Its exact selector remained an oracle/input to OPT-0052/0054/0066, which superseded this allocation |
record | standalone gate incomplete; superseded by OPT-0052/0054/0066 |
| OPT-0041 | VAE K7 Conv1D | Bounded K8/K16 FP16 local partials can reduce OPT-0024's remaining FP32 update chain without shared memory or an unbounded low-precision accumulator | negative | abandoned | Numerics and lifecycle passed. K8 improved the production-weighted score 4279.35 -> 3532.05 ms (1.21158x) but slightly lost C512, while K16 regressed to 5273.40 ms; neither passed the frozen all-tier-win gate, so retain general K4 and stop before C512. The heterogeneous K8 wins are follow-up evidence only |
record, result | receipt 8c638a562e6a00ac9c49e26fb51b5ed6caf81c08d1ba6c8c1fd2e372a0e5da74; benchmark-only, not integrated |
| OPT-0042 | VAE submission integration | Selecting OPT-0027's exact batch64 schedule in ordinary production can realize its 1.366692x C512 outer-wall win and remove seven eighths of decoder drains/idle without changing dispatch order or arithmetic |
positive | integrated | Production and diagnostic paths now select batch64 with tested reconciliation/cancellation/cleanup. OPT-0027 measured exact C512 outer wall 8479.30 -> 6204.25 ms; OPT-0035 independently proved the complete C4500 batch64 waveform bit-exact, while its thermally unstable long wall supports no speed claim |
record, C512 result, C4500 result | allocation baseline ab88972; integrated in current production checkpoint |
| OPT-0043 | WebGPU utilization diagnosis | Standard GPU timestamp queries plus external AGX utilization sampling can separate actual shader execution from submit/drain/host overhead on the current and K4 dense paths | positive | benchmark-only | Weighted OPT-0009/K4 GPU-to-wall ratios were 0.906326/0.908718; GPU and wall speedups matched at 1.373856x/1.377481x. Dense is GPU-kernel-bound, with only about 9% of wall outside the timestamped pass, so submission tuning alone cannot close the target. Full/adversarial numerics and cleanup passed; coarse external AGX samples are contextual, not interval authority |
record, result, AGX | result 83f6ad1045d6d65fbebbe6ac86197d18bef682ea5e38b148c56e2d75abb7fcb7; external 8592ee84f603649f944887754758fc803955070bf36b7324dcdc67aecb264588; diagnostic only, no production authority |
| OPT-0044 | VAE K7 integration | Routing the 16 biased production K7 operations through OPT-0024's bounded FP16 K4-dot/FP32-running-state owner can realize its authenticated C512 2.15227x family win across the long decoder |
positive | superseded | OPT-0024 reduced C512 K7 4,805.70 -> 2,232.85 ms and outer wall 8,300.95 -> 5,613.75 ms with SNR 54.701 dB. The unchanged K7-only proposal is superseded by OPT-0066/0072's complete revision-7 K7+ConvTranspose owner, C4500 waveform gate, and final 180-second production validation |
record, final cumulative result | positive mechanism; superseded by OPT-0066/0072 |
| OPT-0045 | DiT full-attention integration | Routing the exact M2250 full self-attention operations through OPT-0039's fixed-WG256 dual-query owner can realize its 132.2 -> 97.4 ms primitive win across the 12 full-attention layers × 8 evaluations |
inconclusive | superseded | OPT-0039 remained raw-U32 bit-exact and 1.3572895x, but no OPT-0045 production receipt was completed. OPT-0061 qualified the stronger exact quad owner, OPT-0062 exercised it at graph scope, and OPT-0070 owns production integration |
record | dual-query primitive retained; integration superseded by OPT-0061/0062/0070 |
| OPT-0046 | VAE K7 Conv1D | Separating boundary rows from the dominant all-taps-valid interior can remove eight dynamic validity branches and conditional adds from every Cin4 iteration while remaining bit-identical to OPT-0024 | inconclusive | superseded | The registered correctness/timing gate was never executed. Later output-major and row-reuse K4 layout work measured stronger shape-specific mechanisms and superseded this unmeasured branch-free direction | record | no OPT-0046 kernel/selector integrated; superseded by OPT-0047/0051/0057 |
| OPT-0047 | VAE K7 Conv1D | Reordering OPT-0024 K4 weights so each subgroup's 128 output-channel vectors are contiguous at fixed K/Cin4 can remove its warp-wide 7 * Cin/4 weight stride without changing arithmetic |
negative | abandoned | Raw-U16 exact throughout. C1024/C512 won 2.292x/1.544x, but C256/C128 were 0.848x/0.958x, failing the all-tier rule despite a 1.86987x weighted win. An exact new-ID high-tier selector projects 1.90421x and is the authorized follow-up |
record, result | receipt b78c577b16a41ad68597da882a92dd07b0d7125a7f72b6a3ce2f1a5ed1cd81af; benchmark-only, not integrated |
| OPT-0048 | VAE ConvTranspose1D | Applying bounded FP16 K4 partials with FP32 running state inside OPT-0040's best per-shape reuse geometry can turn the remaining scalar Cin reduction into the same efficient contraction mechanism that won in K7 and DiT | negative | abandoned | Raw-U16 exact over 282,624,000 candidate comparisons. The sum improved 316.20 -> 192.55 ms (1.64217x), but block 0 regressed to 0.9320x, failing the all-shape gate. Blocks 1–4 won by 1.27–2.41x, authorizing only a new-ID exact shape-selector experiment |
record, result | receipt 135f3939eb027849f41fa0915e38ee317a940281303196ba00ae58dd739140b2; benchmark-only, not integrated |
| OPT-0049 | VAE K7 Conv1D | Selecting bounded K8 only for the C1024/C256/C128 tiers it won, while retaining K4 for C512, can realize OPT-0041's heterogeneous gains without its losing shape | inconclusive | superseded | OPT-0041 medians projected 4279.35 -> 3510.90 ms (1.21888x), but the selector gate was not executed. OPT-0051/0057 established a materially stronger shape-selected row-reuse direction using bounded K4 arithmetic |
record | no OPT-0049 route integrated; superseded by OPT-0051/0057 |
| OPT-0050 | DiT dense GEMM | Bounded K4 FP16 fma(vec4) across independent output channels can retain OPT-0032's FP32 running state while using the vector direction behind Parakeet's fast path, instead of horizontal FP16 dot |
negative | abandoned | Raw-U32 identical across 25,344,000 production and 17,408 adversarial outputs, but weighted time regressed 168.50 -> 170.55 ms (0.987980x) and three of four shapes lost. Parakeet's vector direction adds no general win under bounded K4/FP32-running arithmetic |
record, result | benchmark cab046e; receipt 51a61af657046f77300290842b2f81da1eb476f2063c7e5803942ffa4816cd01; not integrated |
| OPT-0051 | VAE K7 Conv1D | Rebalancing OPT-0024's unchanged 32 FP32 accumulators/lane from 8×4 to 16 rows×2 outputs can double K4 weight reuse while preserving its bounded FP16-dot/FP32-running-state mechanism | negative | abandoned | Raw-U16 exact across all 16 cases. C1024/C512/C128 won 2.3325x/1.6455x/1.6786x, but C256 regressed to 0.875x, failing the all-tier rule despite a 1.977639x weighted win. The universal owner is closed; an exact same-K4 new-ID selector projects 1.999629x and is the authorized follow-up |
record, result | allocation baseline 93b5f9c; candidate core 59e144c1316d642d362d206222888177cd4e792743b3e23631ca415e923d770a; receipt 1e05cf8e9a00cf9b9b5d0e0eea5f442b132b27139c4221c4bd540e5f71b3c85a; benchmark-only, not integrated |
| OPT-0052 | VAE ConvTranspose1D | Keep OPT-0040 exact for block 0 and select OPT-0048 K4 for blocks 1–4, retaining raw-bit identity while avoiding the candidate's sole losing shape | inconclusive | superseded | Measured primitive medians projected 316.20 -> 189.50 ms (1.66860x), but no standalone authenticated package/C512 receipt completed. Its block-0 exact plus blocks-1–4 K4 mechanism was subsumed by OPT-0054, reauthenticated under OPT-0066, and promoted under OPT-0072 |
record | standalone gate incomplete; mechanism subsumed by OPT-0054/0066/0072 |
| OPT-0053 | VAE K7 Conv1D | Select OPT-0047's output-major layout only for C1024/C512 and retain OPT-0024 native K4 for C256/C128, avoiding both losing tiers while preserving OPT-0024 arithmetic | inconclusive | superseded | Direct medians projected 4660.05 -> 2447.23 ms (1.90421x), but the package/C512 gate did not complete. OPT-0051 measured the stronger row-reuse geometry and OPT-0057 selected it by shape, superseding this output-major direction |
record | no OPT-0053 package/route selected; superseded by OPT-0051/0057 |
| OPT-0054 | VAE mixed K4 package integration | One revision-7 VAE package/profile can combine OPT-0052's raw-exact transpose layout with OPT-0057's shape-selected row-reuse K7 layout without duplicate package generation, while explicitly retaining K4's quality risk | inconclusive | superseded | The original preparation did not reach READY, so its literal result remains inconclusive. OPT-0066 allocated and passed the corrected complete same-arithmetic oracle and C512 quality/timing gate; OPT-0072 owns production promotion after the complete C4500 waveform gate | record, preserved failed setup | original gate superseded by OPT-0066/0072; failed setup retained |
| OPT-0055 | DiT sampler schedule | Six source-derived shift-3 Euler evaluations can remove one quarter of repeated denoiser work while retaining the pinned solver, DCW, BF16 materialization, and all model/VAE math | inconclusive | benchmark-only | The 8/6/5 direct-12 diagnostic reached listening-ready with raw-U32 evaluation-0 identity and three complete blind WAVs, but formal timing was absent and no genuine owner instrumental/vocal listening attestation or mapping reveal occurred. Removing two evaluations is a whole-trajectory quality change, so exact eight remains production | record | diagnostic implemented; no selection, exact eight retained |
| OPT-0056 | DiT selective dense precision | Retaining exact FP32 accumulation only for the unique K6144 MLP down projection may remove OPT-0037's trajectory outlier while preserving K4 on the other 73.9% of repeated-dense MACs | negative | abandoned | The selective exact-down arm was deterministic and improved final maximum error 0.9955760 -> 0.6395987, while NRMSE 0.0078057, SNR 42.1517 dB, and Pearson 0.9999693 passed. Maximum error still exceeded the unchanged 0.25 cap, so correctness failed and timing was skipped; all arms stayed finite with clean sequential disposal |
record, result | receipt 9d84fcb6daa6c18a702f34ac645e47d3611691eb45e13c5ab6b9308238bbdf96; no integration |
| OPT-0057 | VAE K7 Conv1D | Select OPT-0051 row reuse for C1024/C512/C128 and retain unchanged OPT-0024 native K4 for C256, avoiding the candidate's sole losing tier while preserving same-K4 output bits | inconclusive | superseded | The first C512 preparation failed before READY because its benchmark oracle retained scalar revision-6 ConvTranspose arithmetic. OPT-0066 allocated the corrected complete same-arithmetic oracle, authenticated the row-reuse mechanism, and passed the joint C512 gate; OPT-0072 owns production promotion | record, failed setup | literal result inconclusive; superseded by OPT-0066/0072 |
| OPT-0058 | DiT dense GEMM | Standardized packed signed-int8 DP4a can replace four scalar products per instruction, shorten cross-group FP32 updates, and halve dense weight bytes if dynamic activation quantization remains cheap enough | negative | abandoned | Stock Chrome 151 / Apple metal-3 compiled the standardized packed-dot feature and all G32/G64/G128 full production aggregates passed, but every group failed the frozen adversarial gate: finite-to-zero collapses were 104/107/110 of 21,504 (0.0048363/0.0049758/0.0051153) versus at most 0.000001; K4 had 13/21,504, so int8 was materially worse. Correctness stopped before timing; three earlier setup attempts did no timed work and are not evidence |
record, result, feature probe | receipt 4dc0a49c5897fe6684bb227c7ca02876dd44ba3958dd5f35f849e1994c8bfbf9; clean 398/398 cleanup; zero GPU errors; not integrated |
| OPT-0059 | VAE window geometry | Revisit C2378 with short balanced C512/C2314 plus edge-shape timings after the final revision-7 owners land, avoiding OPT-0035's multi-minute thermal/order confound | negative | benchmark-only | Correctness, lifecycle, memory, and projected wall were strongly positive: aggregate C4500 projection 38.4331 -> 29.4480 s (1.305117x, 8.9851 s saved), both directions exceeded 1.30x, normalized decoder ratio was 0.986905, and peak live storage was 3,667,109,696 bytes. The literal gate remains negative (performanceGatePassed: false, decision retain C512) because frozen route-family buckets failed after shape-dependent batch64 pure/mixed composition shifts; no post-hoc waiver, selection, or integration. Three rejected setup/correctness preflights produced no timing, including the invalid OPT-0063 slot-liveness seam |
record, result | result 18ea3b278675ab9ac40e77cab005af3cdcb7e4458c169dee9204de56329cdf3a; core/harness 3edc51fe; benchmark-only, C512 retained, not integrated |
| OPT-0060 | DiT full attention | Integrate OPT-0039's exact WG256 dual-query owner for production full self-attention while leaving sliding/cross attention and every other graph owner unchanged | inconclusive | superseded | Post-registration audit found this unchanged integration direction and its gates were already frozen under OPT-0045. No OPT-0060 code or GPU work started; retain this ID only as duplicate-registration provenance and continue exclusively under OPT-0045 | record, prior authority | superseded before implementation |
| OPT-0061 | DiT full attention | Keep WG256 and extend OPT-0039 from two to three/four independent query streams per subgroup, reducing workgroups and shared K/V/barrier events again without OPT-0030's WG512/WG1024 occupancy loss | positive | benchmark-only | All 32,256,000 correctness comparisons were raw-U32 exact with deterministic complete writes and intact canaries. Authoritative eight-round medians were query8/dual/triple/quad 133.95/111.90/91.75/77.25 ms; quad cleared both gates at 1.44854x over dual and 1.73398x over query8, while triple missed the query8 gate. Eight-evaluation 96-call arithmetic is 5.4432 s; six-evaluation 72-call arithmetic is separately 4.0824 s. One setup launch-delay rejection and two no-click preflights produced no timing. The earlier missing-samples receipt is retained as explicitly non-authoritative |
record, result, non-authoritative provisional | result aa94b429d026d8e2093589b8664be24dbd64ffc14f51160ced4682521a3b95e6; benchmark-only, not integrated |
| OPT-0062 | DiT full-attention integration | Route only the exact 96 odd-layer M2250 full-self labels through OPT-0061's qualified fixed-WG256 quad-query32 owner while retaining query8 for all sliding and cross attention | inconclusive | benchmark-only | Exact integration correctness was positive: 442,368,000 actual-layer U32 words had zero mismatches, and all eight taps, final latent, and quad repeat were exact. The literal performance gate failed: aggregate self-full was 1.560740x and graph-savings sum was 12,896.1 ms, but forward graph/stage/slice regressed, non-attention families drifted, and both long traces spent most observations at thermal levels 1-2. Keep query8 default; revisit only with short balanced/interleaved full-self or bounded graph/layer slices after cool starts, not another unchanged multi-minute full graph |
record, result, failed setup, primitive authority, preserved OPT-0045 | candidate aa46aa5f4f7b1f8ca15453c336418b9cfe471348; benchmark-only, not integrated; query8 default retained |
| OPT-0063 | VAE exact epilogue fusion | Preserve every former FP16 storage round explicitly while fusing adjacent K1/Add/Snake producer chains to remove intermediate traffic and dispatches | positive | benchmark-only | All 13 production/adversarial cases were raw-U16 exact at the former Add boundary and final Snake output with deterministic complete writes and intact canaries. Aggregate affected-chain speedup was 1.207401682x; C512 improved 636.0000 -> 522.0000 ms (1.218390832x) and C2314 2810.0187 -> 2332.0781 ms (1.204941929x). Two-C2314 C4500 planning saving is 955.881176 ms, above the 750 ms escalation threshold. Cleanup was 260/260 with zero live resources. Positive isolated evidence authorizes only a joint revision-7/C2378 decoder screen, not integration, default selection, or product claims |
record, result | fused core b50fa754171b243e410d427034867dd06329f5f1ff1d5cb2c372cae5e80ed825; isolated benchmark-only, not integrated |
| OPT-0064 | Direct product pipeline | Attribute and overlap only dependency-independent request-scoped cache/hash, upload, compilation, and graph-construction work while retaining the hard DiT-before-VAE residency boundary | inconclusive | benchmark-only | Capture-only attribution was positive: initialization 29.310 s, generation 16.4371 s, and all 459 events reconciled; 5.7318 GB across 102 uploads took 3.3311 s summed file wall (2.3813 s stream, 0.3895 s write, 0.1494 s drains, 0.0528 s gaps), while construction was 2.3568/0.1419/0.1127 s and execution 2.2661/6.7245/3.7227 s for conditioner/DiT/VAE. The literal receipt remains failed because the unapproved OPT-0028 FP16 path produced 409b… instead of owner-approved FP32 d085…; byte length, lifecycle, diagnostics, and the 137-observation all-level-0 thermal trace were clean. No mapped upload, overlap, optimization, production selection, or speed claim is authorized |
record, failed receipt / positive capture, thermal trace | receipt b3be2509a3b81522229cbb69f743aa4d65662f907604f085c5f2269151198498; trace file 7de13a65826e4596b73d31ba1a3a617cc4855053107428b849e98eeb721b3144; raw trace 7ecfbbb72a6c51dd193571aea507a8bdae9b47ee5448f4cae44dc8e76a424fcf; no optimization authorized |
| OPT-0065 | DiT sampler schedule | A five-point schedule derived from the pinned shift-3 formula can remove one further complete denoiser evaluation after OPT-0055 while retaining Euler, DCW, BF16 materialization, and every model/VAE operation | inconclusive | benchmark-only | The joint 8/6/5 direct-12 diagnostic reached listening-ready and proved evaluation-0 identity, but no owner listening attestation, mapping reveal, vocal fixture, balanced timing, or five-versus-six product gate occurred. Five evaluations are not a minuscule arithmetic drift; exact eight remains production | record | diagnostic implemented; no selection, exact eight retained |
| OPT-0066 | VAE revision-7 joint quality gate | Treat both selected K7 and selected ConvTranspose FP16-K4 partials as approximate versus scalar revision 6, while a complete same-arithmetic oracle isolates physical layout/routing identity | positive | benchmark-only | The corrected complete K7+ConvT same-arithmetic oracle reached READY and the C512 gate passed. All 16 selected physical tensors matched across 35,880,960 U16 words; the candidate passed the unchanged scalar envelope at joint NRMSE 0.0015684427168221327, SNR 56.090626762434 dB, and Pearson 0.9999987699871865. Balanced scalar/candidate/candidate/scalar medians were K7 1.950901x, ConvT 2.000510x, decoder 1.714172x, and outer 1.677501x, with both directions passing and clean 10/10 sequential-owner cleanup. This authorizes only OPT-0059 C2314 then C4500 waveform/listening gates, not production or an under-minute claim |
record, result, thermal rejection, prior failed setup | result 3062c4ca30e346fc3a4d0bd8e7dcf1258c76d021e808c47c516c27c963c71b63; first level-1 preflight 78fe992709013c5f567a7247968a4f56245f791185bfa15dd33c0a2fe5bd414b; candidate acf8dd6c; benchmark-only, production default retained |
| OPT-0067 | DiT quad-query evaluation slice | Isolate OPT-0062's exact quad owner on only the first complete 24-layer M2250 evaluation, using four independently cooled ABBA arms so no candidate/control pair shares a multi-minute heating trace | inconclusive | benchmark-only | Correctness passed with zero mismatches across 55,296,000 selected actual-layer U32 values and exact identity for all 288,000 evaluation-result words. Quad won full-self and complete evaluation wall in both directions, reached 1.5808922936358476x aggregate full-self speedup, and projected 5828.400000095367 ms saving over eight evaluations. The literal frozen result remains status: failed, passed: false: individual non-target family ratios exceeded the strict 2% gate, including sampler-DCW +0.5 ms and input +7.100000023841858 / +39.59999990463257 ms, even though aggregate non-full-self improved 178.00000023841858 / 247.90000009536743 ms. This is strongly positive causal quad evidence but inconclusive/non-pass overall, not a candidate regression; it authorizes no optimized-stack selection and query8 remains default |
record, result, OPT-0062 result | receipt 66d96c21ddf9d2dc8c30fb87d8759c5709899b66166268fe91751d75b84a5b95 (150,537 bytes); core/harness e8ea8b0e; decision evaluation-slice-non-pass-keep-query8-default; all four nominal through-cleanup traces and two rejected untimed A1 preflights retained; no production/default selection |
| OPT-0068 | Native Metal/MPS/MLX ceiling | Native SIMD-group matrix MMA unavailable to stock WGSL can provide enough exact-shape dense and complete-graph headroom to make the M3 target robust rather than thermally marginal | inconclusive | benchmark-only | The static Phase-1 harness passed 7 Swift tests, a release build, its no-device describe path, 2 Python compile checks, 3 Python tests, and JSON-schema syntax parsing. The authentic six-activation/nine-output M2250 fixture is absent, so no native device, MPS/MLX execution, timing, receipt, or ACE result exists; same-machine MPS/MLX/Parakeet figures remain ceiling references only. Native selection is outside the stock-browser WebGPU/WASM scope, so freeze this investigation without a production/default change | record | static benchmark harness under benchmark/opt-0068-native-metal/; no native GPU/Metal execution, result, or production integration; frozen |
| OPT-0069 | Warm-cache authentication | Replace scalar TypeScript SHA-256 over the exact warm-cache acquisition inventory with a bounded stock-browser hashing mechanism while preserving full manifest authentication | inconclusive | benchmark-only | Literal failed-or-inconclusive / performance.passed: false is retained: WebCrypto authenticated all 158/156 records/digests and 7,325,999,133 physical bytes exactly under eight nominal traces, stayed within 365,005,824 transient bytes, reached a 6,700.95 ms median and 1.0933 GB/s, and saved 22,220.95/22,603.85 ms directionally, but maximumReadCopyRegression = 2.965945 exceeded the frozen 20% component gate. No post-hoc waiver or integration; the complete-wall product objective moves to OPT-0071 |
record, failed/inconclusive result, raw traces | result 46d4a308741e60dce68a86bbfad3c1b5f17b2ce367898bdc138758044b451dac; core/harness bbc7230d1edbac04b84786eefefbd6d495650b2f; benchmark-only, not integrated |
| OPT-0070 | Exact cumulative promotion | Promote exact quad full-self attention and exact C2378 VAE windowing together under causal production gates instead of rewriting their frozen noisy/composition-sensitive non-pass receipts | positive | integrated | Quad is raw-U32 exact through actual layers, all eight taps, and final latent and projects 5.8284 s saving; C2378 is complete-waveform exact and projects 8.9851 s C4500 saving. The final production run used shape-safe OPT70 attention plus two C2314 windows and completed 180 seconds of output in 94,774.8 ms |
record, final result | integrated in 78d7346; shape-safe owner ee2a5c9/75302f4; final validated |
| OPT-0071 | Warm-cache authentication integration | Replace only the scalar cached-file verifier with bounded one-file-at-a-time WebCrypto, judging complete authentication and ordinary cache-only READY wall while retaining exact package trust and lifecycle | inconclusive | benchmark-only | Literal failed-or-inconclusive is retained because tiny unrelated-stage ratios and the A2 aggregate 16.8 ms (3.174%) regression exceeded the frozen 2% rule. All four arms were exact, safe, cache-only, mutation-free, and nominal; sequential WebCrypto had a 6,290.35 ms authentication median, bounded 365,005,824-byte transient, and causal READY savings of 22,724.3/22,098.1 ms. The narrower owner-authorized selection moves to OPT-0073 without relabelling this receipt |
record, failed/inconclusive result, OPT-0069 authority | receipt ae1e63150e1834010a4fdc3a338fdcec0afb7db1af4f83e5eb15a1fc26f4e594; benchmark-only, OPT-0073 follow-up |
| OPT-0072 | VAE production quality promotion | Promote the measured revision-7 dual-K4 VAE under a new product identity and pair it only with OPT-0070 C2378 after complete C4500 waveform evidence | positive | integrated | All 17,280,000 raw stereo samples were finite; NRMSE 0.0015226894, SNR 56.3478 dB, Pearson 0.999998841, max error 0.00787494, repeat raw-U32 exact, seams finite, peak 3,758,347,792 GPU bytes, and 62/62 clean destroys. The final 180-second receipt identifies public OPT72 mapped to physical OPT66, two C2314 windows, all samples finite/nonzero, and no device loss |
record, long result, final result | integrated in f45254e; final validated |
| OPT-0073 | Warm-cache authentication selection | Select the exact, bounded sequential WebCrypto verifier for the current revision-7 production tuple under the owner's delegated best judgment, without rewriting OPT-0071's literal non-pass | positive | integrated | Sequential one-file WebCrypto remains exact and bounded, with no scalar fallback; focused tests passed 24/24. The final OPT70/72 production run completed Generate-to-WAV in 94,774.8 ms, produced 17,280,000/17,280,000 finite/nonzero samples, stayed below 4 GB tracked GPU memory, and released output without device loss |
record, final result, OPT-0071 authority | integrated in 33feaf6; final product validated |
| OPT-0074 | DiT dense GEMM | Two-term native-FP16 dot partials widened immediately into an increasing-K FP32 running sum can recover a material part of K4's throughput while reducing its local rounding exposure enough to survive the complete denoise trajectory | inconclusive | benchmark-only | Correctness passed and K2 improved on K4's full/adversarial error. Weighted K2 means were 1.15058x wall and 1.21494x timestamped GPU, but one of six weighted wall rounds lost and exact/K2 ranges overlapped; the frozen mixed/variance gate failed, so no trajectory, package, or integration work occurred |
record, result | candidate/harness 3ad45dade32cb5d53c37a685ef1f2ba01427fef0; benchmark-only |
| OPT-0075 | DiT RMSNorm | A width-128 WG128 owner can remove half the lane invocations, half the shared allocation, and the redundant stride-128 zero half while retaining the exact WG256 arithmetic tree seen by every finite production row | inconclusive | benchmark-only | Raw-U32 identity passed across 7,014,400 production plus edge outputs with exact reruns and clean lifecycle. WG128 reached mean 1.82695x GPU / 1.80119x wall, but saved only 6.96832/7.14688 ms per layer mix versus the frozen 10.5 ms, projected only 1.006–1.372 s over eight evaluations, and lost one weighted pair. Stop without DiT profile or integration |
record, result | candidate/harness 03456cfef8d4eb42a91564808efc8532f682b246; benchmark-only, WG256 production retained |
| OPT-0076 | VAE C256 K7 promotion | Route only the three C256 residual K7 labels from native scalar FP32 to the existing native-layout OPT-0024 FP16-dot4/FP32-running-state owner | inconclusive | benchmark-only | Selector/oracle identity and the numerical envelope passed over 245,760 U16 values. Mean/median GPU speedups were 3.60390x/3.80198x and wall speedups 2.94194x/3.67089x; all six GPU pairs won, but one candidate-first fenced-wall pair lost 11.7 vs 10.6 ms, so the frozen all-pairs gate stopped before a diagnostic decoder profile |
record, result | selector/harness 6968ed5; benchmark-only, production scalar C256 retained |
| OPT-0077 | VAE selected K7 Conv1D | A condition-1 radix-2 RFFT16 overlap-save pipeline can replace 70 direct K4 dot products per ten outputs with 30 real spectral K4 dot products while preserving the current twelve-route scope and the established K7 numerical envelope | negative | abandoned | Correctness passed across 888,448 candidate/scalar U16 comparisons with zero rerun differences and aggregate NRMSE 0.0000784971, but the transform was neutral-to-slower: mean/median GPU 0.9990x/0.9308x, wall 1.0007x/0.9309x, only 1/6 paired wins, and three regressing stratum medians. It also adds 72,548,352 package bytes. The exact mechanism stops before C512/package escalation; a setup-only thermal-launch rejection ran zero timing work and is disclosed separately |
record, result | allocation 47c4dfed1a2b4826cafda14123282ce72d852cb2; registration 3539a87; candidate/harness f78d0d24c9b205a9614703e17b7e870793d6107e; not integrated |
| OPT-0078 | DiT dense GEMM | A WG256 M32xN256xK32 owner can multicast each native packed weight tile once across eight subgroups and halve per-lane FP32 accumulator pressure while preserving OPT-0009's workgroups, increasing-K arithmetic, and package bytes | inconclusive | benchmark-only | Exact correctness, thermal, and lifecycle gates passed; every shape mean/median improved and the weighted score won all 8/8 GPU and wall pairs. The literal gate failed narrowly: mean GPU/wall speedups were 1.108390x/1.110871x versus 1.12x, and mean GPU saving projected 3,988.783 ms versus 4,000 ms. Median projections were 6.449/6.403 s, but no conjunct is waived; stop before profile or integration |
record, result | registration 23e0947; candidate/harness bfd286002654dc67b85d2986686ad917e497d073; benchmark-only, OPT-0009 production retained |
| OPT-0079 | DiT dense GEMM | An M64xN128xK32/WG256 owner can reuse each converter-native N128 half-tile across 64 rows, decode its packed weights once into an 8 KiB typed-FP16 panel, and remove hot-loop unpacking while preserving increasing-K arithmetic and package bytes | inconclusive | benchmark-only | Exact correctness, thermal, and lifecycle gates passed, but h-h mean GPU and h-1024 median GPU regressed and the weighted score won only 6/8 GPU and wall pairs. Mean GPU/wall speedups were 1.049890x/1.073385x; all 1.15x, 25 ms, and 4.8 s gates failed, and mean wall/GPU saving ratio 1.544278 disagreed. Stop before profile or integration |
record, result | registration 6313bfb; candidate/harness aade2c0223383bab99cf477e7697e2394c05a380; benchmark-only, OPT-0009 production retained |
| OPT-0080 | DiT/VAE scheduling | Preserve every singleton FIFO command buffer and completion fence, but keep at most two submitted within phase-aligned four-completion epochs so JS encoding, promise handling, and driver work can overlap while three quarters of the cooperative queue-empty intervals disappear | positive | integrated | The full-M2250 candidate preserved all eight taps and final latent raw-U32, reduced graph drains 2,553 → 639, saved 7,437.1 ms, and reached 1.138737x; its fail-closed direct DiT selector and exact 96-second product gate are integrated. The actual-C2314 VAE ABBA screen retained all 556 fences, reduced drains/idles 556/555 → 139/138, reproduced all 8,885,760 waveform U32 words and 12 guard regions exactly, and authorized narrow integration. The later direct 96-second product-exact gate proved forced control/candidate final-latent, full-raw, seam, and WAV identity; seam-free ordinary production selected depth two only for C2314, kept the C214 remainder depth one, and matched the candidate WAV byte-exactly. That gate was untimed, captured no thermal trace, did not execute the planner, and makes no timing, thermal, or planner claim |
record, evaluation-slice result, full-graph result, DiT product result, VAE scheduler result, VAE product-selector result | full DiT candidate/harness 2a9be026c6f36363ca73d44fb3ede9c49b5aaacc; DiT base integration 7d9ae463aa580d0290421b067467f799cf1b0b75; direct-only correction/evidence seam 023bdecbf670b9309db37b6ac3030293ffe3b463; DiT product core/harness dfb2a24c979f840f13909b6baee0742bd7ee4f40; VAE scheduler screen core/harness f8fde273ad57d21c8410710518c63e501acaa1ec; VAE production integration/product-exact core/harness 8d443f6aff4f6a9b06df6db77b3f88b9401123a7 |
| OPT-0081 | DiT dense input storage/GEMM | Folding the dense operand's existing FP16 rounding into six producer stores can halve repeated activation traffic and graph residency, while applying the same typed-F16 boundary to OPT-0078's materially different WG256 weight-multicast owner can compound that gain without changing weights, increasing-K FP32 accumulation, or FP32 dense outputs | inconclusive | benchmark-only | The primitive selected B and the sole actual-MLP run retained it. B/A GPU/wall means were gate 47.423488 → 36.962304 / 48.912500 → 38.587500 ms, up 40.108032 → 37.855232 / 41.625000 → 39.100000 ms, down 42.582016 → 41.107456 / 43.437500 → 42.062500 ms, and complete chain 123.052032 → 108.068864 / 124.150000 → 108.737500 ms; every panel mean/median improved and pair wins were 8/8, 7/8, 7/8, 7/8 on both clocks. Chain savings were 14.983168/15.412500 ms, ratio 1.028654290844513, projecting 2,876.768256/2,959.200010 ms over 24 * 8. All 73,728,000 boundary U16 words and 193,536,000 dense U32 comparison instances were exact; the 98-observation trace was wholly nominal with 1,013 ms maximum gap, and cleanup destroyed 16/16 buffers with 132/132 maps balanced and zero live resources. C remains a diagnostic non-pass and nonselectable: chain C/B saved only 4.857856/4.725000 ms, won 6/8 pairs, and its gate GPU median regressed. The representative graph-prefix gate was exact and clean but stopped escalation: B mean/median wall improved 548.4625 → 519.7500 / 538.5000 → 523.6500 ms, yet it won only 6/8 pairs and 2/4 reverse pairs; forward saving was only 24.05 ms, and the directional 95% lower endpoints were -0.8357/-42.6067 ms versus the required 31.25 ms. Complete-evaluation, trajectory, product, package, and integration work did not run; FP32 activation storage remains production |
record, primitive result, actual-MLP result, representative result | allocation bbe180bf7feb59272a5d5f7afbafb3877afee416; primitive registration/implementation 70a5e4a29c5455ec00a4b757dcdf5cdcc70a5e91 / 312d67024978a64b77d2563dd9386b4328f17d33; MLP registration/correction 606d1e29f56867bfda637c117b58778c634c4ee9 / 0f13bcc486569819df7587349b8b1e049b924ccd; MLP harness/binding 436355ff16fb971d11a959e99e1550abc6186480 / 608bdbca56a428fa243842368631754a62ee67dc; primitive/MLP receipts 8cb2c7c30dec7d179729d5644608fb1f0b9ad5b0478ba92125476936644c775c / 92c27035c18ecddd32ebc6a15e8e732f14f36437ee471cbbc0f698d3ab107bfd; MLP shader aggregate 080fff1d8c115c8d748d93e4f62d4285456fd1e9e7fc7a327799398a5e6e97c3; thermal 27ab8f4c8dd2e338fca7f89672acb3bfec70aab6e852900604684083070f19ea; representative receipt eddb46919d5f281d64c3babe4f4de7d68eabb70d955e19d0f690ba10739270de; thermal 55b2e0099c6aa8cf880a929ce35e8a6848f609caca9423f7d1aa616112a0807d; benchmark-only, experiment closed |
| OPT-0082 | Planner semantic head/sampling | Production reference-BF16 semantic M2 can score and read only the exact 64,000 audio-code rows plus forced EOS, then sample in ascending global-token order without changing any retained logit, browser-v1 decision, emitted token, or Philox cursor | inconclusive | benchmark-only | Current reference-BF16 retained logits/filtering/samples were exact; compact improved complete-token median 379.15 → 340.50 ms (1.11351x) and projected 34.785 s over 900 tokens, with all three cache medians lower. It won only 26/36 pairs versus the frozen 30/36 rule, however, and the developer screen had no attached thermal trace. Stop before trajectory/product/production integration; retain the mechanism only for a materially changed compound experiment |
record, result | allocation 2c062119eb36e4cdf8ae13c7275181857b016244; benchmark code/harness 0a63635a01880aac1b019ffb9f38c027bb5d4d87; production Arm A retained |
| OPT-0083 | Planner dense GEMV | A planner-only one/two-row packed-BF16 GEMV can preserve increasing-K FP32 contractions while eliminating the generic M16 kernel's measured 16x/8x padded arithmetic and approaching model-bandwidth limits without a package repack | negative | abandoned | Exact across 610,344 raw-U32 output visits; direct B won 16/16 and reached 3.605206087x authoritative wall / 7.612403101x diagnostic GPU speedup, but its 10.917928005 GB/s wall effective packed-weight bandwidth missed the frozen 20 GB/s gate despite 29.767441860 GB/s diagnostically on GPU. Panel C also failed. Stop before package-native escalation or production integration; revisit only under a newly allocated, materially changed package-native/no-per-sample-drain experiment |
record, result | allocation 552e977be6b1b5c8b01c346d4aeaa7f63c0edbf2; kernel 1ddb65e751529936b3ef3cd48a6360386c7dd205; harness/protocol be3e9a142bcf28db0a1669cbd6bb8f7393f0a8f5 / 4289adcddd26a91e11399dbd870917b96be9c0a9 / 818d0d3c2be7767db0369f47f4220a93f3f7a67f / 9b96b9f73c70ac7f4aaaf62d6bcc7b4a43539d99 / 211e08aa626b957f44e80b8139690922dd4e037d; accepted receipt 69822104f277df43d976a66e8e931db7e9d09621f20d5a56a6da4e05412bfa1e; invalid first attempt aea1889c55b22dbd0634a902e460e49474f9920c1aa10a795f33421b2022328d; not integrated |
| OPT-0084 | Planner sampling | A reusable candidate-domain browser-v1 sampler with stable typed-array radix ordering can remove boxed O(n log n) sorting, full-vocabulary masking, and multi-megabyte per-token allocation while preserving every FP32 reduction order and categorical decision |
positive | pending-integration | The standalone actual-Chrome sampler gate was exact, won 16/16, reduced its aggregate seven-state median 297.900 → 73.600 ms (4.04755x), and projected 28.395 s/900. The later package-native compact-head plus fused-sampler compound gate was exact across full/compact dispatch, next-token cache witnesses, filtering, weights, tokens, cursors, cancellation, and cleanup; C beat production A in 16/16, reduced the three-position complete-token median 1,238.550 → 1,025.200 ms (1.208106x), improved every position, and projects 60.900 s/900. Its fresh nominal launch passed; all later non-nominal thermal observations are disclosed. Complete trajectory and planner-enabled product gates remain mandatory |
record, standalone result, compound result | candidate 245b5fe3347c370390eff990aa1ed45cb2b869ba; standalone harness 2aaa0454587f24ede91a62dcccb702ba9c06cf62; compound harness 3de2853ff85bc069782361ed2fba592ffdc93ef4; pending integration gates |
| OPT-0085 | Planner scheduling | Reusing the integrated depth-two/four-completion scheduler for the planner's singleton model and readback commands can reduce full-token true drains from 34 to 9 without changing a command buffer, dispatch, arithmetic operation, or FIFO dependency | positive | pending-integration | Actual Chrome reproduced every M1/M2 full/compact/EOS logit raw-U32, status, sample, cursor, topology, cancellation, and lifecycle result; candidate won 16/16, lowered all path medians, and reduced aggregate complete-token median 337.200 → 259.500 ms (1.29942x, 77.700 ms, projected 78.477 s/1,010 draws) |
record, result | implementation/harness 85b20ae6c010d4f9b0f5bd28ae4bdbc70734428f; browser gate passed, trajectory/product gates pending |
| OPT-0086 | Planner-enabled downstream scheduling | The already integrated exact OPT-0080 DiT/VAE depth-two policies depend on authenticated graph/package topology, not on whether conditioning was produced by the planner, so removing the planner-only selector exclusion can recover the proven downstream scheduling saving without changing planner or model math | pending | benchmark-only | Registered before implementation; require selector-negative tests plus planner-enabled control/candidate final-latent, waveform, WAV, cancellation, and topology identity before widening production | record | allocation 1ddb65e751529936b3ef3cd48a6360386c7dd205; no implementation yet |
| OPT-0087 | Planner package-native dense GEMV | Route the already exact direct M1/M2 packed-BF16 kernel through all real planner decode layers and tied-head slices so the package can realize its useful bandwidth without OPT-0083's artificial submit-and-drain boundary around every isolated layer sample | positive | pending-integration | Actual Chrome preserved every full logit, appended K/V-cache evidence word, status, token, Philox cursor, topology, cancellation, and lifecycle result. Direct B won 16/16, lowered every M1/M2 layer/head/model/complete median, reached 2.229363x aggregate layer speedup, and reduced aggregate model-through-readback median 272.150 → 156.200 ms (115.950 ms, projected 117.110 s/1,010). Cleanup balanced 3,683/3,683 buffers and 133/133 maps with zero live resources. The fresh nominal launch passed; all later level-1/2 thermal observations are disclosed. Complete trajectory and planner-enabled product gates remain mandatory |
record, result | allocation 84470d877ba69c6ef7821871f9bdb82d8c757747; unchanged OPT-0083 direct kernel; implementation/harness 980623e0f9ec918f7a537b6a90e341adaad6d4b0; pending integration gates |
| OPT-0088 | Portable device support | Every subgroup-dependent production owner (OPT-0032/0037 dense K4, OPT-0051 K7 row-reuse, OPT-0048 ConvTranspose K4, attention query8/quad-query) can gain a workgroup-memory counterpart consuming the unchanged hosted packages, selected by the existing execution-profile machinery, so adapters without subgroups (Safari, Firefox, iOS) run the production graph instead of failing FEATURE_UNAVAILABLE; compatibility experiment, bounded slowdown expected and reported, shader-f16 stays fail-closed |
pending | pending-integration | Portable dense/K7/ConvTranspose owners landed with test-enforced bit-identical arithmetic (byte-equal WGSL arithmetic sections, re-exported rev7/rev8 index math); attention routes to the existing portable oracle (reordered-rounding vs subgroup reduction). End-to-end masked-subgroups waveform and timing gates pending | record | kernels c272b2d/c48c050/373e90a; selection wiring pending |
| OPT-0089 | DiT weight quantization | Weight-only symmetric int8 (per-32-K-block fp16 scales, round-to-nearest, clamp ±127) fake-quantization of all 264 rev7 DiT GEMM tensors, dequantized in place and run through the completely unchanged production graph, preserves end-to-end 30 s waveform quality within a small numerical envelope, so an int8-resident DiT (~1.51 GB + scales) is a credible answer to the observed iPhone 17 Safari OOM kill at 1,789,925,376 / 3,020,808,192 uploaded bytes (layer 14/24); pure quantization-damage gate, zero kernel changes, distinct mechanism from abandoned OPT-0058 activation-quantized DP4a |
positive | benchmark-only | Per-tensor damage uniform and small: NRMSE 0.00515–0.00634 (median 0.00559), min SNR 43.96 dB, no outlier tensor/family, so no fp16-retention map needed. Determinism gate reproduced the pinned fp16 baseline WAV byte-exactly, then fake-quant vs fp16 on identical seeds gave lo-fi/12345 waveform NRMSE 0.0669 (Pearson 0.99777, LSD ≈3.5 dB, RMS Δ −0.050 dB) and latin/424242 NRMSE 0.2268 (Pearson 0.97460, LSD ≈4.7 dB, RMS Δ +0.089 dB; per-second max 5.19 is a near-silent-ending small-denominator artifact) — trajectory divergence of the 8-evaluation sampler, not noise-like corruption; zero non-finite samples and exact peak parity. Projected int8 DiT phase peak ≈1.862 GB tracked GPU (1.51 GB int8 + 94 MB scales + fp16 norms/shared + measured 127 MB overhead) versus the observed iPhone 17 kill at ≈1.920 GB — plausibly fits, marginal ≈58 MB margin; int8 kernel work justified, listening gate mandatory before any product claim |
record, quant result, waveform result | scripts/requantize-dit-int8.py (repo root); fake-quant package ef8355b9… (models-local, not hosted); benchmark-only, no kernel or production change |
| OPT-0090 | DiT quantization listening preview | The authenticated OPT-0089 fake-quant package can be exposed as an explicit opt-in listening comparison without changing the production default, weakening package identity checks, or claiming packed-int8 size/runtime benefits | positive | integrated | Reproduced the recorded ef8355b9… package and byte-identical quant report, published the immutable full-size artifact, then added an Advanced listening selector with orderly model switching and explicit size/non-production disclosure. 43 web integration, 4 unit, and 2,029 ACE tests plus typecheck, formatting, production build, and all PR checks passed; fresh browser/GPU listening remains external |
record | implementation 2816d6b83185cbde5abd6780a9f5ac903a655154; artifact ac50b5c854fb044ce058acb91d4cd9ab82d99cfa |
| OPT-0091 | Packed weight-only INT8 DiT runtime | Keep the 216 repeated-layer dense matrices as signed INT8 with per-output/K32 FP16 scales and dequantize inside subgroup and portable kernels, preserving the current activation rounding and increasing-K FP32 accumulation while materially reducing the 3.02 GB DiT package | pending | experimental-preview | Deterministic packed package is 1.70 GB (43.7% smaller); manifest, hosted identity, automated tests, and build passed. Actual Chrome/WebGPU, waveform, and listening gates remain external, so production stays default and no quality/mobile claim is made | record | artifact a3233c9f…; HF bc43ba20409825c13d7ef25694d39ac47dd8c9a4 |
Experiment IDs are allocated before code changes, never reused, and never removed from this table.
On 2026-08-13 the owner approved a shader-f16, FP16-first mixed-precision
direction for the optimized M3 profile. Heavy kernels may use FP16 while
reductions and range-sensitive islands remain FP32 as evidence requires; the
packed-BF16/FP32 path remains the oracle until the new profile passes its
quality gates. Direct and default-planner performance are tracked separately.
This decision allocates no experiment and does not erase OPT-0008's
reference-profile attribution.