| title | Structured-Output Reliability Across Sampling Configs (NPU LLM stack) | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| date | 2026-06-14 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| kernel | 6.17.0-35-generic | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| hardware | AMD RYZEN AI MAX+ 395 (Strix Halo, gfx1151) | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| xrt_version | 2.21.75 | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| headline | Greedy decoding is the most reliable for structured (JSON) output across small models; aggressive repetition penalties (rep≈1.3) collapse valid output to near zero. Failures are structural, not creative. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| viz |
|
Honest method caveat: isolated structured-output probe, N=3 trials per cell, mean valid-output rate across 4 input spans — not an end-to-end task-correctness measure.
Greedy decoding is the most reliable for structured (JSON) output across small models; aggressive repetition penalties (rep≈1.3) collapse valid output to near zero. Failures are structural, not creative.