15 real-inference learned policies — VLA, diffusion, and world-model — all score 0/68 on RoboGate's 68-scenario industrial Pick&Place suite (NVIDIA Isaac Sim, Franka Panda), despite near-perfect scores on academic benchmarks. The collapse is architecture-independent.
Scripted analytic-IK baseline: 83.8% (57/68) — 100% on nominal — proving the task is solvable in-harness and every 0/68 is a genuine policy failure, not a pipeline artifact.
| Policy | SR | Result | Grasp-miss | Inference | Run date |
|---|---|---|---|---|---|
| Scripted IK baseline (control) | 83.8% | 57/68 | — | — | measured |
| cosmos-policy-2b | 0.0% | 0/68 | 100% | 492 ms | 2026-06-13 |
| gr00t-n1.6 | 0.0% | 0/68 | 100% | — | 2026-04-03 |
| groot-n1.7-libero10 | 0.0% | 0/68 | 100% | 49 ms | 2026-06-07 |
| groot-n1.7-liberogoal | 0.0% | 0/68 | 100% | 48 ms | 2026-06-07 |
| groot-n1.7-liberoobject | 0.0% | 0/68 | 100% | 49 ms | 2026-06-07 |
| groot-n1.7-liberospatial | 0.0% | 0/68 | 100% | 48 ms | 2026-06-07 |
| octo-base | 0.0% | 0/68 | 81% | 291 ms | 2026-03-24 |
| octo-small | 0.0% | 0/68 | 96% | 157 ms | 2026-03-23 |
| openvla-7b | 0.0% | 0/68 | 100% | 311 ms | 2026-03-24 |
| openvla-oft-libero10 | 0.0% | 0/68 | 100% | — | — |
| openvla-oft-liberogoal | 0.0% | 0/68 | 100% | — | — |
| openvla-oft-liberoobject | 0.0% | 0/68 | 100% | — | — |
| openvla-oft-liberospatial | 0.0% | 0/68 | 100% | — | — |
| pi0-base | 0.0% | 0/68 | 100% | 60 ms | 2026-04-02 |
| smolvla-base | 0.0% | 0/68 | 3% | 18 ms | 2026-03-29 |
Real runs only. Mock / unverified excluded. Source: https://robogate.io/vla
Open a submission issue with your model's HF ID and how to run it. We run it on the real 68-scenario suite and publish the result — whatever it is.
Full method + data: paper · failure explorer.