Skip to content

Latest commit

 

History

History
32 lines (25 loc) · 2.03 KB

File metadata and controls

32 lines (25 loc) · 2.03 KB

RoboGate Leaderboard — Does Your Policy Survive Industrial Isaac Sim?

15 real-inference learned policies — VLA, diffusion, and world-model — all score 0/68 on RoboGate's 68-scenario industrial Pick&Place suite (NVIDIA Isaac Sim, Franka Panda), despite near-perfect scores on academic benchmarks. The collapse is architecture-independent.

Scripted analytic-IK baseline: 83.8% (57/68) — 100% on nominal — proving the task is solvable in-harness and every 0/68 is a genuine policy failure, not a pipeline artifact.

Policy SR Result Grasp-miss Inference Run date
Scripted IK baseline (control) 83.8% 57/68 measured
cosmos-policy-2b 0.0% 0/68 100% 492 ms 2026-06-13
gr00t-n1.6 0.0% 0/68 100% 2026-04-03
groot-n1.7-libero10 0.0% 0/68 100% 49 ms 2026-06-07
groot-n1.7-liberogoal 0.0% 0/68 100% 48 ms 2026-06-07
groot-n1.7-liberoobject 0.0% 0/68 100% 49 ms 2026-06-07
groot-n1.7-liberospatial 0.0% 0/68 100% 48 ms 2026-06-07
octo-base 0.0% 0/68 81% 291 ms 2026-03-24
octo-small 0.0% 0/68 96% 157 ms 2026-03-23
openvla-7b 0.0% 0/68 100% 311 ms 2026-03-24
openvla-oft-libero10 0.0% 0/68 100%
openvla-oft-liberogoal 0.0% 0/68 100%
openvla-oft-liberoobject 0.0% 0/68 100%
openvla-oft-liberospatial 0.0% 0/68 100%
pi0-base 0.0% 0/68 100% 60 ms 2026-04-02
smolvla-base 0.0% 0/68 3% 18 ms 2026-03-29

Real runs only. Mock / unverified excluded. Source: https://robogate.io/vla

Submit your model

Open a submission issue with your model's HF ID and how to run it. We run it on the real 68-scenario suite and publish the result — whatever it is.

Full method + data: paper · failure explorer.