Robot-eval-bench
Which models can tell whether a robot did the task? We rank frontier VLMs as robot judges, keyframes versus video, across 16 open datasets.
Leaderboard
Bars scaled from 70% to 100% so the spread reads clearly. Hover any bar for the full keyframes / video / cost breakdown. The leader for the selected metric is tinted in its maker's brand color.
Accuracy vs cost per episode
Up and to the left is better: more accurate, cheaper to run. Cost is on a log scale.
How it works
Two approaches
Each judge sees the episode as 4 keyframes, or as video: native where the model accepts it, ~16 frames where it doesn't.
Balanced labels
Open datasets are almost all successes, so we add a negative counterpart for each one, giving the judge a balanced set to tell apart.
Reproducible
Ground truth, per-episode judgments, the cost model, and the chart code all ship in the repo.