Robot-eval-bench

Which models can tell whether a robot did the task? We rank frontier VLMs as robot judges, keyframes versus video, across 16 open datasets.

Leaderboard

94%
Gemini 3.6 flash
93%
Gemini 3.7 flash
90%
GPT 5.6 Sol
88%
Gemini 3.1 pro
86%
Kimi 3
83%
Muse Spark 1.1
76%
Claude Opus 5

Bars scaled from 70% to 100% so the spread reads clearly. Hover any bar for the full keyframes / video / cost breakdown. The leader for the selected metric is tinted in its maker's brand color.

Accuracy vs cost per episode

Up and to the left is better: more accurate, cheaper to run. Cost is on a log scale.

60%
80%
100%
$0.001
$0.005
$0.01
$0.05
3.6 flash
3.7 flash
GPT 5.6
3.1 pro
Kimi 3
Muse
Opus 5
cost per episode (USD)

How it works

01

Two approaches

Each judge sees the episode as 4 keyframes, or as video: native where the model accepts it, ~16 frames where it doesn't.

02

Balanced labels

Open datasets are almost all successes, so we add a negative counterpart for each one, giving the judge a balanced set to tell apart.

03

Reproducible

Ground truth, per-episode judgments, the cost model, and the chart code all ship in the repo.

Read more