Introducing Kite Evals

Evaluation is the slowest part of building a robot. We made it the fastest — a thousand scenarios in under an hour, with failures that explain themselves.

Luigi D'Introno

Founder, Kite

Aug 14, 2026 · 3 min read

Save a month of robot experiments in hours

You changed one thing in the policy. Now you need to know if it's better.

So you book the robot. You run it. You sit there and watch. Fifty episodes later you have a number you only half-trust, and the week is gone. Meanwhile the four ideas you had on Monday are still waiting, because there's only one robot and one of you.

Evaluation is the slowest step in robotics, and it's slow in the most demoralizing way. Hours of watching. Days of scrubbing video.

Today we're opening Kite Evals: run thousands of scenarios in under an hour, and get back failures and edge cases to make your policy better.

Apply for early access

A month of experiments in a week

Kite Evals runs in the cloud, in parallel. A thousand scenarios in parallel in less than one hour. When an eval costs a week, you only test what you're already fairly sure about. When it costs an hour, you test the long shot.

TraditionalWith Kite1,000-scenario sweep~1 week55 min20 checkpoints compared~2 weeks2 hFinding why it failed~3 days of videoone query
What our design partners were doing before Kite Evals, and what the same work costs now.

Teams are using these reports to compare dozens of architectures and fine-tuned checkpoints, then deciding which policies have earned more real-world testing, more post-training, and more money. The loop stops moving at the speed of your robot's calendar and starts moving at the speed of compute.

Failures you can query

A pass/fail number tells you a policy is worse. It doesn't tell you why, and the why is the critical part.

So every rollout comes back densely labeled — every atomic sub-goal, every detected action, marked and timestamped. Then agents read across all of them and cluster the failure modes for you.

It drops the object when the lighting is low and the object is reflective. That's the sentence you were going to spend three days earning by hand.

A humanoid crossing a warehouse aisle.
An arm reaching for a ball on a kitchen floor.
A patrol in a grocery aisle that ends on its back. Better found here than on site.

Bring your own simulator

NVIDIA's Isaac Lab Arena and MuJoCo are supported natively, so you can run sensitivity analysis across conditions, objects, environments, and embodiments — and isolate the one environment variable that breaks your policy.

Change the floor, the lighting, the clutter, the robot itself, and run the sweep again. Living room, kitchen, hotel room, grocery store, greenhouse, library — plus eight more presets, or a world you bring yourself:

Living room
Living room
Kitchen
Kitchen
Hotel room
Hotel room
Grocery store
Grocery store
Greenhouse
Greenhouse
Library
Library

And there's nothing to stand up. No GPU cluster to provision, no simulators to configure. Sharded, vectorized GPU simulations scale with your workload and disappear when they're done.

We're opening it to a small group

Kite Evals is in early access with a selected group of companies across manufacturing, logistics, lab automation, and retail. They're the ones deciding what we build next.

We're ready to take on a few more.

Similar articles