Introducing Kite Evals
Evaluation is the slowest part of building a robot. We made it the fastest — a thousand scenarios in under an hour, with failures that explain themselves.
Luigi D'Introno
Founder, Kite
Aug 14, 2026 · 3 min read
You changed one thing in the policy. Now you need to know if it's better.
So you book the robot. You run it. You sit there and watch. Fifty episodes later you have a number you only half-trust, and the week is gone. Meanwhile the four ideas you had on Monday are still waiting, because there's only one robot and one of you.
Evaluation is the slowest step in robotics, and it's slow in the most demoralizing way. Hours of watching. Days of scrubbing video.
Today we're opening Kite Evals: run thousands of scenarios in under an hour, and get back failures and edge cases to make your policy better.
A month of experiments in a week
Kite Evals runs in the cloud, in parallel. A thousand scenarios in parallel in less than one hour. When an eval costs a week, you only test what you're already fairly sure about. When it costs an hour, you test the long shot.
Teams are using these reports to compare dozens of architectures and fine-tuned checkpoints, then deciding which policies have earned more real-world testing, more post-training, and more money. The loop stops moving at the speed of your robot's calendar and starts moving at the speed of compute.
Failures you can query
A pass/fail number tells you a policy is worse. It doesn't tell you why, and the why is the critical part.
So every rollout comes back densely labeled — every atomic sub-goal, every detected action, marked and timestamped. Then agents read across all of them and cluster the failure modes for you.
It drops the object when the lighting is low and the object is reflective. That's the sentence you were going to spend three days earning by hand.
Bring your own simulator
NVIDIA's Isaac Lab Arena and MuJoCo are supported natively, so you can run sensitivity analysis across conditions, objects, environments, and embodiments — and isolate the one environment variable that breaks your policy.
Change the floor, the lighting, the clutter, the robot itself, and run the sweep again. Living room, kitchen, hotel room, grocery store, greenhouse, library — plus eight more presets, or a world you bring yourself:






And there's nothing to stand up. No GPU cluster to provision, no simulators to configure. Sharded, vectorized GPU simulations scale with your workload and disappear when they're done.
We're opening it to a small group
Kite Evals is in early access with a selected group of companies across manufacturing, logistics, lab automation, and retail. They're the ones deciding what we build next.
We're ready to take on a few more.