Basic skills
Can it carry out a simple instruction?
Grasping, placing, opening, putting away on request, in a familiar environment.
Intelligent robotics observatory
Computed from 3,023 published results for 1,320 models. Percentages are success ratesSuccess rateOut of 100 attempts, how many times the robot completed the task. on benchmarksBenchmarkA standardized test: the same tasks for everyone, so that models can be compared fairly. in simulationSimulationA virtual world (like a realistic video game) where tests can be run thousands of times without breaking real hardware..
level reliably reached, in the real world
Level 2 (follows instructions): 85–99% (sim.) · average 90% · threshold 99% and real-world evidence.
From the programmed factory robot to the science-fiction robot. A level counts as reached only with two kinds of evidence: product-grade reliability (at least 99% success, the vertical line) on all of its capabilities, and a demonstration outside simulation, with real customers.
Repeats a pre-coded movement, without perceiving or understanding anything.
Welding robot on a car assembly line.
Reproduces a movement demonstrated by a human, in a specific place.
Folding a towel after about a hundred demonstrations.
Understands a written instruction and performs varied tasks in a familiar environment.
“Put the mug in the top drawer.”
Remains effective when the scenery, the objects or the robot change; uses both hands; manages in an unfamiliar home.
Clearing the table in a kitchen it has never seen.
Remembers, plans long missions and corrects its errors without human help.
Tidying the whole house over an afternoon, unsupervised.
The science-fiction robot (the Terminator, Asimov's androids): any task, anywhere, for days, improvising in the face of the unknown.
No existing benchmark measures it.
embodiedrank's reading grid, for guidance only: it is not an official standard. Records often come from different models, fine-tunedFine-tuningRetraining a model on the data of a specific test before taking it. for each benchmark — no robot yet holds all of these records at once.
A mission is a chain of steps, and each step can fail. Probability of finishing it without help = (success per step)number of steps. Between 90% and 100%, each point gained changes the outcome by an order of magnitude.
| Success per step | 10-step mission | 30-step mission | 100-step mission |
|---|---|---|---|
| 90% | 35% | 4.2% | < 0.01% |
| 95% | 60% | 21% | 0.59% |
| 99% | 90% | 74% | 37% |
| 99.9% | 99% | 97% | 90% |
Going from 90% to 99% cuts errors by a factor of 10; from 99% to 99.9%, by another 10. Each additional “9” represents the same factor of 10.
With a 90% success rate per step, over 30 steps, the robot finishes on its own 4 times out of 100.
Simplified calculation: independent steps and no error correction. A robot that can recover does better; a robot that causes damage does worse. The number of steps per mission is an order of magnitude.
| Number of steps | Product target (99%) | Simulation record (90%) | Real, unfamiliar home (56%) |
|---|---|---|---|
| 1 | 99% | 90% | 56% |
| 5 | 95% | 59% | 5.5% |
| 10 | 90% | 35% | 0.30% |
| 15 | 86% | 21% | 0.02% |
| 20 | 82% | 12% | < 0.01% |
| 25 | 78% | 7.2% | < 0.01% |
| 30 | 74% | 4.2% | < 0.01% |
| 35 | 70% | 2.5% | < 0.01% |
| 40 | 67% | 1.5% | < 0.01% |
| 45 | 64% | 0.87% | < 0.01% |
| 50 | 61% | 0.52% | < 0.01% |
| 55 | 58% | 0.30% | < 0.01% |
| 60 | 55% | 0.18% | < 0.01% |
| 65 | 52% | 0.11% | < 0.01% |
| 70 | 49% | 0.06% | < 0.01% |
| 75 | 47% | 0.04% | < 0.01% |
| 80 | 45% | 0.02% | < 0.01% |
| 85 | 43% | 0.01% | < 0.01% |
| 90 | 40% | < 0.01% | < 0.01% |
| 95 | 38% | < 0.01% | < 0.01% |
| 100 | 37% | < 0.01% | < 0.01% |
Human reference values and the best robot result published in 2025–2026, with sources.
100% on every axis = the ideal autonomous robot. Shaded area = best published results, in simulation only.
| Capability | Ideal autonomous robot | Best published result |
|---|---|---|
| Basic skills | 100% | 90% |
| Robustness | 100% | 81% |
| Realism | 100% | 93% |
| Two arms | 100% | 96% |
| Home | 100% | 74% |
| Memory | 100% | 52% |
| Long missions | 100% | 25% |
Each spoke is a capability; 100% = perfect success. The shaded area = best result published in simulation for each capability. Records come from different models, often fine-tuned for each benchmark.
| Capability | Published record | Gap to 100% |
|---|---|---|
| Two arms | 96% | 4 pts to go |
| Realism | 93% | 7 pts to go |
| Basic skills | 90% | 10 pts to go |
| Robustness | 81% | 19 pts to go |
| Home | 74% | 26 pts to go |
| Memory | 52% | 48 pts to go |
| Long missions | 25% | 75 pts to go |
What each benchmark checks, the current record and how it has progressed.
Can it carry out a simple instruction?
Grasping, placing, opening, putting away on request, in a familiar environment.
Does it hold up when conditions change?
Different lighting, camera, scenery or instruction: does it truly understand, or does it repeat by rote?
Would its results hold on a real robot?
Simulations calibrated to predict performance on real machines.
Can it coordinate two hands?
Handing over, holding with one hand while acting with the other, manipulating with both arms.
Can it manage in an unfamiliar home?
Varied kitchens, everyday objects, common sense.
Does it remember what it has seen or done?
Remembering a position, counting, resuming a task where it left off.
Can it carry out a long mission on its own?
Planning dozens of steps, handling the unexpected, replanning.