Skip to content
embodiedrank – home

Intelligent robotics observatory

How far are we from a truly autonomous robot?

record success rate in simulation on simple instructions
90%
tasks completed by the best humanoid tested in 30 unfamiliar homes (Figure 03, Sep. 2026) source
56%
time taken by an autonomous robot compared with a human on everyday tasks (Jan. 2026) source
4× to 10×
long household missions completed end to end, in simulation (BEHAVIOR Challenge 2025) source
12.4%
best public-evidence score among 71 humanoids (Figure AI Figure 03)
10 / 18
benchmarks measuring human-level general autonomy
0

Computed from 3,023 published results for 1,320 models. Percentages are success rates on benchmarks in simulation.

Distance to an autonomous robot
1/ 5

level reliably reached, in the real world

  1. 0
  2. 1
  3. 2
  4. 3
  5. 4
  6. 5
programmed robot“Terminator”

Level 2 (follows instructions): 85–99% (sim.) · average 90% · threshold 99% and real-world evidence.

The autonomy ladder

From the programmed factory robot to the science-fiction robot. A level counts as reached only with two kinds of evidence: product-grade reliability (at least 99% success, the vertical line) on all of its capabilities, and a demonstration outside simulation, with real customers.

  1. Level 0Achieved

    Programmed

    Repeats a pre-coded movement, without perceiving or understanding anything.

    Welding robot on a car assembly line.

  2. We are here
    Level 1Achieved

    Learns a skill

    Reproduces a movement demonstrated by a human, in a specific place.

    Folding a towel after about a hundred demonstrations.

  3. Level 285–99% (sim.)

    Follows instructions

    Understands a written instruction and performs varied tasks in a familiar environment.

    “Put the mug in the top drawer.”

    In simulation
    90% on averageweak spot: basic skills (90%)
    In the real worldLimited pilots

    In factories, humanoids repeat a tightly defined task (Figure at BMW, Digit at GXO). But in 30 real, unfamiliar homes, the best humanoid tested (Figure 03) completes only 56% of tasks (Sep. 2026).

  4. Level 350–85% (sim.)

    Adapts

    Remains effective when the scenery, the objects or the robot change; uses both hands; manages in an unfamiliar home.

    Clearing the table in a kitchen it has never seen.

    In simulation
    86% on averageweak spot: home (74%)
    In the real worldLab demos

    π0.5 cleaned kitchens and bedrooms in homes it had never seen, on tasks lasting 2 to 5 minutes, with frequent errors (April 2025). No commercial deployment.

  5. Level 4< 50% (sim.)

    Autonomous over time

    Remembers, plans long missions and corrects its errors without human help.

    Tidying the whole house over an afternoon, unsupervised.

    In simulation
    39% on averageweak spot: long missions (25%)
    In the real worldNo real-world evidence

    No real-world demonstration of a multi-hour mission. Even in simulation, the winner of the 2025 BEHAVIOR Challenge completes only 12.4% of household tasks end to end.

  6. Level 5Not measurable

    General autonomy

    The science-fiction robot (the Terminator, Asimov's androids): any task, anywhere, for days, improvising in the face of the unknown.

    No existing benchmark measures it.

    No benchmark can measure it yet.

embodiedrank's reading grid, for guidance only: it is not an official standard. Records often come from different models, fine-tuned for each benchmark — no robot yet holds all of these records at once.

Why 90% is not enough: the last 10 percent

A mission is a chain of steps, and each step can fail. Probability of finishing it without help = (success per step)number of steps. Between 90% and 100%, each point gained changes the outcome by an order of magnitude.

Chances of completing a mission without a single error, by success rate per step
Success per step10-step mission30-step mission100-step mission
90%35%4.2%< 0.01%
95%60%21%0.59%
99%90%74%37%
99.9%99%97%90%
Errors over a day of 200 actions
  • 90%20 errors
  • 95%10 errors
  • 99%2 errors
  • 99.9%0.2 errors (1 every 5 days)

Going from 90% to 99% cuts errors by a factor of 10; from 99% to 99.9%, by another 10. Each additional “9” represents the same factor of 10.

Simulator

90%
30
Chances of finishing the mission without help
4.2%

With a 90% success rate per step, over 30 steps, the robot finishes on its own 4 times out of 100.

Simplified calculation: independent steps and no error correction. A robot that can recover does better; a robot that causes damage does worse. The number of steps per mission is an order of magnitude.

Mission success rate by mission length
  • Product target (99%)
  • Simulation record (90%)
  • Real, unfamiliar home (56%)
Show chart data: Chances of finishing the mission without help, by number of steps
Chances of finishing the mission without help, by number of steps
Number of stepsProduct target (99%)Simulation record (90%)Real, unfamiliar home (56%)
199%90%56%
595%59%5.5%
1090%35%0.30%
1586%21%0.02%
2082%12%< 0.01%
2578%7.2%< 0.01%
3074%4.2%< 0.01%
3570%2.5%< 0.01%
4067%1.5%< 0.01%
4564%0.87%< 0.01%
5061%0.52%< 0.01%
5558%0.30%< 0.01%
6055%0.18%< 0.01%
6552%0.11%< 0.01%
7049%0.06%< 0.01%
7547%0.04%< 0.01%
8045%0.02%< 0.01%
8543%0.01%< 0.01%
9040%< 0.01%< 0.01%
9538%< 0.01%< 0.01%
10037%< 0.01%< 0.01%

Human vs. best robot, point by point

Human reference values and the best robot result published in 2025–2026, with sources.

Human vs. best published robot, dimension by dimension
DimensionHumanBest published robotDetails
Endurance / energy autonomy8-hour shift≈ 4–5 h per charge
Details: Endurance / energy autonomy

Human: A worker does an 8-hour shift, including breaks, and works about 5 shifts a week with just a meal to "recharge." An elite marathoner runs for 2 hours without stopping.

Robot: Manufacturer-claimed range per charge: Figure 03 about 5 hours (manufacturer, Oct. 2025). 1X NEO about 4 hours (manufacturer, 2025). Agility Digit 5: 90 minutes of operation, 9 minutes of charging, and more than 20 productive hours per day claimed (unveiled Sept. 2026, deliveries planned for 2027). UBTech Walker S2: swaps its own battery in about 3 minutes, with no human intervention or downtime, hence the promise of 24/7 operation (July 2025). UBTech does not publish per-battery range. Counterexample: according to an Understanding AI report, Unitree G1 units overheat after 5 to 15 minutes of intense effort and need to rest for about 10 minutes.

Task execution speedbaseline4× to 10× human time
Details: Task execution speed

Human: This is the benchmark: a human folds a t-shirt, makes a sandwich, or opens a door in a few seconds to a few dozen seconds.

Robot: "Humanoid Olympics" challenges completed by Physical Intelligence with its π0.6 model, autonomously (Jan. 2026). Ratio to human time: self-closing door 4 times slower, peanut butter sandwich 4 times, key 5 times, window 5 times, turning a sock right-side out 6.7 times, folding a t-shirt 10 times. According to Understanding AI, these robots took 4 to 10 times longer than a human on nearly every task and only succeeded 52% of the time. In a warehouse, a Galbot took about 40 seconds per item. In a factory, Figure 02 at BMW had to keep to an imposed 84-second cycle, including 37 seconds of loading.

Reliability under real-world conditions—no audited intervention rate published
Details: Reliability under real-world conditions

Human: A trained worker makes very few errors on a repetitive task and handles the unexpected alone. I did not find a standardized, comparable rate.

Robot: Figure 02 at BMW Spartanburg (results published Nov. 19, 2025): 11 months of deployment, more than 1,250 hours of operation, 10-hour shifts Monday through Friday, more than 90,000 parts loaded, contributing to more than 30,000 BMW X3s. The stated targets were more than 99% success per shift and zero interventions per shift, but the rates actually achieved are not published in detail. The part that failed most often: the forearm. Agility Digit at GXO: more than 100,000 totes moved (Nov. 2025), with no published availability rate or number of interventions. No public, audited "mean time between interventions" was found.

Hand dexterity≈ 27 joints20–22 joints per hand
Details: Hand dexterity

Human: The human hand has about 27 degrees of freedom, a figure often cited. Its glabrous skin contains about 17,000 mechanoreceptors (the classic figure from Johansson and Vallbo, 1979). It threads a needle, peels an orange, and uses almost any tool without specific training.

Robot: Tesla Optimus V3: a tendon-driven hand with 22 degrees of freedom, according to announcements and patents. The unveiling was pushed back several times in 2026, and the claim has not been independently verified. Figure 03 (Oct. 2025): tactile sensors detecting about 3 grams (the weight of a paperclip) and cameras in the palms. 1X NEO: hands with a claimed 22 degrees of freedom. Notably, Physical Intelligence succeeded at fine-motor challenges (using a key, peeling an orange, turning a sock right-side out, cleaning a greasy pan) with simple two-finger grippers and vision alone, with no force feedback.

Locomotion (running, terrain, falls)57:20 (record)half marathon 50:26 (2026)
Details: Locomotion (running, terrain, falls)

Human: World records: 100 m in 9.58 s (Usain Bolt), half marathon in 57:20 (Jacob Kiplimo, Lisbon, March 2026). An ordinary human walks on any terrain, climbs stairs, and gets back up without thinking about it.

Robot: Beijing E-Town half marathon. In April 2026, Honor's "Lightning" robot ran it in 50:26 under autonomous navigation, faster than the human record. Only about 40% of the more than 300 robots entered were autonomous; the rest were remote-controlled. In April 2025, Tiangong Ultra had won in 2 hours 40 minutes with 3 battery swaps, and it was the only robot to finish within the 3-hour-10-minute limit. Sprint: Tiangong Ultra ran the 100 m in 8.64 s at the 2nd World Humanoid Robot Games (Beijing, Aug. 26, 2026); whether it was autonomous or remote-controlled is not specified. In 2025, Unitree H1 had won the 1,500 m (6:34.40) and the 400 m (1:28.03) while teleoperated. Falls and collisions remain frequent: a robot fell at the start, a robot ran into a barrier.

Learning a new task1 to a few demos≈ 176 demos / 8 h of data
Details: Learning a new task

Human: An adult learns a simple manual task after one or a few demonstrations, or even from a spoken instruction alone.

Robot: Physical Intelligence: turning a sock right-side out took 176 successful demonstrations, about 8 hours of data. Most "Humanoid Olympics" challenges each required less than 9 hours of data, via fine-tuning of π0.6 (late 2025-Jan. 2026). π0.5 (April 2025) was pretrained on about 400 hours of mobile-manipulator data, collected in about 100 homes, plus web data and data from other robots. According to Understanding AI, a warehouse company needed about 5 minutes of training per new item. π0.7 (April 2026) claims a degree of "zero-shot" generalization to never-seen tasks; these results have not been independently verified.

Decision autonomy / long-horizon tasks—12.4% in simulation
Details: Decision autonomy / long-horizon tasks

Human: A human strings together hours of related tasks, for example tidying an entire house or preparing a full meal, adapting to the unexpected along the way.

Robot: π0.5 (Physical Intelligence, April 2025): cleaning kitchens and bedrooms in 3 real, never-seen homes, with tasks lasting 2 to 5 minutes. The metric used is "task progress" (share of steps completed successfully), not full success; the paper acknowledges frequent errors. The often-cited "94%" figure comes from a different evaluation, run in test environments, and does not measure full success in a real home. BEHAVIOR Challenge 2025 (NeurIPS, in simulation, 50 household tasks averaging about 6.6 minutes): the winning team (Robot Learning Collective) achieved 12.4% full completion and a partial score of 0.26. According to Understanding AI, Physical Intelligence's models "remember" about 15 minutes at most.

Safety around humans—ISO 25785-1 standard not finalized
Details: Safety around humans

Human: A human coworker perceives, anticipates, and holds back their strength. Workplace accidents happen but are governed by mature law and standards.

Robot: No dedicated standard has been finalized. ISO 25785-1 (dynamically stable mobile industrial robots, legged or otherwise) was still at committee-draft stage (ISO/CD) as of September 2026, and Part 2, on integration, has not yet been drafted. Digit 5 advertises an independent safety controller, human detection, and a seated fallback position in case of danger. Gruendel v. Figure AI complaint (Nov. 21, 2025): the company's former head of product safety alleges that Figure 02 impact tests measured force more than 20 times the pain threshold and more than 2 times what fractures an adult skull. He also alleges that one strike left about a 6 mm gash in a steel refrigerator door near an employee. Figure disputes this and says he was terminated for poor performance.

Cost≈ $45/h (US, fully loaded cost)from ≈ $5,000
Details: Cost

Human: In the United States, a private-sector employee costs about $45 per hour, including benefits, according to the BLS (a rough figure, to verify for the exact year). A low-skill warehouse position costs more like $20 to $30 per hour.

Robot: Cheaper options: Unitree R1 at about $4,900 to $5,900 (R1-D version at $4,290), Unitree G1 starting at about $13,500, Unitree H2 at $29,900. Consumer: 1X NEO at $20,000 or $499 per month, deliveries announced for 2026. Industrial: Walker S2 has no public price; UBTech's 2025 accounts suggest about $106,000 in revenue per unit sold, service included. Digit and Figure are estimated at more than $150,000 or are leased to companies. Tesla Optimus is not commercially available.

Hidden or partial teleoperation—≈ 60% of robots remote-controlled (2026 half marathon)
Details: Hidden or partial teleoperation

Human: Not applicable.

Robot: Tesla "We, Robot" event (Oct. 2024): the bartending, chatty Optimus units were remotely teleoperated, according to Bloomberg, The Verge, and Morgan Stanley analyst Adam Jonas; one robot even admitted it. In January 2026, Elon Musk admitted that no Optimus was doing useful work at Tesla. 1X NEO (Oct. 2025): "Expert mode" has a human operator remotely pilot the robot for unfamiliar tasks. The Wall Street Journal reporter never saw NEO do anything autonomously during its demo. Unofficial estimates put autonomy at about 60-70% at launch. Competitions: at the 2025 World Humanoid Robot Games, many robots were joystick-controlled, including Unitree H1 for its race wins. At the 2026 half marathon, about 60% of robots were remote-controlled.

In simulation: the ideal robot vs. the records

100% on every axis = the ideal autonomous robot. Shaded area = best published results, in simulation only.

  • Ideal autonomous robot
  • Best published result
Show chart data: Scores by capability (0 to 100)
Scores by capability (0 to 100)
CapabilityIdeal autonomous robotBest published result
Basic skills100%90%
Robustness100%81%
Realism100%93%
Two arms100%96%
Home100%74%
Memory100%52%
Long missions100%25%

How to read this chart

Each spoke is a capability; 100% = perfect success. The shaded area = best result published in simulation for each capability. Records come from different models, often fine-tuned for each benchmark.

Published record by capability
CapabilityPublished recordGap to 100%
Two arms96%4 pts to go
Realism93%7 pts to go
Basic skills90%10 pts to go
Robustness81%19 pts to go
Home74%26 pts to go
Memory52%48 pts to go
Long missions25%75 pts to go

Capability by capability

What each benchmark checks, the current record and how it has progressed.

Basic skills

Can it carry out a simple instruction?

90%
85–99% (sim.)
record in simulationline: product reliability (99%)
Record trend since Jun 2023

Grasping, placing, opening, putting away on request, in a familiar environment.

  • Instruction following · LIBEROrecord: LaST-R1100%
  • Precision skills · RLBenchrecord: BridgeVLA++94%
  • Grasping varied objects · ManiSkill2record: GeoVLA77%

Robustness

Does it hold up when conditions change?

81%
50–85% (sim.)
record in simulationline: product reliability (99%)
Record trend since Oct 2025

Different lighting, camera, scenery or instruction: does it truly understand, or does it repeat by rote?

  • Robustness to the unexpected · LIBERO-Plusrecord: RoboHarness92%
  • Traps and perturbations · LIBERO-Prorecord: QuoVLA70%

Realism

Would its results hold on a real robot?

93%
85–99% (sim.)
record in simulationline: product reliability (99%)
Record trend since May 2024

Simulations calibrated to predict performance on real machines.

  • Small arm: realism · SimplerEnv · WidowX VMrecord: StARe-VLA IPI (OpenVLA)98%
  • Google robot: realism · SimplerEnv · Google Robot VMrecord: TBD-VLA91%
  • Google robot: varied scenery · SimplerEnv · Google Robot VArecord: GRACE (GPT-4o)90%

Two arms

Can it coordinate two hands?

96%
85–99% (sim.)
record in simulationline: product reliability (99%)
Record trend since Jun 2025

Handing over, holding with one hand while acting with the other, manipulating with both arms.

  • Two arms, clean scene · RoboTwin 2.0 · Easyrecord: MotuBrain96%
  • Two arms, cluttered scene · RoboTwin 2.0 · Hardrecord: MotuBrain96%

Home

Can it manage in an unfamiliar home?

74%
50–85% (sim.)
record in simulationline: product reliability (99%)
Record trend since Jun 2024

Varied kitchens, everyday objects, common sense.

  • Virtual kitchen (mobile arm) · RoboCasa · Pandarecord: Z-1 RL81%
  • Virtual kitchen (humanoid) · RoboCasa · GR1 (humanoïde)record: ACE-Ego-073%
  • Common sense and general knowledge · VLABenchrecord: Xiaomi-Robotics-169%

Memory

Does it remember what it has seen or done?

52%
50–85% (sim.)
record in simulationline: product reliability (99%)
Record trend since Feb 2025

Remembering a position, counting, resuming a task where it left off.

  • Working memory · RoboMMErecord: GroundSG+Oracle84%
  • Short-term memory · MIKASA-Roborecord: ELMUR58%
  • Long-term memory · LIBERO-Memrecord: Naive Embodied-SlotSSM15%

Long missions

Can it carry out a long mission on its own?

25%
< 50% (sim.)
record in simulationline: product reliability (99%)

Planning dozens of steps, handling the unexpected, replanning.

  • Long, unpredictable missions · RoboCerebrarecord: GT-plan + OpenVLA* (RoboCerebra Oracle)25%