Full leaderboard
Every published result, test by test. Each row = one modelVLA (Vision-Language-Action)A robot “brain” that looks (camera), reads an instruction and directly decides which movements to make. on one testBenchmarkA standardized test: the same tasks for everyone, so that models can be compared fairly., with its success rateSuccess rateOut of 100 attempts, how many times the robot completed the task.. These are tests run in simulationSimulationA virtual world (like a realistic video game) where tests can be run thousands of times without breaking real hardware.; for real robots, see the Real world page.
Is the average success rate improving?
By month of publication: average of each model's latest known scores, across all tests and then for the 3 most-used tests
Show chart data: Average published success rate, by month
| Month | All benchmarks | Instruction following | Small arm: realism | Two arms, clean scene |
|---|---|---|---|---|
| Sep 23 | 48.5% | — | — | — |
| Oct 23 | 48.5% | — | — | — |
| Nov 23 | 48.5% | — | — | — |
| Dec 23 | 48.5% | — | — | — |
| Jan 24 | 44.2% | 40.0% | — | — |
| Feb 24 | 47.3% | 40.0% | — | — |
| Mar 24 | 47.3% | 40.0% | — | — |
| Apr 24 | 47.3% | 40.0% | — | — |
| May 24 | 44.0% | 40.0% | — | — |
| Jun 24 | 49.7% | 53.3% | — | — |
| Jul 24 | 50.9% | 56.0% | — | — |
| Aug 24 | 51.4% | 56.7% | — | — |
| Sep 24 | 51.4% | 56.7% | — | — |
| Oct 24 | 52.4% | 58.4% | — | — |
| Nov 24 | 51.5% | 58.8% | 19.5% | — |
| Dec 24 | 51.3% | 60.4% | 19.5% | — |
| Jan 25 | 54.5% | 64.3% | 20.5% | — |
| Feb 25 | 55.0% | 70.1% | 27.8% | — |
| Mar 25 | 56.0% | 72.9% | 25.8% | — |
| Apr 25 | 56.8% | 73.0% | 25.8% | — |
| May 25 | 59.7% | 77.2% | 31.4% | — |
| Jun 25 | 58.3% | 78.9% | 36.2% | — |
| Jul 25 | 59.0% | 79.2% | 39.1% | 71.4% |
| Aug 25 | 60.0% | 80.3% | 45.5% | 71.4% |
| Sep 25 | 62.3% | 82.0% | 46.3% | 71.4% |
| Oct 25 | 63.7% | 83.9% | 50.5% | 63.9% |
| Nov 25 | 65.0% | 84.7% | 51.2% | 66.2% |
| Dec 25 | 65.3% | 85.7% | 53.5% | 56.9% |
| Jan 26 | 66.5% | 86.5% | 53.2% | 66.5% |
| Feb 26 | 67.5% | 87.6% | 50.8% | 67.7% |
| Mar 26 | 68.4% | 87.5% | 53.3% | 65.3% |
| Apr 26 | 69.1% | 87.8% | 53.6% | 68.6% |
| May 26 | 70.1% | 87.5% | 54.1% | 68.6% |
| Jun 26 | 71.5% | 88.1% | 53.6% | 71.4% |
| Jul 26 | 72.6% | 88.6% | 53.8% | 73.3% |
| Aug 26 | 72.9% | 89.0% | 53.9% | 73.8% |
All results
Click a column header to sort · hover over a test name to see what it measures
| # | Source | ||||
|---|---|---|---|---|---|
| 1 | LaST-R14B Chen et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.9% | Apr 2026 | |
| 2 | QuoVLA Wang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.6% | May 2026 | |
| 3 | ABot-M0.5 Chen et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.4% | Jul 2026 | |
| 4 | AcceRL Lu et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.4% | Mar 2026 | |
| 5 | CORAL Luo et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.3% | Jul 2026 | Paper Third-party measurement |
| 6 | CORAL (SimVLA)0.8B Luo et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.3% | Mar 2026 | |
| 7 | FUTURE-VLA Fan et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.2% | Feb 2026 | |
| 8 | SRPO (Online) Fei et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.2% | Nov 2025 | |
| 9 | MindLWPI4B Xia et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.1% | Jul 2026 | |
| 10 | SAM3D-VLA Liu et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.1% | Jul 2026 | |
| 11 | PriorVLA Guo et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.1% | May 2026 | |
| 12 | DeVA Zhang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.0% | Jul 2026 | |
| 13 | π₀ + DLAM Tang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.0% | Jul 2026 | |
| 14 | SaiVLA-0 Shi et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.0% | Jul 2026 | Paper Third-party measurement |
| 15 | GeoAlign Chen et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.0% | Jun 2026 | |
| 16 | SimpleVLA-RL7B Li et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.0% | Sep 2025 | |
| 17 | pi_0.5 with MoH3B Jing et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 99.0% | Nov 2025 | |
| 18 | Faster-WAM Zhao et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.9% | Aug 2026 | |
| 19 | DreamWAM Yuan et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.9% | Aug 2026 | |
| 20 | MVUCF Yang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.9% | Aug 2026 | |
| 21 | ChainVLA1.2B Huang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.8% | Aug 2026 | |
| 22 | InternVLA-A1.5 Ma et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.8% | Jul 2026 | |
| 23 | TFP3.3B Liang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.8% | Jul 2026 | |
| 24 | FoMoVLA Li et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.8% | Jul 2026 | |
| 25 | DualCoT-VLA Zhong et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.8% | Mar 2026 | |
| 26 | LingBot-VLA + BCP Xu et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Aug 2026 | |
| 27 | Hermite-VLAReg~3.3B Lv et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Aug 2026 | |
| 28 | CoRE-VLA Zhang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Jul 2026 | |
| 29 | ST-WAM Wang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Jul 2026 | |
| 30 | PearlVLA (K=4) + CRG-PRL Yang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Jun 2026 | |
| 31 | 3DThinkVLA Shi et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Jun 2026 | |
| 32 | Xiaomi-Robotics-04.7B Cai et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.7% | Feb 2026 | |
| 33 | SimVLA0.5B Luo et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Feb 2026 | |
| 34 | IntentVLA Lian et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Jul 2026 | Paper Third-party measurement |
| 35 | Self-Evolving (criticality routing)0.8B He et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Jul 2026 | |
| 36 | ω-EVA Stage 31.2B Sun et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Jun 2026 | Paper Third-party measurement |
| 37 | LaWAM2.3B Chen et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Jun 2026 | |
| 38 | Multi-view VLA + AML4B Xiao et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | May 2026 | Paper Third-party measurement |
| 39 | FocusVLA0.5B Zhang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Mar 2026 | |
| 40 | Cosmos Policy + World2Act Vuong et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Mar 2026 | |
| 41 | DiT4DiT Ma et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Mar 2026 | |
| 42 | ABot-M0 Yang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | Feb 2026 | |
| 43 | Hume Song et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.6% | May 2025 | |
| 44 | Joint-WAM Zhao et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | Aug 2026 | |
| 45 | CofactVLA Zhang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | Aug 2026 | |
| 46 | MindWPI4B Xia et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | Jul 2026 | |
| 47 | PearlVLA (K=4) Yang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | Jun 2026 | |
| 48 | GeoSem-WAM Ma et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | Jun 2026 | |
| 49 | A2World-policy3.0B Huang et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | Jun 2026 | |
| 50 | ElasticFlow Chen et al.VLA model | Instruction following LIBERO Fine-tuned for this test | 98.5% | May 2026 |
Sources: AllenAI · VLA Evaluation Harness · last ingested on Sep 25, 2026, 4:45 PM