Best Simulation Benchmarks For VLA Models In 2026
The best VLA simulation benchmarks in 2026 are LIBERO-Plus or LIBERO-PRO, RoboCasa365, RoboTwin 2.0, RoboLab and SimplerEnv. Plain LIBERO is saturated at 97.1%.
The Short Answer
As of October 2026, no single simulation benchmark ranks vision-language-action (VLA) models well on its own. Plain LIBERO is saturated (OpenVLA-OFT reports a 97.1% average across its four suites), so pair it with LIBERO-Plus or LIBERO-PRO for robustness, RoboCasa365 for long-horizon household tasks (the best baseline averages 20.0% across 300 tasks), RoboTwin 2.0 for bimanual work, and SimplerEnv or RoboLab when you need simulated scores that track real-world rankings. BEHAVIOR-1K is the hardest benchmark here, and CALVIN, Meta-World and ManiSkill3 still answer narrower questions well.
The table below compares all of them on the same columns. Every task count, engine and score comes from the project page, repository or paper linked in the Sources list at the end. Where we could not verify a cell, it says n/a. For how to run these benchmarks once you have picked them (rollout counts, seeds, what to publish), see our guide on how to evaluate a VLA policy.
Comparison Table
| Benchmark | Simulator / physics | Tasks | Embodiment | What it tests | Saturation (Oct 2026) | Rough compute |
|---|---|---|---|---|---|---|
| LIBERO | robosuite on MuJoCo | 130 in 4 suites | Franka Panda | Knowledge transfer across spatial, object and goal shifts | Saturated. OpenVLA-OFT 97.1% average | Low. CPU physics, 50 rollouts per task is the usual protocol |
| LIBERO-PRO | LIBERO (MuJoCo) | LIBERO tasks under 4 perturbation dimensions | Franka Panda | Memorization: swapped objects, initial states, instructions, environments | Not saturated. Models above 90% on LIBERO fall to 0.0% | Low |
| LIBERO-Plus | LIBERO (MuJoCo) | 10,030 | Franka Panda | Robustness to 7 factors (layout, camera, robot initial state, language, light, background, sensor noise) | Not saturated. Leaderboard: pi0 53.6%, OpenVLA-OFT 69.6% | Low per rollout, large task count |
| RoboCasa365 | robosuite on MuJoCo | 365 (65 atomic, 300 composite) | PandaOmron mobile manipulator (Franka arm on holonomic base) | Household skills and long-horizon chaining across 2,500 kitchens | Not saturated. GR00T N1.5 20.0% average multi-task | Moderate. Full object registry about 30 GB, 20 episodes per task in LeRobot's protocol |
| SimplerEnv | SAPIEN (ManiSkill2, with a ManiSkill3 GPU port) | 4 Google Robot task groups, 4 WidowX Bridge tasks | Google Robot, WidowX 250 | Whether sim scores rank real-trained policies the way real tests do | Validated on 2024-era policies (RT-1, RT-2-X, Octo); no wrist camera support | Low. ManiSkill3 port runs 10 to 15x faster |
| ManiSkill3 | SAPIEN, GPU-parallel | 12 task domains | Arms, mobile manipulators, humanoids, quadrupeds, dexterous hands | Throughput for RL and imitation learning, plus Bridge digital twins | n/a | Low per rollout. Up to 30,000+ FPS with rendering |
| RoboTwin 2.0 | SAPIEN | 50 dual-arm tasks | Aloha-AgileX, ARX-X5, Piper, Franka, UR5 | Bimanual manipulation, clean vs domain-randomized | Not saturated. Pi0 46.4% Easy, 16.3% Hard | Moderate. 100 rollouts per task per setting |
| CALVIN | PyBullet | 34 tasks, 4 environments (A to D) | Franka Panda | Chaining 5 language instructions in an unseen environment (ABC to D) | Near ceiling. FLOWER 4.53 of 5 average length | Low |
| BEHAVIOR-1K | OmniGibson on Isaac Sim | 1,000 activities; 50 in the 2025 Challenge | Galaxea R1 Pro (Challenge) | Long-horizon household mobile manipulation | Far from saturated. 2025 winner 26% q-score | High. Isaac Sim needs an RTX GPU |
| Meta-World | MuJoCo | 50 | Sawyer | Multi-task and meta-RL (MT1/MT10/MT50, ML1/ML10/ML45) | n/a | Low |
| RoboLab-120 | Isaac Lab on Isaac Sim | 120 (65 simple, 38 moderate, 18 complex) | DROID Franka Panda | Visual, procedural and relational competency of real-trained policies | Not saturated. pi0.5 28.0% overall | High. Isaac Sim needs an RTX GPU |
| Isaac Lab-Arena | Isaac Lab on Isaac Sim | Framework, ships task catalogs including RoboLab | Franka, DROID, GR1, G1 | Authoring benchmarks and running them in parallel | n/a (framework, alpha v0.3) | High. 2,390 env-steps/s with 1,024 environments on one RTX 5880 Ada, camera-free |
"Saturated" here means the top published scores sit close enough to the ceiling that new methods mostly differ by noise. The rollout math makes this concrete: at 97% success, a few points of improvement is inside the confidence interval of a standard 500-rollout suite.
LIBERO And Its Robustness Successors
LIBERO is still the most reported VLA benchmark, which is why it stays on the list. It is cheap to run, every open checkpoint has a LIBERO number, and it is the fastest way to confirm a policy loads and acts sensibly. The problem is that the top of the leaderboard has stopped moving. OpenVLA-OFT reports a 97.1% average across the four suites, and the PolaRiS authors found that nearly every policy they tested scored between 90% and 95% on LIBERO while spanning the full range from strong to weak in real-world tests. A LIBERO score above 90% no longer tells you much about how a policy compares to its peers.
Two follow-up benchmarks keep the LIBERO tasks and assets but perturb them. LIBERO-PRO varies the manipulated objects, initial states, task instructions and environments, and reports that models scoring above 90% on standard LIBERO collapse to 0.0% in its generalized setting, persisting with learned actions even when the target object is replaced. LIBERO-Plus is larger: 10,030 tasks across seven perturbation factors and 21 components, with a public leaderboard that breaks results down per factor. Its headline finding is that policies are most sensitive to camera viewpoint and robot initial state, with performance dropping from 95% to below 30% under modest perturbations, and that many models largely ignore the language instruction.
If you report one LIBERO-family number in 2026, make it a perturbed one, and keep the per-suite and per-factor breakdown. We also found that the most widely used LIBERO training datasets store horizontally mirrored images, so check image orientation before trusting any LIBERO score near zero.
RoboCasa365 For Household Skill
RoboCasa365 (ICLR 2026) extends the original 100-task RoboCasa into 365 tasks: 65 atomic skills and 300 composite tasks organized into 60 activities across six families such as cooking and cleaning. It ships 2,500 pretraining kitchens plus 10 held-out target kitchens, more than 600 hours of human demonstrations and about 1,600 hours of synthetic data, all on a PandaOmron mobile manipulator in MuJoCo. Of the 365 tasks, 220 require mobile manipulation, such as fetching an object at one kitchen station and operating a fixture at another.
This is the benchmark with the most room left. In the paper's multi-task experiment across 300 tasks, GR00T N1.5 averaged 20.0%, pi0.5 16.9%, pi0 15.0% and Diffusion Policy 6.1%, with composite tasks far harder than atomic ones. LeRobot now ships RoboCasa365 as a built-in environment with benchmark-group shortcuts such as atomic_seen and composite_unseen, and recommends 20 episodes per task for reproducible results.
SimplerEnv And RoboLab For Real-World Correlation
Most simulation benchmarks only tell you how a policy does in that simulator. Two on this list were built and validated against real robots.
SimplerEnv (CoRL 2024) recreates the Google Robot and WidowX BridgeData V2 setups that many open policies were trained on, so you can evaluate real-trained policies without fine-tuning them in sim. Its visual matching setup reported an average Pearson correlation of 0.924 with real Google Robot evaluations and 0.890 on Bridge tasks. The limits are the task set (four Google Robot task groups and four Bridge tasks), the 2024-era policies it was validated on, and the lack of wrist camera support, which the PolaRiS authors point out rules out most current VLAs. We cover what to use alongside it in SimplerEnv alternatives.
RoboLab, from NVIDIA Research, takes the DROID route instead. RoboLab-120 is 120 tasks on a simulated DROID Franka with matched wrist and external cameras, tagged by visual, procedural and relational competency and split into 65 simple, 38 moderate and 18 complex tasks. Policies are fine-tuned only on real DROID data, so the benchmark measures real-world generalization rather than in-sim fitting. The best policy, pi0.5, reached 28.0% overall and 13.5% on complex tasks, and dropped to 15.3% when instructions were made vague. The authors compared RoboLab-120 rankings against the real-world RoboArena Elo scores and found the ordering preserved (Spearman 1.00 across five policies), which is strong evidence at the benchmark level from a small number of policies. RoboLab ships as a task catalog in Isaac Lab-Arena, NVIDIA's open-source framework for authoring benchmarks and running them in parallel, which is still labeled alpha (v0.3) with APIs that will change.
ManiSkill3 And RoboTwin 2.0
ManiSkill3 is a simulator first and a benchmark second. Built on SAPIEN with GPU-parallel simulation and rendering, it reports up to 30,000+ FPS and 10 to 1000x speedups over other platforms, across 12 task domains covering tabletop arms, mobile manipulation, humanoids, quadrupeds and dexterous hands. For VLA evaluation, the most relevant part is its set of four BridgeData V2 digital twin tasks derived from SimplerEnv, which the docs say evaluate up to 60x faster than the real world and 10x faster than CPU simulation. It is a good choice when rollout throughput is the bottleneck, less so when you need a standard leaderboard number.
RoboTwin 2.0 is the bimanual benchmark to use. It covers 50 dual-arm tasks across five embodiments, an object library of 731 instances in 147 categories, and a protocol of 100 rollouts per task under Easy (clean) and Hard (clutter, lighting, texture and table-height randomization) conditions. In the paper's own evaluation on the Aloha-AgileX embodiment, pi0 averaged 46.4% on Easy and 16.3% on Hard, and RDT 34.5% and 13.7%, so the randomized setting is a meaningful robustness test.
CALVIN, Meta-World And BEHAVIOR-1K
CALVIN tests long-horizon language following: 34 tasks across four PyBullet environments, evaluated as 1,000 chains of five instructions, with the ABC to D split measuring transfer to an unseen environment. The metric is average completed chain length out of 5, and FLOWER reports 4.53 on ABC to D, which puts the benchmark close to its ceiling.
Meta-World is 50 MuJoCo tasks on a Sawyer arm, organized into multi-task (MT1, MT10, MT50) and meta-learning (ML1, ML10, ML45) splits and now maintained by the Farama Foundation as Meta-World+. It remains a standard for multi-task and meta reinforcement learning, but it was not designed around language-conditioned visual policies, so it says little about a VLA's generalization.
BEHAVIOR-1K is the opposite end of the scale: 1,000 household activities in 50 interactive scenes with more than 10,000 objects, simulated in OmniGibson on Isaac Sim. The 2025 BEHAVIOR Challenge used 50 of these tasks with 10,000 teleoperated demonstrations (more than 1,200 hours) on the Galaxea R1 Pro, and the winning entry, built on pi0.5, reached a 26% q-score. It is the right benchmark if your claim is about long-horizon mobile manipulation, and the most expensive one to run, since Isaac Sim requires an RTX GPU and does not support the A100 or H100.
Which Benchmark For Which Question
| If you want to know... | Run | Why |
|---|---|---|
| Does the policy load and act at all, and how does it compare to older papers? | LIBERO | Every checkpoint has a number, and it is cheap |
| Did it memorize the training scenes? | LIBERO-PRO or LIBERO-Plus | Same tasks, perturbed objects, cameras, instructions and initial states |
| Can it do multi-step household chores? | RoboCasa365 | 300 composite tasks with headroom (best average 20.0%) |
| Will the simulated ranking hold up on a real robot? | SimplerEnv (Google Robot, WidowX) or RoboLab (DROID) | Both report correlation with real-world evaluation |
| How does it handle two arms and domain randomization? | RoboTwin 2.0 | 50 dual-arm tasks, Easy vs Hard protocol |
| Does it follow chained language instructions? | CALVIN | Five-instruction chains in an unseen environment |
| Can it handle long-horizon mobile manipulation? | BEHAVIOR-1K | 50 challenge tasks, winning q-score 26% |
| I need millions of rollouts for RL or ablations | ManiSkill3 | GPU-parallel simulation and rendering |
| I need to build my own benchmark on Isaac Lab | Isaac Lab-Arena | Composable scenes, embodiments and tasks |
A reasonable default for a new VLA paper in 2026 is one perturbed LIBERO variant, RoboCasa365, and one benchmark with published real-world correlation that matches your embodiment. Add RoboTwin 2.0 for bimanual claims and BEHAVIOR-1K for mobile manipulation claims. Choosing training data is a separate decision, covered in our list of the top datasets for robot manipulation.
Where Manifold Fits
Running four or five of these benchmarks usually means four or five harnesses with different simulators, success checks and Python environments. Manifold is Bifrost's evaluation platform that puts them behind one harness. Its supported academic benchmarks include LIBERO, LIBERO-Plus, RoboLab, RoboCasa, RoboMimic, CALVIN, SIMPLER, RoboTwin 2.0, RoboMemory and RLBench, on simulators including NVIDIA Isaac Sim, MuJoCo, ManiSkill and Genesis, with native support for Isaac Lab-Arena. Rollouts are sharded across GPUs, every run gets a manifold:// URI that pins the checkpoint, simulator version, benchmark version and seeds, and failed episodes are clustered by failure mode instead of being reduced to one success rate. We have also precomputed baselines for pi0.5, gr00t-n1.5, openvla-7b and octo-base on LIBERO and RoboCasa. Manifold is in early access, and you can join the waitlist from the Manifold page. More background is in Introducing Manifold.
Sources
- Liu et al., LIBERO, NeurIPS 2023
- Kim et al., OpenVLA-OFT, 2025
- LIBERO-PRO, 2025
- LIBERO-Plus, 2025, and its leaderboard
- Nasiriany et al., RoboCasa365, ICLR 2026, and the LeRobot RoboCasa365 docs
- Li et al., Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER), CoRL 2024
- Tao et al., ManiSkill3, 2024
- Chen et al., RoboTwin 2.0, 2025
- Mees et al., CALVIN, RA-L 2022, and FLOWER, 2025
- BEHAVIOR-1K and the 1st place 2025 BEHAVIOR Challenge solution
- Meta-World, Farama Foundation
- RoboLab, 2026
- Isaac Lab-Arena documentation, NVIDIA
- Jain et al., PolaRiS, 2025
Frequently Asked Questions
Is LIBERO still a good benchmark for VLA models?
LIBERO is still useful as a smoke test and for comparing against older papers, but its standard suites are saturated. OpenVLA-OFT reports a 97.1% average across the four suites, and the PolaRiS authors found that nearly every policy they tested scored between 90% and 95% on LIBERO while spanning the full range in real-world tests. Report LIBERO alongside a perturbed variant such as LIBERO-Plus or LIBERO-PRO rather than on its own.
What is the difference between LIBERO-PRO and LIBERO-Plus?
Both perturb the original LIBERO tasks to test whether a policy memorized its training scenes. LIBERO-PRO perturbs four dimensions (manipulated objects, initial states, task instructions and environments) and reports models above 90% on standard LIBERO falling to 0.0% in its generalized setting. LIBERO-Plus is larger, with 10,030 tasks across seven perturbation factors including camera viewpoint, robot initial state, lighting, background and sensor noise, and it publishes a per-factor leaderboard.
Which simulation benchmark correlates best with real-world robot performance?
SimplerEnv and RoboLab are the two benchmarks on this list with published real-world comparisons. SimplerEnv's visual matching setup reported an average Pearson correlation of 0.924 against real Google Robot evaluations, and RoboLab-120 preserved the policy ranking from the real-world RoboArena evaluation with a Spearman correlation of 1.00 across five DROID policies. Most other benchmarks here have no published sim-to-real correlation.
Which benchmark should I use for bimanual manipulation?
RoboTwin 2.0 is the most complete option, with 50 dual-arm tasks, five dual-arm embodiments and a clean (Easy) versus domain-randomized (Hard) protocol of 100 rollouts per task. Pi0 averaged 46.4% on Easy and 16.3% on Hard in the paper's own evaluation, so it is far from saturated. BEHAVIOR-1K also requires bimanual control, but it targets long-horizon household mobile manipulation and needs much more compute.
Do I need an RTX GPU to run these benchmarks?
Only for the Isaac Sim based ones. BEHAVIOR-1K runs on OmniGibson, and RoboLab and Isaac Lab-Arena run on Isaac Lab, all of which sit on Isaac Sim, and NVIDIA's requirements state that GPUs without RT cores such as the A100 and H100 are not supported. LIBERO, RoboCasa, CALVIN, Meta-World, SimplerEnv, ManiSkill3 and RoboTwin 2.0 run on MuJoCo, PyBullet or SAPIEN, so the RT-core requirement does not apply to them.