Robot Policy Evaluation Platforms Compared (2026)
Isaac Lab-Arena, RoboLab, LeRobot and SimplerEnv evaluate robot policies in simulation, RoboArena and AutoEval on real robots. The right one depends on your benchmark.
Which Robot Policy Evaluation Platform Should You Use?
As of October 2026, no single platform covers every robot policy evaluation need, so the choice depends on whether you need simulated or real-world results, which benchmarks you report, and whether you want to run the infrastructure yourself. For simulation, the main open tools are NVIDIA's Isaac Lab-Arena (parallel evaluation across thousands of environments on Isaac Sim), RoboLab (a 120-task DROID benchmark that tracks failure events), LeRobot's lerobot-eval (LIBERO, Meta-World and RoboCasa365), SimplerEnv (real-to-sim Google Robot and WidowX tasks), XPolicyLab with RoboDojo, and RLinf. For real robots, RoboArena ranks DROID policies through double-blind comparisons at partner institutions, and AutoEval runs WidowX policies autonomously around the clock. Bifrost's Manifold is a hosted option that runs many simulation benchmarks from one harness with automated failure analysis, and it is in early access.
The table below compares them on the questions that usually decide the choice. Every cell is taken from the tool's own documentation, repository or paper as of October 2026, and cells we could not confirm are marked "Not documented" rather than guessed.
Comparison Table
| Tool | Real or sim | Benchmarks covered | GPU parallelism | Hosted or self-run | Bring your own policy | Baselines | Failure analysis | License and access |
|---|---|---|---|---|---|---|---|---|
| Isaac Lab-Arena (NVIDIA) | Sim (Isaac Sim) | RoboLab-inspired tasks (Arena counterparts of RoboLab's 120 tasks where available), Lightwheel RoboCasa and LIBERO task ports, RoboTwin 2.0, Lightwheel RoboFinals, LeRobot EnvHub environments, your own composed tasks | Thousands of heterogeneous environments per GPU; multi-node through OSMO | Self-run | Yes, behind a policy server, with integrations for GR00T and pi0.5 | None built in | Subtask predicates, sensitivity analysis | Apache 2.0; alpha (v0.3), APIs will change; Isaac Sim has its own terms |
| RoboLab (NVIDIA Research) | Sim (Isaac Lab) | RoboLab-120: 120 tasks on a DROID Franka | Multiple episodes in parallel across environments | Self-run | Yes, policy runs as a server; any Isaac Lab compatible robot | Five policies in the paper; pi0.5 best at 28.0% | Wrong-object grasp, drop and gripper collision events; dashboard with episode replay; sensitivity analysis | Apache 2.0; RTX GPU with 48 GB+ VRAM recommended |
LeRobot lerobot-eval (Hugging Face) | Sim | LIBERO (130 tasks), LIBERO-plus, Meta-World (MT50), RoboCasa365, RoboTwin 2.0, RoboCerebra, RoboMME, VLABench, Isaac Lab-Arena environments via EnvHub | Vectorized environments on one machine (--eval.batch_size) | Self-run | LeRobot-format policies; new benchmarks via a gym wrapper | pi0.5 LIBERO reproduction at 97.5%; SmolVLA RoboCasa checkpoint | Success rate with 95% Wilson intervals, and reward, per task and suite | Apache 2.0 |
| SimplerEnv | Sim (real-to-sim) | Google Robot tasks (pick coke can, move near, drawers) and four WidowX Bridge tasks | CPU-based ManiSkill2 (SAPIEN still needs an NVIDIA GPU); the Bridge tasks on a ManiSkill3 branch are GPU-parallel and 10 to 15x faster | Self-run | Yes, by writing a policy inference script | RT-1 checkpoints and Octo included | None; reports sim-to-real correlation (Pearson, MMRV) | MIT |
| RoboArena | Real | DROID Franka, tasks and scenes chosen by evaluators at seven institutions in the paper | Not applicable | Hosted evaluator network; you host the policy server | Yes, any DROID-compatible policy | Initial pool of seven DROID policies (pi0 and PaliGemma variants) | Free-form evaluator explanations; LLM-assisted policy reports that cite episodes | Open to community submissions; code MIT |
| AutoEval | Real | Bridge WidowX; two public stations with four drawer and eggplant tasks | Cells run in parallel, 24/7 | Hosted; you host the policy server and submit through a dashboard | Yes, matching its observation and action spec | None listed; an example OpenVLA policy server is provided | Success rates, episode videos and frames | Code MIT; the project page lists its four tasks as available until January 1, 2026, with the 2026 tasks to be announced |
| XPolicyLab | Sim and real | RoboTwin 2.0, RoboDojo sim and real, RMBench | Batched policy queries across environments | Self-run | Yes, through an adapter added by pull request | 44 integrated policies as of September 2026 | Not documented | Apache 2.0 |
| RoboDojo | Sim and real | 42 sim tasks across five capability dimensions; 18 real tasks on ARX X5, Piper and Piper X | Heterogeneous parallel simulation on Isaac Sim | Sim self-run; official leaderboard and real-world runs through RoboDojo's cloud evaluation submission | Yes, through XPolicyLab | 30 policies in the paper; public leaderboard | Scores by capability dimension | README states a non-commercial research license, while the repository's LICENSE file is MIT; confirm with the maintainers |
| RLinf | Sim and real | LIBERO (with LIBERO-PRO and LIBERO-Plus), ManiSkill OOD, RoboTwin, BEHAVIOR-1K, PolaRiS, RoboCasa365, Franka real-world | Parallel rollouts; multi-GPU and multi-node framework | Self-run | OpenVLA, OpenVLA-OFT, pi0, pi0.5, GR00T and others | Not documented for evaluation | Success rate, logs and videos | Apache 2.0; built mainly for RL training |
| Lightwheel (LW-BenchHub, RoboFinals) | Sim (built on Isaac Lab-Arena) | 130 LIBERO-derived and 138 RoboCasa-derived tasks; RoboFinals-100, 100 household, factory and retail tasks | Through Isaac Lab-Arena | LW-BenchHub self-run; RoboFinals access by contacting Lightwheel | LW-BenchHub yes, through a server-client policy API | Not documented | Not documented | LW-BenchHub Apache 2.0; RoboFinals license not stated |
| Manifold (Bifrost) | Sim | LIBERO, LIBERO-Plus, RoboLab, RoboCasa, RoboMimic, CALVIN, SIMPLER, RoboTwin 2.0, RoboMemory, RLBench, custom tasks; native Isaac Lab-Arena support | Sharded and vectorized across GPUs | Hosted | Yes, point it at a policy checkpoint | pi0.5 on LIBERO (published) | Per-subgoal rubric scoring; agents cluster failures and rank them by impact | Early access via waitlist; not open source today |
Simulation Frameworks You Run Yourself
Isaac Lab-Arena is NVIDIA's framework for building and running policy evaluations on Isaac Lab. Its main strength is scale, and the documentation reports 2,390 environment-steps per second with 1,024 parallel environments on one RTX 5880 Ada GPU in a preliminary camera-free benchmark, and environments in one batch can differ at the object level, so a single run covers many scene variations. Policies run behind a server, which keeps their dependencies out of the simulator's environment. The trade-offs are that it is labeled alpha (v0.3) with APIs that will break, it needs an RTX GPU because Isaac Sim does not support GPUs without RT cores such as the A100 and H100, and its LIBERO and RoboCasa suites are Lightwheel's adaptations on Isaac Lab rather than the original MuJoCo benchmarks, so treat their scores as a separate benchmark. We cover the throughput side in How To Evaluate Robot Policies In Simulation At Scale.
LeRobot is the lightest path when your policy is already in LeRobot format. lerobot-eval runs every benchmark through the same gym interface, reports success per task and suite, and documents recommended protocols, such as 10 episodes per task across the four standard LIBERO suites and 20 per task on RoboCasa365. Its pi0.5 reproduction reports 97.5% on LIBERO against 96.85% from openpi. Parallelism is per machine, through vectorized environments.
SimplerEnv is narrower and older, and it answers a different question, which is whether simulated scores rank real-trained policies the way real robots do. It covers Google Robot and WidowX Bridge setups only, and as the PolaRiS authors point out, it does not support wrist cameras, which is why newer DROID-focused tools exist. Our SimplerEnv alternatives post compares it with PolaRiS, RoboLab and others.
RLinf is primarily a reinforcement learning framework for VLAs, with an evaluation entry point that runs parallel rollouts in simulation or on a real Franka. It makes the most sense when evaluation is one step inside an RL post-training loop rather than a separate activity.
Fixed Benchmarks With Their Own Tooling
RoboLab is the most diagnostic of the open tools. Policies are fine-tuned only on real DROID data and tested on 120 new tasks, so the benchmark measures generalization rather than in-sim fitting, and its ranking of the four policies measured in both places matched the real-world RoboArena ordering exactly (Spearman 1.00). It records wrong-object grasps, drops and gripper collisions automatically, gives partial credit for completed subtasks, and uses neural posterior estimation to find which scene variables (such as wrist camera pose) drive failures. The paper's main evaluation used 10 episodes per task, and the authors note that this gives a confidence interval of roughly plus or minus 30% on a single task near 50% success.
XPolicyLab and RoboDojo are designed to work together, and RoboDojo uses XPolicyLab as its policy interface. XPolicyLab defines a common adapter so that N policies and M environments need N plus M integrations instead of N times M, and it ships adapters for more than 40 policies. RoboDojo adds 42 simulation tasks and 18 real-world bimanual tasks on standardized hardware that outside teams can access remotely. Check the license before using RoboDojo commercially, since its README describes a non-commercial research license even though the LICENSE file in the repository is MIT.
Real-World Evaluation Networks
RoboArena is the real-world reference that RoboLab and other simulation benchmarks compare their rankings against. Evaluators at partner institutions run double-blind pairwise comparisons on DROID robots with tasks of their choosing, and each comparison records a progress score, a preference and a written explanation. The paper reports more than 600 pairwise episodes across seven policies and finds this ranks policies more accurately than conventional centralized evaluation. It is the most trustworthy signal in this list and the slowest, since every data point costs real robot time.
AutoEval nearly removes the human from real-world evaluation on WidowX arms, using learned success classifiers and reset policies so that stations run around the clock. You host your policy as a server and submit jobs through a web dashboard, and results come back with videos. Its coverage is small (two public stations and four tasks on the project page, which lists those tasks as available until January 1, 2026 and the 2026 set as not yet announced), so check what is running before you plan around it. It suits Bridge-style policies rather than general benchmarking.
How To Choose
| Your situation | Best fit |
|---|---|
| One benchmark (LIBERO, Meta-World, RoboCasa365), a LeRobot policy and one GPU machine | LeRobot lerobot-eval |
| Ranking DROID policies on real robots | RoboArena |
| Real-world numbers for a WidowX Bridge policy without a lab robot | AutoEval, if a station and task match your policy |
| Google Robot or Bridge policies trained on real data, needing a sim proxy | SimplerEnv |
| A DROID policy and a sim benchmark that reports why it failed | RoboLab |
| Authoring your own tasks on Isaac Sim and running thousands of environments | Isaac Lab-Arena |
| Bimanual policies, or the same adapter for sim and real | XPolicyLab with RoboDojo (check the RoboDojo license for commercial use) |
| Evaluation inside an RL post-training loop | RLinf |
| Many benchmarks and simulators, many checkpoints, failure analysis, no infrastructure to run | Manifold |
Several of these combine well. A common pattern is a fast simulated loop on every checkpoint, a perturbation suite to catch memorization, and real-world evaluation on the few candidates that survive. Simulated scores still need checking against real robots, and Best Simulation Benchmarks For VLA Models In 2026 covers which benchmarks have published real-world correlations.
Where Manifold Fits
Manifold is Bifrost's hosted platform for evaluating robot policies in simulation. One harness runs a policy across benchmarks including LIBERO, LIBERO-Plus, RoboLab, RoboCasa, CALVIN, SIMPLER and RoboTwin 2.0, on simulators including NVIDIA Isaac Sim, MuJoCo, ManiSkill and Genesis, with native support for Isaac Lab-Arena. Rollouts are sharded and vectorized across GPUs, so a LIBERO sweep that takes 8 hours comes back in about 30 minutes (roughly 16x), and there is no simulator, cloud or GPU setup on your side. Every task is broken into atomic subgoals scored against rubrics you can define, agents cluster failures and rank them by impact, and each run records the policy and benchmark container images it used, pinned by digest. Published reference results so far cover pi0.5 on LIBERO (pi0.5 LIBERO results on Manifold).
Manifold does not replace real-world evaluation, and it is not open source today. If you run one benchmark on a machine you control, LeRobot or Isaac Lab-Arena will do the job, and RoboArena is the better tool for a real-world ranking. Manifold is in early access, and you can join the waitlist from the Manifold page or read more on the robotics page.
Sources
- NVIDIA, Isaac Lab-Arena repository and documentation
- Yang et al. (NVIDIA), RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies, RSS 2026, and the RoboLab repository
- Hugging Face, LeRobot repository, LIBERO docs, RoboCasa365 docs and Adding a New Benchmark
- Li et al., Evaluating Real-World Robot Manipulation Policies in Simulation, CoRL 2024, and the SimplerEnv repository
- Atreya, Pertsch et al., RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies, 2025
- Zhou, Atreya, Tan et al., AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World, and the AutoEval project page
- XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment, 2026
- RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies, 2026, and the RoboDojo repository
- RLinf repository and evaluation docs
- Lightwheel, LW-BenchHub and RoboFinals Industrial Benchmark
- NVIDIA Isaac Sim, system requirements
Frequently Asked Questions
What is the best open-source robot policy evaluation framework?
It depends on the benchmark. LeRobot's lerobot-eval is the simplest way to run LIBERO, Meta-World or RoboCasa365 on a LeRobot-format policy. Isaac Lab-Arena is the stronger choice if you want to author your own tasks on Isaac Sim and run thousands of environments in parallel, and SimplerEnv remains the standard for Google Robot and WidowX policies trained on real data.
Can I evaluate a robot policy on a real robot without owning one?
Yes, through RoboArena and AutoEval. RoboArena distributes double-blind pairwise evaluations of DROID policies to evaluators at partner institutions, and AutoEval runs WidowX policies on public Bridge setups with learned success detection and automatic resets. In both cases you host your policy on a server that the evaluation system queries remotely.
What is the difference between Isaac Lab-Arena and RoboLab?
Isaac Lab-Arena is a framework for composing simulation environments and evaluating policies in parallel on Isaac Lab. RoboLab is a fixed benchmark from NVIDIA Research, with 120 tasks on a simulated DROID Franka, failure-event tracking and a results dashboard, and Isaac Lab-Arena's documentation lists those tasks with Arena counterparts where they are available.
Which evaluation tools record why a robot policy failed?
RoboLab records wrong-object grasps, dropped objects and gripper collisions and timestamps them in a dashboard. Isaac Lab-Arena provides subtask predicates and sensitivity analysis, and RoboArena collects free-form evaluator explanations that are summarized into policy reports. Manifold scores atomic subgoals and uses agents to cluster failures and rank them by impact.
Do I need an RTX GPU to evaluate robot policies?
Only for tools built on Isaac Sim, which include Isaac Lab-Arena, RoboLab and RoboDojo's simulation benchmark. RoboLab's documentation recommends an NVIDIA RTX GPU with 48 GB or more of VRAM. LeRobot's LIBERO, Meta-World and RoboCasa365 environments run on MuJoCo, and SimplerEnv runs on SAPIEN, which needs an NVIDIA GPU but works on non-RTX cards, more slowly for ray-traced scenes.