BIFROST
All posts

Robot Policy Evaluation Platforms Compared (2026)

Isaac Lab-Arena, RoboLab, LeRobot and SimplerEnv evaluate robot policies in simulation, RoboArena and AutoEval on real robots. The right one depends on your benchmark.

Which Robot Policy Evaluation Platform Should You Use?

As of October 2026, no single platform covers every robot policy evaluation need, so the choice depends on whether you need simulated or real-world results, which benchmarks you report, and whether you want to run the infrastructure yourself. For simulation, the main open tools are NVIDIA's Isaac Lab-Arena (parallel evaluation across thousands of environments on Isaac Sim), RoboLab (a 120-task DROID benchmark that tracks failure events), LeRobot's lerobot-eval (LIBERO, Meta-World and RoboCasa365), SimplerEnv (real-to-sim Google Robot and WidowX tasks), XPolicyLab with RoboDojo, and RLinf. For real robots, RoboArena ranks DROID policies through double-blind comparisons at partner institutions, and AutoEval runs WidowX policies autonomously around the clock. Bifrost's Manifold is a hosted option that runs many simulation benchmarks from one harness with automated failure analysis, and it is in early access.

The table below compares them on the questions that usually decide the choice. Every cell is taken from the tool's own documentation, repository or paper as of October 2026, and cells we could not confirm are marked "Not documented" rather than guessed.

Comparison Table

ToolReal or simBenchmarks coveredGPU parallelismHosted or self-runBring your own policyBaselinesFailure analysisLicense and access
Isaac Lab-Arena (NVIDIA)Sim (Isaac Sim)RoboLab-inspired tasks (Arena counterparts of RoboLab's 120 tasks where available), Lightwheel RoboCasa and LIBERO task ports, RoboTwin 2.0, Lightwheel RoboFinals, LeRobot EnvHub environments, your own composed tasksThousands of heterogeneous environments per GPU; multi-node through OSMOSelf-runYes, behind a policy server, with integrations for GR00T and pi0.5None built inSubtask predicates, sensitivity analysisApache 2.0; alpha (v0.3), APIs will change; Isaac Sim has its own terms
RoboLab (NVIDIA Research)Sim (Isaac Lab)RoboLab-120: 120 tasks on a DROID FrankaMultiple episodes in parallel across environmentsSelf-runYes, policy runs as a server; any Isaac Lab compatible robotFive policies in the paper; pi0.5 best at 28.0%Wrong-object grasp, drop and gripper collision events; dashboard with episode replay; sensitivity analysisApache 2.0; RTX GPU with 48 GB+ VRAM recommended
LeRobot lerobot-eval (Hugging Face)SimLIBERO (130 tasks), LIBERO-plus, Meta-World (MT50), RoboCasa365, RoboTwin 2.0, RoboCerebra, RoboMME, VLABench, Isaac Lab-Arena environments via EnvHubVectorized environments on one machine (--eval.batch_size)Self-runLeRobot-format policies; new benchmarks via a gym wrapperpi0.5 LIBERO reproduction at 97.5%; SmolVLA RoboCasa checkpointSuccess rate with 95% Wilson intervals, and reward, per task and suiteApache 2.0
SimplerEnvSim (real-to-sim)Google Robot tasks (pick coke can, move near, drawers) and four WidowX Bridge tasksCPU-based ManiSkill2 (SAPIEN still needs an NVIDIA GPU); the Bridge tasks on a ManiSkill3 branch are GPU-parallel and 10 to 15x fasterSelf-runYes, by writing a policy inference scriptRT-1 checkpoints and Octo includedNone; reports sim-to-real correlation (Pearson, MMRV)MIT
RoboArenaRealDROID Franka, tasks and scenes chosen by evaluators at seven institutions in the paperNot applicableHosted evaluator network; you host the policy serverYes, any DROID-compatible policyInitial pool of seven DROID policies (pi0 and PaliGemma variants)Free-form evaluator explanations; LLM-assisted policy reports that cite episodesOpen to community submissions; code MIT
AutoEvalRealBridge WidowX; two public stations with four drawer and eggplant tasksCells run in parallel, 24/7Hosted; you host the policy server and submit through a dashboardYes, matching its observation and action specNone listed; an example OpenVLA policy server is providedSuccess rates, episode videos and framesCode MIT; the project page lists its four tasks as available until January 1, 2026, with the 2026 tasks to be announced
XPolicyLabSim and realRoboTwin 2.0, RoboDojo sim and real, RMBenchBatched policy queries across environmentsSelf-runYes, through an adapter added by pull request44 integrated policies as of September 2026Not documentedApache 2.0
RoboDojoSim and real42 sim tasks across five capability dimensions; 18 real tasks on ARX X5, Piper and Piper XHeterogeneous parallel simulation on Isaac SimSim self-run; official leaderboard and real-world runs through RoboDojo's cloud evaluation submissionYes, through XPolicyLab30 policies in the paper; public leaderboardScores by capability dimensionREADME states a non-commercial research license, while the repository's LICENSE file is MIT; confirm with the maintainers
RLinfSim and realLIBERO (with LIBERO-PRO and LIBERO-Plus), ManiSkill OOD, RoboTwin, BEHAVIOR-1K, PolaRiS, RoboCasa365, Franka real-worldParallel rollouts; multi-GPU and multi-node frameworkSelf-runOpenVLA, OpenVLA-OFT, pi0, pi0.5, GR00T and othersNot documented for evaluationSuccess rate, logs and videosApache 2.0; built mainly for RL training
Lightwheel (LW-BenchHub, RoboFinals)Sim (built on Isaac Lab-Arena)130 LIBERO-derived and 138 RoboCasa-derived tasks; RoboFinals-100, 100 household, factory and retail tasksThrough Isaac Lab-ArenaLW-BenchHub self-run; RoboFinals access by contacting LightwheelLW-BenchHub yes, through a server-client policy APINot documentedNot documentedLW-BenchHub Apache 2.0; RoboFinals license not stated
Manifold (Bifrost)SimLIBERO, LIBERO-Plus, RoboLab, RoboCasa, RoboMimic, CALVIN, SIMPLER, RoboTwin 2.0, RoboMemory, RLBench, custom tasks; native Isaac Lab-Arena supportSharded and vectorized across GPUsHostedYes, point it at a policy checkpointpi0.5 on LIBERO (published)Per-subgoal rubric scoring; agents cluster failures and rank them by impactEarly access via waitlist; not open source today

Simulation Frameworks You Run Yourself

Isaac Lab-Arena is NVIDIA's framework for building and running policy evaluations on Isaac Lab. Its main strength is scale, and the documentation reports 2,390 environment-steps per second with 1,024 parallel environments on one RTX 5880 Ada GPU in a preliminary camera-free benchmark, and environments in one batch can differ at the object level, so a single run covers many scene variations. Policies run behind a server, which keeps their dependencies out of the simulator's environment. The trade-offs are that it is labeled alpha (v0.3) with APIs that will break, it needs an RTX GPU because Isaac Sim does not support GPUs without RT cores such as the A100 and H100, and its LIBERO and RoboCasa suites are Lightwheel's adaptations on Isaac Lab rather than the original MuJoCo benchmarks, so treat their scores as a separate benchmark. We cover the throughput side in How To Evaluate Robot Policies In Simulation At Scale.

LeRobot is the lightest path when your policy is already in LeRobot format. lerobot-eval runs every benchmark through the same gym interface, reports success per task and suite, and documents recommended protocols, such as 10 episodes per task across the four standard LIBERO suites and 20 per task on RoboCasa365. Its pi0.5 reproduction reports 97.5% on LIBERO against 96.85% from openpi. Parallelism is per machine, through vectorized environments.

SimplerEnv is narrower and older, and it answers a different question, which is whether simulated scores rank real-trained policies the way real robots do. It covers Google Robot and WidowX Bridge setups only, and as the PolaRiS authors point out, it does not support wrist cameras, which is why newer DROID-focused tools exist. Our SimplerEnv alternatives post compares it with PolaRiS, RoboLab and others.

RLinf is primarily a reinforcement learning framework for VLAs, with an evaluation entry point that runs parallel rollouts in simulation or on a real Franka. It makes the most sense when evaluation is one step inside an RL post-training loop rather than a separate activity.

Fixed Benchmarks With Their Own Tooling

RoboLab is the most diagnostic of the open tools. Policies are fine-tuned only on real DROID data and tested on 120 new tasks, so the benchmark measures generalization rather than in-sim fitting, and its ranking of the four policies measured in both places matched the real-world RoboArena ordering exactly (Spearman 1.00). It records wrong-object grasps, drops and gripper collisions automatically, gives partial credit for completed subtasks, and uses neural posterior estimation to find which scene variables (such as wrist camera pose) drive failures. The paper's main evaluation used 10 episodes per task, and the authors note that this gives a confidence interval of roughly plus or minus 30% on a single task near 50% success.

XPolicyLab and RoboDojo are designed to work together, and RoboDojo uses XPolicyLab as its policy interface. XPolicyLab defines a common adapter so that N policies and M environments need N plus M integrations instead of N times M, and it ships adapters for more than 40 policies. RoboDojo adds 42 simulation tasks and 18 real-world bimanual tasks on standardized hardware that outside teams can access remotely. Check the license before using RoboDojo commercially, since its README describes a non-commercial research license even though the LICENSE file in the repository is MIT.

Real-World Evaluation Networks

RoboArena is the real-world reference that RoboLab and other simulation benchmarks compare their rankings against. Evaluators at partner institutions run double-blind pairwise comparisons on DROID robots with tasks of their choosing, and each comparison records a progress score, a preference and a written explanation. The paper reports more than 600 pairwise episodes across seven policies and finds this ranks policies more accurately than conventional centralized evaluation. It is the most trustworthy signal in this list and the slowest, since every data point costs real robot time.

AutoEval nearly removes the human from real-world evaluation on WidowX arms, using learned success classifiers and reset policies so that stations run around the clock. You host your policy as a server and submit jobs through a web dashboard, and results come back with videos. Its coverage is small (two public stations and four tasks on the project page, which lists those tasks as available until January 1, 2026 and the 2026 set as not yet announced), so check what is running before you plan around it. It suits Bridge-style policies rather than general benchmarking.

How To Choose

Your situationBest fit
One benchmark (LIBERO, Meta-World, RoboCasa365), a LeRobot policy and one GPU machineLeRobot lerobot-eval
Ranking DROID policies on real robotsRoboArena
Real-world numbers for a WidowX Bridge policy without a lab robotAutoEval, if a station and task match your policy
Google Robot or Bridge policies trained on real data, needing a sim proxySimplerEnv
A DROID policy and a sim benchmark that reports why it failedRoboLab
Authoring your own tasks on Isaac Sim and running thousands of environmentsIsaac Lab-Arena
Bimanual policies, or the same adapter for sim and realXPolicyLab with RoboDojo (check the RoboDojo license for commercial use)
Evaluation inside an RL post-training loopRLinf
Many benchmarks and simulators, many checkpoints, failure analysis, no infrastructure to runManifold

Several of these combine well. A common pattern is a fast simulated loop on every checkpoint, a perturbation suite to catch memorization, and real-world evaluation on the few candidates that survive. Simulated scores still need checking against real robots, and Best Simulation Benchmarks For VLA Models In 2026 covers which benchmarks have published real-world correlations.

Where Manifold Fits

Manifold is Bifrost's hosted platform for evaluating robot policies in simulation. One harness runs a policy across benchmarks including LIBERO, LIBERO-Plus, RoboLab, RoboCasa, CALVIN, SIMPLER and RoboTwin 2.0, on simulators including NVIDIA Isaac Sim, MuJoCo, ManiSkill and Genesis, with native support for Isaac Lab-Arena. Rollouts are sharded and vectorized across GPUs, so a LIBERO sweep that takes 8 hours comes back in about 30 minutes (roughly 16x), and there is no simulator, cloud or GPU setup on your side. Every task is broken into atomic subgoals scored against rubrics you can define, agents cluster failures and rank them by impact, and each run records the policy and benchmark container images it used, pinned by digest. Published reference results so far cover pi0.5 on LIBERO (pi0.5 LIBERO results on Manifold).

Manifold does not replace real-world evaluation, and it is not open source today. If you run one benchmark on a machine you control, LeRobot or Isaac Lab-Arena will do the job, and RoboArena is the better tool for a real-world ranking. Manifold is in early access, and you can join the waitlist from the Manifold page or read more on the robotics page.

Sources

Frequently Asked Questions

What is the best open-source robot policy evaluation framework?

It depends on the benchmark. LeRobot's lerobot-eval is the simplest way to run LIBERO, Meta-World or RoboCasa365 on a LeRobot-format policy. Isaac Lab-Arena is the stronger choice if you want to author your own tasks on Isaac Sim and run thousands of environments in parallel, and SimplerEnv remains the standard for Google Robot and WidowX policies trained on real data.

Can I evaluate a robot policy on a real robot without owning one?

Yes, through RoboArena and AutoEval. RoboArena distributes double-blind pairwise evaluations of DROID policies to evaluators at partner institutions, and AutoEval runs WidowX policies on public Bridge setups with learned success detection and automatic resets. In both cases you host your policy on a server that the evaluation system queries remotely.

What is the difference between Isaac Lab-Arena and RoboLab?

Isaac Lab-Arena is a framework for composing simulation environments and evaluating policies in parallel on Isaac Lab. RoboLab is a fixed benchmark from NVIDIA Research, with 120 tasks on a simulated DROID Franka, failure-event tracking and a results dashboard, and Isaac Lab-Arena's documentation lists those tasks with Arena counterparts where they are available.

Which evaluation tools record why a robot policy failed?

RoboLab records wrong-object grasps, dropped objects and gripper collisions and timestamps them in a dashboard. Isaac Lab-Arena provides subtask predicates and sensitivity analysis, and RoboArena collects free-form evaluator explanations that are summarized into policy reports. Manifold scores atomic subgoals and uses agents to cluster failures and rank them by impact.

Do I need an RTX GPU to evaluate robot policies?

Only for tools built on Isaac Sim, which include Isaac Lab-Arena, RoboLab and RoboDojo's simulation benchmark. RoboLab's documentation recommends an NVIDIA RTX GPU with 48 GB or more of VRAM. LeRobot's LIBERO, Meta-World and RoboCasa365 environments run on MuJoCo, and SimplerEnv runs on SAPIEN, which needs an NVIDIA GPU but works on non-RTX cards, more slowly for ray-traced scenes.

Get access More Posts