SimplerEnv Alternatives For Real-To-Sim Policy Evaluation
The main SimplerEnv alternatives are PolaRiS for DROID policies, RoboLab in Isaac Lab-Arena, Gaussian-splat soft-body twins, RobotArena ∞, and real-world RoboArena.
The Short Answer
SimplerEnv is still the cheapest way to score Google Robot and WidowX BridgeData V2 policies in simulation, but it covers eight task groups, no wrist cameras and only those two robots. As of October 2026, the strongest alternative for DROID-style Franka policies is PolaRiS, which builds Gaussian-splat scenes from short phone videos and reports an average Pearson correlation of 0.9 against real-world evaluations across six paired environments. RoboLab-120 in NVIDIA's Isaac Lab-Arena gives a fixed 120-task DROID suite whose policy ranking matched the real-world RoboArena ranking. Gaussian-splat soft-body twins cover deformable tasks, RobotArena ∞ scales scene creation from existing robot videos, and RoboArena itself is the real-world reference most of these validate against.
LIBERO-PRO, LIBERO-Plus and ManiSkill3 are useful alongside SimplerEnv, but they answer different questions, and the table below says which.
What SimplerEnv Is
SimplerEnv is the code release for SIMPLER, Evaluating Real-World Robot Manipulation Policies in Simulation (Li et al., CoRL 2024). Its goal is to evaluate policies trained on real robot data, such as RT-1, RT-1-X, RT-2-X and Octo, in simulation without fine-tuning them on simulated data, and to show that the simulated ranking matches the real one.
It recreates two widely used setups. The Google Robot environments cover four task groups (pick Coke can, move near, open or close drawer, and open drawer and place apple). The WidowX environments recreate four BridgeData V2 tasks (spoon on towel, carrot on plate, stack green block on yellow block, eggplant in basket). It is built on the SAPIEN simulator and the CPU-based ManiSkill2 benchmark, and a ManiSkill3 version with GPU parallelization runs 10 to 15 times faster.
SimplerEnv offers two ways of closing the visual gap. Visual matching overlays real-world images onto the simulated background and adjusts robot and object textures so the simulated camera view resembles the real one. Variant aggregation builds simulated variants with different backgrounds, lighting, distractors, table textures and camera poses, then averages the results. The authors also tuned the simulated controller with system identification to close the control gap.
The validation is what made SimplerEnv the default. Across six Google Robot checkpoints (three RT-1 checkpoints, RT-1-X, RT-2-X and Octo-Base), visual matching reached an average Pearson r of 0.924 and a Mean Maximum Rank Violation (MMRV, lower is better) of 0.056, compared with 0.778 and 0.143 for variant aggregation and 0.308 and 0.375 for ranking by validation action MSE. On Bridge tasks with RT-1-X, Octo-Base and Octo-Small, visual matching reached r = 0.890 and MMRV = 0.014.
Why Teams Look For Alternatives
SimplerEnv's limits come from its scope.
- Two robots. If your policy runs on a Franka, a bimanual rig or a humanoid, there is no SimplerEnv environment for it.
- No wrist cameras. The PolaRiS paper points out that SIMPLER and earlier real-to-sim methods do not support wrist cameras, which rules out most current DROID-trained VLAs.
- Hand-built scenes. Each environment was matched by hand, so adding a new scene or task takes real effort, and the task set has stayed small.
- 2024-era validation. The correlation numbers were measured on RT-1, RT-2-X and Octo checkpoints. They are strong evidence for those policies, but the paper predates pi0-class models and did not test them.
- Rigid bodies only. None of the tasks involve cloth, rope or soft objects.
SimplerEnv Alternatives Compared
| Option | What it is | Embodiment | Engine / rendering | Published sim-real correlation | Best for |
|---|---|---|---|---|---|
| SimplerEnv (baseline) | Hand-matched real-to-sim replicas | Google Robot, WidowX 250 | SAPIEN (ManiSkill2, ManiSkill3 port) | Yes. Avg r = 0.924 (Google Robot), 0.890 (Bridge) | Policies trained on Google Robot or Bridge data, third-person camera |
| PolaRiS | Scenes reconstructed from 2 to 5 minute videos with 2D Gaussian splatting, plus a short sim co-finetune | DROID Franka, wrist and external cameras | Isaac Sim physics, Gaussian-splat rendering | Yes. Avg r = 0.9 over 6 paired environments (worst 0.81); r = 0.98 vs RoboArena scores | DROID policies, evaluating in your own scenes |
| Gaussian-splat soft-body twins (Zhang et al.) | PhysTwin spring-mass objects rendered with 3D Gaussian splatting | UFactory xArm 7 | PhysTwin dynamics, 3DGS rendering | Yes. r = 0.944 (toy packing), 0.901 (rope routing), 0.915 (T-block pushing) | Deformable objects |
| RobotArena ∞ | Sim scenes generated automatically from Bridge, DROID and RH20T videos, scored by VLMs and crowdsourced pairwise preferences | Robots from the Bridge, DROID and RH20T videos | Physics sim; controller gains identified in Genesis | Limited. One real task, three policies, outcomes matched | Large-scale ranking of many open policies |
| RoboLab-120 on Isaac Lab-Arena | 120 tasks tagged by visual, procedural and relational competency | DROID Franka | Isaac Lab on Isaac Sim | Ranking level. Spearman 1.00 vs RoboArena Elo across 5 policies | A fixed DROID suite you can rerun and compare |
| LIBERO-PRO / LIBERO-Plus | Perturbed versions of LIBERO tasks | Franka Panda | robosuite on MuJoCo | None published | Testing memorization and robustness in sim |
| ManiSkill3 Bridge digital twins | GPU port of SimplerEnv's four Bridge tasks | WidowX 250 | SAPIEN, GPU-parallel | No separate study; based on SimplerEnv's Bridge environments | Fast Bridge evaluation at scale |
| RoboArena | Distributed, double-blind pairwise evaluation on real robots | DROID Franka | Real world | Not applicable. It is the real-world reference | Final rankings of generalist DROID policies |
PolaRiS
PolaRiS (Jain et al., 2025) is the closest thing to a successor for DROID policies. You record a 2 to 5 minute monocular video of a real scene, reconstruct it with 2D Gaussian splatting, extract a mesh for collisions and import it into Isaac Sim, with the splats providing the rendering. Robot splats are anchored to the robot's links, which is what makes wrist cameras work. The paper reports that environment creation typically completes within 30 minutes on an RTX 4090.
Reconstruction alone did not give reliable correlation, so PolaRiS adds a co-training step. Policies are fine-tuned for 1,000 steps on about 350 simulated demonstrations collected in 15 separate co-training scenes, and then evaluated zero-shot in unseen PolaRiS environments with 50 rollouts per task. Across six paired real and simulated environments, with 20 real rollouts per policy per environment, PolaRiS reached an average Pearson r of 0.9 for pi0, pi0-FAST, pi0.5 and a PaliGemma binning policy, and scores also correlated with RoboArena results at r = 0.98.
PolaRiS has two limits. The fine-tuning step means you are evaluating a lightly co-trained copy of your policy, not the exact checkpoint. And the authors say the tasks cover rigid-body tabletop manipulation, a small subset of what RoboArena tests, so they present PolaRiS as a fast estimate during development rather than a replacement for real-world testing. The code and a Hugging Face hub of shared environments are open.
Gaussian-Splat Soft-Body Twins
If your tasks involve deformable objects, SimplerEnv and PolaRiS do not help. Zhang et al. (2025) build digital twins from interaction videos using PhysTwin, which models deformable objects as dense spring-mass systems, and render them with 3D Gaussian splatting. They evaluated ACT, Diffusion Policy, SmolVLA and pi0 on plush toy packing, rope routing and T-block pushing and report correlations of r = 0.944, 0.901 and 0.915. An Isaac Lab baseline on the same rope and push-T tasks reached only 0.237 and 0.649. The scope is three tasks on one xArm 7 setup, so treat it as a strong method for building your own soft-body evaluation, not a ready benchmark.
RobotArena ∞
RobotArena ∞ (Carnegie Mellon and National Taiwan University) automates the step SimplerEnv does by hand. It converts demonstration videos from BridgeData V2, DROID and RH20T into simulated scenes using VLMs, 2D-to-3D generation and differentiable rendering, then scores policies with VLM-judged task progress and crowdsourced pairwise preferences (8,749 comparisons in its BridgeSim environments). It also perturbs textures and object placements to test robustness, and it evaluated Octo, RoboVLM, SpatialVLA, CogACT, X-VLA and an open pi0 reimplementation. The authors say plainly that full real-world correlation would be very hard to measure; they tested one task, put the carrot on the plate, and real and simulated outcomes matched for all three policies tried. It is a good tool for broad comparisons across many open policies, with weaker real-world evidence than SimplerEnv or PolaRiS.
RoboLab And Isaac Lab-Arena
RoboLab, from NVIDIA Research, is a fixed benchmark rather than a scene-capture pipeline. RoboLab-120 has 120 tasks on a simulated DROID Franka whose cameras match the real robot, and policies are fine-tuned only on real DROID data. The best of five policies, pi0.5, reached 28.0% success, so there is plenty of headroom. RoboLab's per-policy success rates preserved RoboArena's Elo ordering exactly (Spearman 1.00). That is a benchmark-level result on five policies, not a per-scene correlation like PolaRiS reports. RoboLab ships as a task catalog in Isaac Lab-Arena, NVIDIA's open-source framework for authoring benchmarks and running them in parallel, which is still labeled alpha (v0.3). Both need an RTX GPU, since Isaac Sim does not support GPUs without RT cores such as the A100 and H100.
LIBERO-PRO, LIBERO-Plus And ManiSkill3
These are often mentioned in the same breath as SimplerEnv but do a different job. LIBERO-PRO and LIBERO-Plus perturb LIBERO's objects, instructions, cameras, lighting and initial states, and they are excellent at exposing memorization (LIBERO-PRO reports models above 90% falling to 0.0%). They do not claim to predict real-world performance, and the PolaRiS authors found plain LIBERO scores poorly correlated with real results. ManiSkill3 ships SimplerEnv's four Bridge tasks as GPU digital twins that its docs say evaluate up to 60 times faster than the real world and 10 times faster than CPU simulation, so it is the fastest way to run the Bridge half of SimplerEnv rather than a replacement for it.
RoboArena
RoboArena is not a simulator. It is a network of evaluators at seven academic institutions running double-blind pairwise comparisons of DROID policies on real robots, on tasks and scenes each evaluator chooses. The paper reports more than 600 pairwise episodes across seven generalist policies and finds the crowd-sourced ranking more accurate than conventional centralized evaluation. It is the most trustworthy signal on this list and the slowest, since every data point costs real robot time. It is also the reference that PolaRiS and RoboLab validate against.
How To Choose
- Your policy was trained on Google Robot or Bridge data with a third-person camera. Keep SimplerEnv. It is cheap, validated and comparable to prior papers. Use the ManiSkill3 port for speed.
- Your policy is a DROID Franka VLA with a wrist camera. Use RoboLab-120 for a fixed, comparable suite and PolaRiS to build environments that match your own lab.
- Your tasks involve cloth, rope or soft objects. Build on the Gaussian-splat soft-body approach.
- You want to rank many open policies across datasets. RobotArena ∞, keeping in mind its limited real-world validation.
- You need to know whether the policy memorized its training scenes. Add LIBERO-PRO or LIBERO-Plus, but do not treat them as real-world proxies.
- You are about to make a final claim about a DROID policy. Submit to RoboArena.
Whatever you pick, report the rollout count and an interval. The trial math for telling two policies apart applies to simulated evaluations as much as real ones, and our guide to VLA evaluation covers the rest of the protocol. For simulation benchmarks beyond real-to-sim, see Best Simulation Benchmarks For VLA Models In 2026.
Where Manifold Fits
Manifold is Bifrost's evaluation platform for running robot policies across simulation benchmarks from one harness. Its supported benchmarks include SIMPLER, RoboLab, LIBERO, LIBERO-Plus, RoboCasa, RoboTwin 2.0 and CALVIN, it runs on simulators including NVIDIA Isaac Sim, MuJoCo and ManiSkill, and it has native support for Isaac Lab-Arena. Rollouts are sharded across GPUs, each run gets a manifold:// URI that pins the checkpoint, simulator version, benchmark version and seeds, and failures are clustered by mode rather than reported as one number. Manifold does not replace real-world evaluation such as RoboArena. It makes the simulated half of the loop faster and reproducible. It is in early access, and you can join the waitlist from the Manifold page.
Sources
- Li et al., Evaluating Real-World Robot Manipulation Policies in Simulation, CoRL 2024, and the SimplerEnv repository
- Jain et al., PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies, 2025
- Zhang et al., Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions, 2025
- Jangir et al., RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation, 2025
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies, NVIDIA, 2026
- Isaac Lab-Arena, NVIDIA
- Atreya, Pertsch, Lee et al., RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies, 2025
- LIBERO-PRO and LIBERO-Plus, 2025
- Tao et al., ManiSkill3, 2024, and the ManiSkill digital twin docs
Frequently Asked Questions
Does SimplerEnv support wrist cameras?
No. SimplerEnv recreates the third-person camera setups of the Google Robot and WidowX BridgeData V2 rigs, and the PolaRiS authors note that SIMPLER and earlier real-to-sim approaches do not support policies that use wrist cameras. That rules out most current DROID-trained VLAs, which is the main reason PolaRiS and RoboLab exist.
What is the difference between visual matching and variant aggregation in SimplerEnv?
Visual matching overlays real-world images onto the simulated background and adjusts the textures of the robot and foreground objects so the simulated view looks like the real one. Variant aggregation instead builds many simulated variants with different backgrounds, lighting, distractors, table textures and camera poses and averages the results. In the SIMPLER paper, visual matching correlated better with real-world results, with an average Pearson r of 0.924 versus 0.778 on Google Robot tasks.
Is SimplerEnv built on ManiSkill2 or ManiSkill3?
The original SimplerEnv is built on the SAPIEN simulator and the CPU-based ManiSkill2 benchmark. A ManiSkill3 version with GPU parallelization runs 10 to 15 times faster according to the repository, and ManiSkill3 itself ships the four WidowX Bridge tasks as GPU digital twins based on SimplerEnv.
Can I use LIBERO instead of SimplerEnv to predict real-world performance?
Not reliably. LIBERO measures in-simulation performance after fine-tuning on simulated demonstrations, and the PolaRiS authors found that nearly every DROID policy they tested scored between 90% and 95% on LIBERO-90 while ranging from strong to weak in real-world tests. LIBERO-PRO and LIBERO-Plus are good robustness tests, but neither publishes a correlation with real robots.
Can a video world model replace a simulator for policy evaluation?
Not yet for diverse scenes. In the PolaRiS comparison, rolling policies out in the open-source Ctrl-World video model produced hallucinations during object interaction and clear policy mis-rankings, with an MMRV of 0.22. The authors note that video models have shown better results in narrower scene settings.