Can You Trust Simulated Robot Policy Evaluation?
For ranking policies, often yes. SIMPLER reports a 0.924 Pearson correlation with real results and PolaRiS 0.90, but sim success rates rarely match real ones.
The Short Answer
You can trust simulated evaluation for ranking policies when the simulator has been checked against paired real-world trials on similar tasks, and much less for predicting a policy's absolute real-world success rate. Validated real-to-sim setups report strong agreement: SIMPLER's visual-matching environments reached a Pearson correlation of 0.924 and an MMRV of 0.056 against real Google Robot evaluations of six policies, and PolaRiS reached an average Pearson correlation of 0.90 across six real DROID scenes. Unvalidated setups can fail badly, though. A June 2026 study of five recent VLAs measured an average Pearson correlation of 0.402 for SIMPLER and 0.725 for VLA-Arena on its own aligned tasks, and PolaRiS found that LIBERO scores barely separated policies that differed widely on real hardware.
The useful question is whether a given simulator, on your tasks and with your policies, has been shown to preserve the real-world ordering. The rest of this post explains the two metrics the field uses to answer that, collects the published numbers in one table, and ends with a checklist.
How Sim-To-Real Correlation Is Measured
The setup is the same in every paper. Take N policies, evaluate each one on real hardware and in the simulator on matched tasks, and compare the two lists of success rates. Two metrics dominate.
Pearson Correlation
Pearson correlation (r) measures whether simulated success rates rise and fall linearly with real ones across policies. It ranges from -1 to 1, and 1 means the points fall on a straight line. It does not require the simulator to match absolute numbers. A simulator that scores every policy 20 points lower than reality can still have r close to 1.
Pearson has two known weaknesses for this job, which the SIMPLER paper illustrates in its Figure 2. It penalizes a simulator that ranks policies correctly but not linearly, and it is oversensitive to small noise when policies perform similarly in the real world. With only 3 to 6 policies, which is typical, a single noisy point can move r a long way.
Mean Maximum Rank Violation (MMRV)
SIMPLER introduced MMRV to measure ranking mistakes directly, weighted by how much they matter. Let R_i be policy i's real success rate and S_i its simulated success rate. A rank violation between policies i and j is the real gap |R_i − R_j| if the simulator orders the pair differently from reality, and zero otherwise. MMRV takes the worst violation for each policy and averages over policies:
RankViolation(i, j) = |R_i - R_j| × 1[ (S_i < S_j) != (R_i < R_j) ] MMRV = (1/N) × Σ_i max_j RankViolation(i, j)
MMRV ranges from 0 to 1, and 0 means the simulator orders every pair correctly. Swapping two policies that are 2 points apart in reality costs little. Swapping two that are 40 points apart costs a lot.
A Worked Example
Four policies score 80%, 70%, 45% and 20% on the real robot. The simulator gives them 55%, 60%, 30% and 5%. Every simulated number is 10 to 25 points lower than reality, and the top two are swapped.
- Pearson r = 0.973, because the sim scores track the real ones closely apart from the offset.
- Spearman rank correlation = 0.80, because one adjacent pair is out of order.
- MMRV = 0.050. The swapped pair is 10 points apart in reality, so each of those two policies carries a worst violation of 0.10, the other two carry 0, and the average is 0.20 / 4.
All three numbers look good, but the simulator would have picked the wrong best policy. That is the case MMRV is meant to surface, and it is why a 0.05 MMRV over four policies should be read as "one close pair might be swapped," not "rankings are right." To check our implementation, we recomputed SIMPLER's Table XIII (seven policies on Pick Coke Can) and got its reported r of 0.959 and MMRV of 0.027.
Published Sim-Versus-Real Results
The table collects the headline correlations from the main real-to-sim evaluation papers. Higher r is better, lower MMRV is better. Each study uses different tasks, policies and real-trial counts, so compare rows within a study rather than across studies.
| Study | Setup | Policies | Pearson r | MMRV |
|---|---|---|---|---|
| SIMPLER, Visual Matching | Google Robot, 3 tasks, SAPIEN | 6 (three RT-1 checkpoints, RT-1-X, RT-2-X, Octo-Base) | 0.924 | 0.056 |
| SIMPLER, Variant Aggregation | Same tasks, averaged over visual variants | 6 | 0.778 | 0.143 |
| Validation action MSE (no simulator), from SIMPLER | Same tasks | 6 | 0.308 | 0.375 |
| SIMPLER, Visual Matching | WidowX + BridgeData V2, 4 tasks | 3 (RT-1-X, Octo-Base, Octo-Small) | 0.575 to 1.000 per task | 0.000 on three tasks, 0.111 on Put Carrot on Plate |
| PolaRiS | DROID Franka, 6 Gaussian-splat scenes scanned from real environments, policies co-trained on sim data | 4 (π0, π0-FAST, PaliGemma-binning, π0.5) | 0.90 average, 0.81 worst scene | Not reported as a single number |
| PolaRiS against RoboArena scores | Same policies, real-world RoboArena progress scores | 4 | 0.98 | Not reported |
| Ctrl-World video model, from PolaRiS | Policies rolled out in a world model, human-scored | 4 | Not reported | 0.22 |
| Gaussian-splat soft-body sim (Zhang et al.) | Toy packing, rope routing, T-block pushing | Checkpoints of ACT, Diffusion Policy, SmolVLA, π0 | 0.944, 0.901, 0.915 | 0.076, 0.174, 0.108 |
| Isaac Lab baseline, from Zhang et al. | Rope routing, T-block pushing | Same | 0.237, 0.649 | 0.270, 0.196 |
| Wang et al., REALM | 9 aligned tabletop tasks, 4 perturbation types, Isaac Sim | 5 (π0, π0-FAST, π0.5, GR00T N1.6, GR00T N1.7) | 0.785 average | 0.030 |
| Wang et al., VLA-Arena | MuJoCo, same alignment protocol | 5 | 0.725 average | 0.060 |
| Wang et al., SIMPLER | SAPIEN, 3 perturbation types | 5 | 0.402 average | 0.128 |
| Wang et al., REALM after sim post-training | Small amount of simulator data per task | 5 | 0.878 average | 0.015 |
Sources for each row are linked in the first column. SIMPLER figures are from its Table I and Table V. PolaRiS figures are from its Section 5.2. The soft-body figures are from Table I of Zhang et al. The Wang et al. figures are from its Tables 2 and 4, averaged over perturbation dimensions.
Two of these rows describe the same benchmark under different conditions. SIMPLER scores strongly in its own paper (six RT-1, RT-2-X and Octo policies on Google Robot tasks) and weakly in Wang et al. (five 2025 to 2026 VLAs, perturbed tasks aligned to a different real setup). Both results can be right at once, because a simulator's correlation is a property of the simulator, the tasks and the policy set together, and it does not carry over automatically when any of them changes.
When Sim Rankings Transfer
Visual inputs are matched to the real scene. SIMPLER's visual-matching approach overlays real background images and tunes object and robot textures. On the drawer tasks, it scored r = 0.942 against 0.486 for variant aggregation, which averages over generic visual variations instead. PolaRiS's ablations point the same way: Gaussian-splat rendering of scanned real scenes gave stronger correlation than ray-traced or textureless versions of the same scenes.
Physics and control are close enough for the task. For rigid-object pick-and-place, SIMPLER found its correlation stayed high across a range of plausible object masses and friction settings, with Pearson r between 0.957 and 0.990 on Pick Coke Can. Deformable objects are harder. In Zhang et al., an Isaac Lab baseline reached only r = 0.237 on rope routing, while their simulator with physics fitted from real video reached 0.901.
Policies differ by a lot in the real world. Large real gaps are easy to preserve. Rankings among policies within a few points of each other are fragile in any simulator, and they are often within the noise of the real-world evaluation as well.
Policies have seen a little simulator data. PolaRiS found that pretrained checkpoints evaluated without simulator co-training did not correlate well enough to rank accurately, and that co-training on scenes completely separate from the evaluation scenes worked about as well as in-domain data. Wang et al. found a similar gain from light post-training in REALM.
When They Do Not
Absolute success rates. Even SIMPLER's best setup does not reproduce absolute numbers. RT-1-X averaged 0.760 on real Pick Coke Can and 0.567 in simulation, and Octo-Base 0.293 against 0.170. Use simulation to compare policies and real trials to estimate deployment success.
Saturated benchmarks. PolaRiS evaluated its policies on all 90 LIBERO-90 tasks at 50 rollouts per task, 4,500 rollouts per policy, and found nearly all of them scored between 90% and 95% while spanning the full range on real hardware. A benchmark where every policy scores near the ceiling cannot rank them.
Observations the simulator does not render. PolaRiS notes that earlier real-to-sim approaches, including SIMPLER, do not support wrist cameras, which most current VLAs use.
Generative world models as simulators, for now. In PolaRiS's tests, the open-source Ctrl-World video model hallucinated during object interaction and misranked policies (MMRV 0.22). Video models have shown better results in narrower scenes, but they are not yet a general substitute.
Too much simulator fine-tuning. Wang et al. found the relationship non-monotonic. Ten demonstrations per task gave the best alignment in REALM, while twenty kept ranking consistency reasonable but made perturbation-sensitivity alignment worse than no fine-tuning at all. PolaRiS likewise reports that fine-tuning for too long reduces correlation.
Perturbation types the simulator handles poorly. In Wang et al., SIMPLER's language-perturbation correlation was r = 0.086 and VLA-Arena's vision-perturbation correlation was r = 0.241, even where other dimensions scored well. An average can hide a dimension that does not transfer.
The Validation Itself Needs Enough Real Trials
Every correlation in the table rests on a small number of real-world trials. PolaRiS ran 20 real rollouts per policy per environment, Zhang et al. ran 16 to 27 episodes per checkpoint, and Wang et al. ran 5 real rollouts per policy, task and perturbation dimension (1,115 real rollouts in total). At 20 trials, a single real success rate near 80% has a 95% Wilson interval of about ±17 points, so the "ground truth" side of the comparison is noisy too. Our post on how many trials you need to evaluate a robot policy has the lookup tables.
Two ways to deal with this:
- Validate on well-separated policies first, then trust the simulator only for comparisons of similar or larger gaps.
- Combine sim and real statistically. SureSim (Badithela et al.) treats the problem as prediction-powered inference, using a small number of paired real and simulated evaluations to correct bias in large-scale simulation and produce confidence intervals on real-world performance. The authors report saving over 20 to 25% of hardware evaluation effort for similar bounds.
A Practical Checklist
- Pick validation policies that span a wide performance range in the real world, and include at least one weak and one strong policy. Four to six is the range most papers use.
- Run paired evaluations on matched tasks: same task definition, same success criterion, same camera views (including wrist cameras if your policy uses them), and initial states as close as you can make them.
- Report Pearson r, MMRV and a rank correlation together, per task and averaged. Report the real and sim success rates themselves, not only the correlations.
- Check each task and perturbation type separately. One dimension with near-zero correlation can hide inside a good average.
- Match visuals to the deployment scene where you can, through background overlays, texture tuning or scanned Gaussian-splat scenes.
- Fit physics where contact matters. Rigid pick-and-place tolerates rough parameters. Deformables and articulated objects need parameters fitted to real behavior.
- Keep simulator fine-tuning light and measure correlation at more than one fine-tuning budget.
- Re-validate when the policy family changes. A simulator validated on one generation of policies is not automatically valid for the next.
- Use simulation to choose and real trials to confirm. Shortlist checkpoints in simulation, then spend the hardware budget on the final comparison and on absolute success estimates.
Where Manifold Fits
Manifold is Bifrost's robot policy evaluation platform. It runs policies across simulator benchmarks, including SIMPLER, LIBERO and RoboCasa, with thousands of rollouts sharded across GPUs, and agents cluster the failures across rollouts. That makes the simulation side of a validation study cheaper to run at the rollout counts these comparisons need. It does not replace the real-world trials, which you still need to establish that a given simulator tracks your hardware. Manifold is in early access through the waitlist. For the wider evaluation workflow, see how to evaluate a VLA policy.
Sources
- Li et al., Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER), 2024, and simpler-env.github.io.
- Jain et al., PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies, 2025, and polaris-evals.github.io.
- Zhang et al., Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions, 2025.
- Wang et al., A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation, June 2026.
- Badithela et al., Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators, 2025.
Frequently Asked Questions
What is MMRV in robot policy evaluation?
MMRV, or Mean Maximum Rank Violation, is a metric introduced by the SIMPLER paper to measure how badly a simulator misorders policies relative to real-world results. For each policy it finds the worst pair that the simulator ranks in the wrong order, weights that mistake by the real-world success gap between the two policies, and averages over policies. Zero means no ranking mistakes, and lower is better.
Is Pearson correlation enough to validate a simulator for policy evaluation?
No. Pearson correlation measures whether simulated scores rise and fall linearly with real scores, so it can stay high while two close policies swap places, and it can be dragged down by a simulator that ranks correctly but compresses scores. It also becomes unstable when policies have similar real success rates. Report it alongside a rank metric such as MMRV or Spearman correlation.
Does a high LIBERO score predict real-world performance?
Not reliably. In the PolaRiS study, nearly all tested VLAs scored between 90% and 95% on LIBERO-90 while spanning the full range from high to low performance on real hardware, and the LIBERO score correlated poorly with real results. LIBERO is useful for studying generalization in simulation, but it was not built as a real-world proxy.
How many real-world trials do I need to validate a simulator?
Enough that the real success rates you compare against are not themselves noise. Published validations use 5 to 27 real trials per policy per condition, which leaves 95% intervals of roughly ±11 to ±20 points per number at 20 trials. Validate with policies whose real performance is well separated, and use more trials where policies are close.
Does fine-tuning on simulator data make simulated evaluation more trustworthy?
Light fine-tuning on simulator data helped in two 2025 to 2026 studies. PolaRiS found that policies without simulator co-training were not ranked accurately, and a June 2026 study raised REALM's Spearman correlation from 0.700 to 0.875 after targeted post-training. The same study found that too much simulator fine-tuning made perturbation sensitivity worse, so keep it small.