Is LIBERO Saturated? What Top Scores Hide And What To Report
Largely, yes. Top VLAs report 96 to 98% averages, but LIBERO-PRO and LIBERO-Plus perturbations drop several to near 0% when objects move, tasks change or cameras shift.
Is LIBERO Saturated?
For the standard protocol, largely yes. As of October 2026, the strongest open VLA models report LIBERO averages between 96.85% and 98.1% across the four standard suites, which leaves less than 4 points of headroom and puts the gaps between top models inside the confidence interval of a 500-episode suite. The perturbation benchmarks show what that ceiling hides. On LIBERO-PRO, OpenVLA, pi0 and pi0.5 score 0% to 1% when the task is modified, and OpenVLA and pi0 score 0% when object positions move, and on LIBERO-Plus OpenVLA-OFT drops from 97.1% to 37.2% when the robot's initial state changes. LIBERO is still a useful regression test, but it no longer separates the best policies, and it says little about generalization on its own.
The rest of this post shows the ceiling with sourced numbers, summarizes what the two perturbation benchmarks found, lists what still discriminates between policies, and ends with a short recommendation for what to report next to a LIBERO score.
The Ceiling On Standard LIBERO
These are the four-suite averages reported by each source, with LIBERO-10 shown separately because it is where most of the remaining spread sits.
| Model | LIBERO-10 (Long) | Four-suite average | Source |
|---|---|---|---|
| X-VLA (0.9B) | 97.6 | 98.1 | X-VLA paper, Table 13 |
| pi0.5 (LeRobot reproduction, 10 episodes per task) | 96.0 | 97.5 | LeRobot LIBERO docs |
| RIPT-VLA | n/a | 97.5 | As listed in LIBERO-Plus, Table 1 |
| OpenVLA-OFT | 94.5 | 97.1 | OpenVLA-OFT paper, Table I |
| pi0.5 (openpi, 30k steps) | 92.4 | 96.85 | openpi LIBERO README |
| pi0 (fine-tuned) | 85.2 | 94.2 | Compiled in OpenVLA-OFT, Table I |
| SmolVLA (0.45B) | 71 | 87.3 | SmolVLA paper, Table 2 |
| OpenVLA (fine-tuned) | 53.7 | 76.5 | OpenVLA README |
OpenVLA's 76.5% average was added to its paper in September 2024, and by early 2025 OpenVLA-OFT reported 97.1%. Two facts about the protocol explain why the top rows bunch up.
The first is sample size. A standard suite is 10 tasks with 50 episodes each, so 500 episodes. At a true success rate of 97%, the 95% confidence interval on a 500-episode estimate is about plus or minus 1.5 points (normal approximation), and at LeRobot's recommended 10 episodes per task it is wider still. The difference between 96.85% and 98.1% is smaller than that margin, which means the ordering among the top five rows above is not established by the published numbers.
The second is what the test measures. The LIBERO-PRO authors point out that LIBERO's evaluation tasks are the same tasks used for training, and that the only difference at test time is a small change in the initial placement of objects. Each task ships a fixed set of initial states, and the standard scripts evaluate the same 50 of them every time. A policy can score in the high 90s by learning a reliable mapping from these scenes to these trajectories. LIBERO was designed to study knowledge transfer in lifelong learning (Liu et al., 2023), and the high-90s scores come from fine-tuning a large pretrained model on each suite and testing on the same suite, which is a different use from the one the benchmark was built for.
What LIBERO-PRO Does To Top Models
LIBERO-PRO (Zhou et al., 2025) keeps the LIBERO scenes and perturbs five things: the manipulated objects (appearance, size, color), their initial positions, the instruction wording, the task itself (new goals built from the same scene), and the environment. It evaluates 50 episodes per task, matching the original protocol. The table below gives the range across the four standard suites, from the original-task results and the leaderboard in the project README.
| Condition | OpenVLA | pi0 | pi0.5 |
|---|---|---|---|
| Original LIBERO tasks | 0.93 to 0.99 | 0.82 to 0.98 | 0.93 to 0.98 |
| Object appearance changed | 0.81 to 0.98 | 0.79 to 0.95 | 0.92 to 0.98 |
| Instruction paraphrased | 0.96 to 0.98 | 0.82 to 0.97 | 0.93 to 0.97 |
| Object positions changed | 0.00 on all four | 0.00 on all four | 0.08 to 0.38 |
| Task modified | 0.00 on all four | 0.00 on all four | 0.00 to 0.01 |
| Environment swapped | 0.00 to 0.98 | 0.27 to 0.60 | 0.46 to 0.73 |
The authors read the stable rows with caution. They report that models keep grasping when the target object is replaced with an irrelevant item, and that outputs stay unchanged when the instruction is corrupted or replaced with nonsense tokens, so the high scores under paraphrased instructions may mean the policy ignores the instruction rather than understands it. The one position result above zero, pi0.5 at 0.38 on LIBERO-Goal, is the kind of difference between models that the standard benchmark hides, since all three score above 0.9 on the original tasks.
LIBERO-PRO's task and position perturbations are large by design, and a 0% score under a modified task partly reflects that the policy was never trained on that task. The finding that holds up is that the same models cannot recover when objects start somewhere new in a scene they have seen thousands of times.
What LIBERO-Plus Does To Top Models
LIBERO-Plus (Fei et al., 2025) takes a finer-grained approach. It perturbs seven factors (objects layout, camera viewpoint, robot initial state, language instruction, lighting, background texture, and sensor noise) across 21 sub-dimensions, and grades each variant into five difficulty levels. The released benchmark has 10,030 tasks. Table 1 of the paper reports success rates (%) for each perturbation type:
| Model | Original | Camera | Robot init | Language | Light | Background | Noise | Layout |
|---|---|---|---|---|---|---|---|---|
| OpenVLA | 76.5 | 1.1 | 4.1 | 26.8 | 4.4 | 25.3 | 19.3 | 31.6 |
| OpenVLA-OFT | 97.1 | 59.7 | 37.2 | 81.5 | 85.8 | 92.4 | 76.7 | 77.1 |
| pi0 | 94.2 | 15.8 | 6.6 | 61.0 | 79.6 | 78.5 | 79.4 | 70.4 |
| pi0-FAST | 85.5 | 66.4 | 24.8 | 63.3 | 73.0 | 67.7 | 75.8 | 70.3 |
| UniVLA | 95.2 | 4.3 | 50.3 | 71.8 | 59.1 | 80.0 | 25.3 | 34.3 |
| RIPT-VLA | 97.5 | 58.3 | 36.7 | 80.1 | 87.9 | 90.4 | 73.8 | 76.5 |
Camera viewpoint and robot initial state are the hardest factors, and the authors attribute that to both requiring an understanding of spatial geometry and proprioception, while lighting and background are easier. Rankings also reverse under perturbation, as pi0-FAST sits 8.7 points below pi0 on original LIBERO but scores 66.4% to pi0's 15.8% under camera changes. In a blank-instruction experiment, OpenVLA-OFT's performance on LIBERO-Object stayed largely unchanged with no language input at all, and dropped substantially only on LIBERO-10. The authors conclude that the model behaves more like a vision-action policy than a vision-language-action one.
LIBERO-Plus also has headroom that standard LIBERO lacks. Across its full benchmark (Table 2), totals range from 15.6% for OpenVLA to 69.6% for OpenVLA-OFT, and the authors' own post-training on more than 20,000 perturbed trajectories reaches 79.6%.
What Still Discriminates Between Policies
LIBERO-10. The long-horizon suite is the one standard suite with a wide spread left. Reported LIBERO-10 scores run from 53.7% (fine-tuned OpenVLA) and 60.2% (pi0-FAST, as compiled by OpenVLA-OFT) through 71% (SmolVLA) to 94.5% (OpenVLA-OFT) and 97.6% (X-VLA). LIBERO-Plus also found it was the only suite where removing the instruction hurt OpenVLA-OFT substantially. If you report one standard suite to show a difference between two policies, this is the one.
Perturbation suites. LIBERO-PRO and LIBERO-Plus reuse LIBERO's scenes and assets, so they cost little extra setup if you already run LIBERO, and both separate models that tie on the original benchmark. LeRobot already supports LIBERO-plus as an environment (--env.type=libero_plus), with the caveat in its docs that installing it replaces vanilla LIBERO in the same Python environment.
Benchmarks with harder task structure. RoboCasa is far from saturated. In the GR00T N1 paper, GR00T-N1-2B averages 32.1% across 24 RoboCasa kitchen tasks with 100 demonstrations per task and 49.6% with 300, and GR00T N1 stays below 15% on individual tasks such as opening a double door. The newer RoboCasa365 release has 365 tasks across 2,500 kitchen environments and is also available through LeRobot. When a policy scores 97% on LIBERO and struggles on RoboCasa, the gap shows limits that LIBERO alone would not have revealed.
Real robots. Simulation scores, perturbed or not, are a proxy. The case for measuring how well sim performance predicts real performance, rather than assuming it, is made well in recent NVIDIA work on sim benchmarking (Yang et al., 2025), which we discuss in How To Evaluate A VLA Policy.
What To Report Alongside LIBERO
If you publish or circulate a LIBERO result, the following makes it informative even when the headline number is near the ceiling.
- Per-suite scores, with LIBERO-10 shown separately, plus the number of episodes per task and the seeds. A four-suite average of 97% says less than a LIBERO-10 score with its confidence interval.
- The evaluation pipeline. Name the script and commit, the max steps per suite, the action chunk length executed per query, and the image preprocessing, including whether frames were rotated 180 degrees. These details move standard LIBERO scores by several points, as covered in How To Run LIBERO Evaluation and Your Robot May Be Learning The World Flipped.
- At least one perturbation result. LIBERO-Plus or LIBERO-PRO, broken out by factor. Camera viewpoint, robot initial state and object position are the factors that have separated models most in published results.
- A blank or corrupted instruction check. If the score does not move when the instruction is removed, the policy is not using language on that suite, and any claim about instruction following needs other evidence.
- One benchmark outside the LIBERO family, such as RoboCasa, for any claim about general manipulation ability.
Where Manifold Fits
Manifold is Bifrost's platform for evaluating robot policies in simulation. Its benchmark list includes LIBERO and LIBERO-Plus alongside RoboCasa, RoboMimic, CALVIN, SIMPLER, RoboTwin 2.0 and RLBench, run through one harness, so adding a perturbation suite or a second benchmark next to LIBERO does not mean building a second evaluation pipeline. Rollouts are sharded across GPUs, which brings a LIBERO sweep that takes 8 hours on a single GPU down to about 30 minutes (roughly 16x faster) and makes larger episode counts practical. Manifold also clusters failed episodes by failure mode, which is the view you want when a policy that scores 97% on LIBERO drops under camera or initial-state perturbations. If you only need LIBERO-Plus for one checkpoint, the LeRobot integration above works well. Manifold is in early access through the waitlist on the Manifold page, and Introducing Manifold explains why we built it.
Sources
- Liu et al., LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- Zhou et al., LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization, and the LIBERO-PRO repository and leaderboard
- Fei et al., LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Kim, Finn and Liang, Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)
- Zheng et al., X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Physical Intelligence, openpi LIBERO example
- Shukor et al., SmolVLA
- NVIDIA, GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Hugging Face, LeRobot LIBERO and LIBERO-plus documentation
Frequently Asked Questions
What is the highest reported LIBERO score?
As of October 2026, several open models report averages above 97% across the four standard suites, including OpenVLA-OFT at 97.1%, the LeRobot pi0.5 reproduction at 97.5%, and X-VLA at 98.1%. These come from different pipelines and episode counts, so treat the ordering among them as noise rather than a ranking.
What is the difference between LIBERO-PRO and LIBERO-Plus?
Both extend LIBERO with perturbations. LIBERO-PRO changes objects, initial positions, instructions, tasks and environments, and reports that OpenVLA, pi0 and pi0.5 fall to 0% under task changes. LIBERO-Plus perturbs seven factors (layout, camera, robot initial state, language, light, background, sensor noise) across 10,030 task variants with five difficulty levels.
Is LIBERO still worth reporting?
Yes, as a sanity check and for comparison with prior work. A policy that cannot reach the mid-90s on standard LIBERO usually has a pipeline bug. It should not be the only evidence for a generalization claim, because the evaluation tasks are the same tasks the policy was trained on.
Which LIBERO suite is hardest?
LIBERO-10, also called LIBERO-Long, which chains multiple steps per task. It still separates models, with reported scores from 53.7% for fine-tuned OpenVLA to 97.6% for X-VLA, while reported LIBERO-Object scores for the same models sit between 88.4% and 99.0%.
What benchmark should I use instead of LIBERO?
Use LIBERO alongside a perturbation suite such as LIBERO-Plus or LIBERO-PRO, and add a benchmark with harder task structure such as RoboCasa, where GR00T N1 reports a 32.1% average across 24 kitchen tasks with 100 demonstrations each. Real-robot evaluation is still the final test.