How To Run LIBERO Evaluation And Reproduce Published Results
Run each LIBERO suite from its 50 fixed initial states, 50 episodes per task, with step limits of 220 to 520, then match the training image rotation and action chunking.
How To Run A LIBERO Evaluation
To run a LIBERO evaluation that matches published numbers, load each task's fixed initial states from the LIBERO package, run 50 episodes per task (one per stored initial state) for each of the four standard suites, wait 10 no-op steps for objects to settle, cap each episode at the suite's step limit (220 for LIBERO-Spatial up to 520 for LIBERO-10), and count an episode as a success when the environment reports the task goal as met. Use the evaluation script from the same codebase that produced the checkpoint, because the image preprocessing (including a 180 degree rotation), the action chunk length and the seed all move the score. Report per-suite success rates, the episode count and the seeds, and average over 3 seeds if you want to compare against the OpenVLA and OpenVLA-OFT papers.
The rest of this post covers the suites, the protocol, verified commands from the four codebases people use most (LIBERO, OpenVLA and OpenVLA-OFT, openpi, and LeRobot), a table of published scores to check against, and the reasons a reproduction usually comes out a few points low or close to zero.
The LIBERO Task Suites
LIBERO (Liu et al., NeurIPS 2023) has 130 tasks. Three suites isolate one kind of distribution shift each, and LIBERO-100 is split into LIBERO-90 for pretraining and LIBERO-10 for held-out testing. VLA papers almost always report the four 10-task suites and call LIBERO-10 "LIBERO-Long."
| Suite | CLI name | Tasks | Max steps (OpenVLA, openpi) | Max steps (LeRobot) | Longest training demo |
|---|---|---|---|---|---|
| LIBERO-Spatial | libero_spatial | 10 | 220 | 280 | 193 steps |
| LIBERO-Object | libero_object | 10 | 280 | 280 | 254 steps |
| LIBERO-Goal | libero_goal | 10 | 300 | 300 | 270 steps |
| LIBERO-10 (Long) | libero_10 | 10 | 520 | 520 | 505 steps |
| LIBERO-90 | libero_90 | 90 | 400 | 400 | 373 steps |
The step limits come from the max_steps values in the OpenVLA evaluation script, which OpenVLA-OFT and the openpi LIBERO client reuse. Each limit carries a comment giving the length of the longest training demonstration in that suite, and each sits a little above it. LeRobot's LIBERO environment uses the same table except for LIBERO-Spatial, where it allows 280 steps instead of 220. A policy that is slow but eventually succeeds will score higher under the longer limit, so this one value is enough to make two "LIBERO-Spatial" numbers disagree.
The Standard Evaluation Protocol
The protocol most VLA papers follow is the one the OpenVLA team wrote into their evaluation script, and the openpi and OpenVLA-OFT scripts copy it closely.
- Fixed initial states. LIBERO ships a file of initial states for every task. The file we inspected for the first LIBERO-Spatial task holds an array of shape 50 by 92, which is 50 stored starting configurations of the scene. The scripts call
task_suite.get_task_init_states(task_id)and set episodeito initial statei, so every lab evaluates the same 50 starting scenes per task. - 50 episodes per task.
num_trials_per_taskdefaults to 50, giving 500 episodes per suite and 2,000 across the four standard suites. - Settling steps. The first 10 steps (
num_steps_wait = 10) send a no-op action so objects can settle after reset. These do not count against the step limit. - Seeds. OpenVLA, OpenVLA-OFT and openpi all default to
seed = 7. The openpi script also seeds the environment and notes in a comment that the seed seems to affect object positions even when a fixed initial state is used. OpenVLA's published table averages 3 random seeds of 500 rollouts each. - Success criterion. LIBERO uses a sparse reward that fires when the task's goal predicates are satisfied. The scripts end the episode and record a success as soon as
env.stepreturnsdone.
LeRobot's LIBERO docs recommend a lighter protocol of 10 episodes per task (400 in total) averaged over 3 seeds, and report a 95% Wilson interval with every success rate. At that size, they note, a task scored 9 out of 10 has an interval of roughly 60% to 98%. If you are comparing against a 50-episode result, run 50.
Commands From The Official Repos
Each command below is copied from the current README or docs of the named repository as of October 2026. Flags not shown here are left at the script defaults, which are the values the authors used.
LIBERO itself
The LIBERO README installs into a Python 3.8.13 conda environment:
conda create -n libero python=3.8.13 conda activate libero git clone https://github.com/Lifelong-Robot-Learning/LIBERO.git cd LIBERO pip install -r requirements.txt pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 --extra-index-url https://download.pytorch.org/whl/cu113 pip install -e .
LIBERO's own libero/lifelong/evaluate.py evaluates the lifelong-learning baselines that ship with the benchmark (BC-RNN, BC-Transformer, BC-ViLT). For VLA checkpoints, use the model's own evaluation script.
OpenVLA
After installing LIBERO, the OpenVLA README installs its extra requirements and runs one suite per fine-tuned checkpoint:
cd openvla pip install -r experiments/robot/libero/libero_requirements.txt python experiments/robot/libero/run_libero_eval.py \ --model_family openvla \ --pretrained_checkpoint openvla/openvla-7b-finetuned-libero-spatial \ --task_suite_name libero_spatial \ --center_crop True
Swap the checkpoint and --task_suite_name for libero_object, libero_goal and libero_10. The README stresses that --center_crop True is required because the model was fine-tuned with random crops covering 90% of the image area. The published numbers used Python 3.10.13, PyTorch 2.2.0, transformers 4.40.1 and flash-attn 2.5.5 on an A100.
OpenVLA-OFT
The OpenVLA-OFT LIBERO guide uses the same script name with fewer flags, because the defaults are already set for the OFT checkpoints (two input images, proprioception, center crop, 8-step action chunks):
python experiments/robot/libero/run_libero_eval.py \ --pretrained_checkpoint moojink/openvla-7b-oft-finetuned-libero-spatial \ --task_suite_name libero_spatial
A single checkpoint trained on all four suites, moojink/openvla-7b-oft-finetuned-libero-spatial-object-goal-10, scored 96.8% against 97.1% for the four per-suite checkpoints. The guide asks you to evaluate on the same GPU type you trained on, or to merge the LoRA weights on the evaluation device, because otherwise "performance may drop substantially."
openpi (pi0, pi0-FAST, pi0.5)
The openpi LIBERO example runs the policy as a server and LIBERO as a client. With Docker:
git submodule update --init --recursive sudo xhost +local:docker SERVER_ARGS="--env LIBERO" docker compose -f examples/libero/compose.yml up --build
To pick a suite or checkpoint, set CLIENT_ARGS="--args.task-suite-name libero_10" or SERVER_ARGS="--env LIBERO policy:checkpoint --policy.config pi05_libero --policy.dir ./my_custom_checkpoint". Without Docker, run uv run scripts/serve_policy.py --env LIBERO in one terminal and python examples/libero/main.py in a second terminal from the LIBERO virtual environment. If you hit EGL errors, prefix the command with MUJOCO_GL=glx. The client defaults to libero_spatial, so pass the suite explicitly for the other three.
LeRobot
LeRobot wraps LIBERO as a Gym environment and evaluates any LeRobot policy with lerobot-eval. LIBERO in LeRobot requires Linux.
pip install -e ".[libero]" export MUJOCO_GL=egl lerobot-eval \ --policy.path="your-policy-id" \ --env.type=libero \ --env.task=libero_spatial,libero_object,libero_goal,libero_10 \ --eval.batch_size=1 \ --eval.n_episodes=10 \ --env.max_parallel_tasks=1
For its pi0.5 reproduction, the docs add --policy.n_action_steps=10. LeRobot preserves LIBERO's hard reset by default and offers --env.init_states=true --env.hard_reset=false as a faster soft reset, but the docs say soft resets are not bit-identical and recommend hard resets when reproducing benchmark results.
Published LIBERO Scores To Check Against
These are success rates (%) as reported by each source. Rows are not all directly comparable: they differ in inputs (third-person camera only, or with wrist camera and proprioception), training data (the filtered OpenVLA datasets or the originals), and episodes per task.
| Model | Spatial | Object | Goal | Long | Average | Source |
|---|---|---|---|---|---|---|
| OpenVLA (fine-tuned) | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | OpenVLA README, 3 seeds x 500 rollouts |
| OpenVLA-OFT | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 | OpenVLA-OFT paper, Table I |
| pi0 (fine-tuned) | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 | Compiled in OpenVLA-OFT Table I |
| pi0-FAST (fine-tuned) | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 | Compiled in OpenVLA-OFT Table I |
| pi0.5 @ 30k steps | 98.8 | 98.2 | 98.0 | 92.4 | 96.85 | openpi LIBERO README |
| pi0.5 (LeRobot reproduction) | 97.0 | 99.0 | 98.0 | 96.0 | 97.5 | LeRobot LIBERO docs, 10 episodes per task |
| SmolVLA (0.45B) | 90 | 96 | 92 | 71 | 87.3 | SmolVLA paper, Table 2, 10 trials per task |
| pi0 (3.3B, re-run by SmolVLA authors) | 90 | 86 | 95 | 73 | 86.0 | SmolVLA paper, Table 2 |
The last row is the useful one for anyone debugging a reproduction. The same model family that its authors' pipeline reports at 94.2% scored 86.0% when another team trained and evaluated it in their own stack, with LIBERO-Object dropping from 98.8% to 86%. Neither number is wrong, since they measure different pipelines, but the table cannot tell you which pipeline choices caused the gap.
Common Reasons Your Numbers Don't Match
When a reproduction is close to zero, the cause is almost always input preprocessing. When it is a few points to a few tens of points low, it is usually one of the protocol details below.
Image orientation
MuJoCo returns frames through OpenGL bottom-row-first, so raw LIBERO images are upside down. The OpenVLA get_libero_image helper and the openpi client both apply img[::-1, ::-1], a 180 degree rotation, with a comment that it matches training preprocessing. A 180 degree rotation is a vertical flip plus a horizontal mirror, so the standard LIBERO training data shows the scene mirrored left to right. If you train on those datasets, you must apply the same rotation at test time. If you "fix" it with a vertical flip only, the policy sees an unmirrored scene it never trained on. We wrote up the history and the evidence (a milk carton label that reads backwards) in Your Robot May Be Learning The World Flipped.
Action chunking and replan steps
Chunked policies predict several actions per query, and the evaluation script decides how many of them to execute before querying again. The openpi LIBERO client defaults to replan_steps = 5. The LeRobot pi0.5 reproduction sets n_action_steps=10 and describes that as matching the original openpi implementation. OpenVLA-OFT executes 8 actions per query and prints a warning when num_open_loop_steps does not match the chunk size the model was trained with. The effect can be large. In the SmolVLA paper's ablation (Table 13), executing 1, 10, 30 and 50 actions per query gave LIBERO averages of 80.3%, 82.8%, 70.8% and 51.8% for the same ablation model.
Max steps and settling steps
As the suite table shows, LIBERO-Spatial is capped at 220 steps in the OpenVLA and openpi scripts and 280 in LeRobot. Skipping the 10 settling steps, or counting them against the limit, also changes results on the tasks where a policy finishes near the cap.
Episodes, seeds and initial states
Fifty episodes per task and 10 per task are both in use, and the spread at 10 is wide enough to hide a 5 point difference. Use the stored initial states in order, keep the seed fixed when comparing two policies, and report the seed list. LeRobot's docs add that comparing two policies on the same episodes needs the same --seed, --env.init_states=true, and each task run in a single batch.
Simulator and package versions
LIBERO's requirements.txt pins robosuite==1.4.0, while the OpenVLA and openpi LIBERO requirements pin robosuite==1.4.1, and openpi's compiled lockfile pins mujoco==3.2.3. Record the exact versions you ran. The OpenVLA and OpenVLA-OFT authors also ask you to keep their PyTorch and transformers versions, and both note that a different GPU can shift results because of nondeterminism in large models.
Rendering, resize and crop
The OpenVLA and openpi scripts render at 256 by 256 and resize to 224 by 224 for the model (openpi pads to keep the aspect ratio). OpenVLA checkpoints need --center_crop True. LeRobot's soft reset path can change camera observations slightly compared with a hard reset. Any of these, and the choice between EGL and GLX rendering backends, is worth checking against the training pipeline before you look at the model.
Training data revision
The OpenVLA datasets on Hugging Face filter out no-op actions and failed demonstrations, and OpenVLA-OFT reports results separately for filtered and unfiltered data. LeRobot's docs recommend pinning --dataset.revision to a commit hash when you report results, because Hub datasets can be re-uploaded.
Where Manifold Fits
Manifold is Bifrost's platform for evaluating robot policies in simulation, and LIBERO is one of the benchmarks it runs. Rollouts are sharded across GPUs, so a LIBERO sweep that takes 8 hours on a single GPU comes back in about 30 minutes (roughly 16x faster), which makes the full 50-episode protocol across all four suites practical to run every time instead of a 10-episode spot check. Manifold has precomputed baselines for pi0.5, GR00T N1.5, OpenVLA-7B and Octo-base on LIBERO to compare against. Policies and benchmarks declare their image orientation conventions through the Manifold SDK, and a run will not start if the declarations do not line up, which is how it catches the rotation mismatch described above. If you only need one LIBERO number for one checkpoint, the official scripts above are the right tool. Manifold is in early access, and you can join the waitlist on the Manifold page.
For the broader question of which benchmarks to run and how many rollouts a result needs, see How To Evaluate A VLA Policy. For why Manifold exists, see Introducing Manifold.
Sources
- Liu et al., LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, and the LIBERO repository
- Kim et al., OpenVLA, and the OpenVLA LIBERO evaluation script
- Kim, Finn and Liang, Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT), and the OpenVLA-OFT LIBERO guide
- Physical Intelligence, openpi LIBERO example
- Hugging Face, LeRobot LIBERO documentation
- Shukor et al., SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Frequently Asked Questions
How many episodes per task does LIBERO evaluation use?
The OpenVLA, OpenVLA-OFT and openpi evaluation scripts default to 50 episodes per task, which is 500 per suite and 2,000 across the four standard suites. LeRobot's docs recommend 10 episodes per task (400 in total) and averaging over 3 seeds. Say which one you used, because the confidence interval at 10 episodes per task is much wider.
What is the difference between LIBERO-10 and LIBERO-Long?
They are the same suite. LIBERO-10 is the 10-task held-out split of LIBERO-100, and most VLA papers label it LIBERO-Long because its tasks are multi-step. The CLI name in every major codebase is libero_10.
Why does my LIBERO success rate drop to near zero?
The most common cause is an image orientation mismatch. The standard LIBERO training datasets store camera frames rotated 180 degrees, and the OpenVLA and openpi evaluation scripts apply the same rotation at test time. If you feed the policy unrotated or vertically flipped frames, a model trained on those datasets sees a different world than it trained on.
Do I need a GPU to run LIBERO evaluation?
LIBERO itself runs on MuJoCo and renders headless with MUJOCO_GL=egl on Linux, but VLA policies such as OpenVLA, pi0 and pi0.5 need a GPU for inference. OpenVLA and OpenVLA-OFT report their published numbers on an NVIDIA A100 and warn that results can shift on a different GPU.
Which LIBERO max episode steps should I use?
The OpenVLA and openpi scripts use 220 steps for LIBERO-Spatial, 280 for LIBERO-Object, 300 for LIBERO-Goal, 520 for LIBERO-10 and 400 for LIBERO-90, plus 10 settling steps. LeRobot uses the same values except 280 for LIBERO-Spatial, so match the limit used by the result you are comparing against.