BIFROST
All posts

How To Run LIBERO Evaluation And Reproduce Published Results

Run each LIBERO suite from its 50 fixed initial states, 50 episodes per task, with step limits of 220 to 520, then match the training image rotation and action chunking.

How To Run A LIBERO Evaluation

To run a LIBERO evaluation that matches published numbers, load each task's fixed initial states from the LIBERO package, run 50 episodes per task (one per stored initial state) for each of the four standard suites, wait 10 no-op steps for objects to settle, cap each episode at the suite's step limit (220 for LIBERO-Spatial up to 520 for LIBERO-10), and count an episode as a success when the environment reports the task goal as met. Use the evaluation script from the same codebase that produced the checkpoint, because the image preprocessing (including a 180 degree rotation), the action chunk length and the seed all move the score. Report per-suite success rates, the episode count and the seeds, and average over 3 seeds if you want to compare against the OpenVLA and OpenVLA-OFT papers.

The rest of this post covers the suites, the protocol, verified commands from the four codebases people use most (LIBERO, OpenVLA and OpenVLA-OFT, openpi, and LeRobot), a table of published scores to check against, and the reasons a reproduction usually comes out a few points low or close to zero.

The LIBERO Task Suites

LIBERO (Liu et al., NeurIPS 2023) has 130 tasks. Three suites isolate one kind of distribution shift each, and LIBERO-100 is split into LIBERO-90 for pretraining and LIBERO-10 for held-out testing. VLA papers almost always report the four 10-task suites and call LIBERO-10 "LIBERO-Long."

SuiteCLI nameTasksMax steps (OpenVLA, openpi)Max steps (LeRobot)Longest training demo
LIBERO-Spatiallibero_spatial10220280193 steps
LIBERO-Objectlibero_object10280280254 steps
LIBERO-Goallibero_goal10300300270 steps
LIBERO-10 (Long)libero_1010520520505 steps
LIBERO-90libero_9090400400373 steps

The step limits come from the max_steps values in the OpenVLA evaluation script, which OpenVLA-OFT and the openpi LIBERO client reuse. Each limit carries a comment giving the length of the longest training demonstration in that suite, and each sits a little above it. LeRobot's LIBERO environment uses the same table except for LIBERO-Spatial, where it allows 280 steps instead of 220. A policy that is slow but eventually succeeds will score higher under the longer limit, so this one value is enough to make two "LIBERO-Spatial" numbers disagree.

The Standard Evaluation Protocol

The protocol most VLA papers follow is the one the OpenVLA team wrote into their evaluation script, and the openpi and OpenVLA-OFT scripts copy it closely.

  1. Fixed initial states. LIBERO ships a file of initial states for every task. The file we inspected for the first LIBERO-Spatial task holds an array of shape 50 by 92, which is 50 stored starting configurations of the scene. The scripts call task_suite.get_task_init_states(task_id) and set episode i to initial state i, so every lab evaluates the same 50 starting scenes per task.
  2. 50 episodes per task. num_trials_per_task defaults to 50, giving 500 episodes per suite and 2,000 across the four standard suites.
  3. Settling steps. The first 10 steps (num_steps_wait = 10) send a no-op action so objects can settle after reset. These do not count against the step limit.
  4. Seeds. OpenVLA, OpenVLA-OFT and openpi all default to seed = 7. The openpi script also seeds the environment and notes in a comment that the seed seems to affect object positions even when a fixed initial state is used. OpenVLA's published table averages 3 random seeds of 500 rollouts each.
  5. Success criterion. LIBERO uses a sparse reward that fires when the task's goal predicates are satisfied. The scripts end the episode and record a success as soon as env.step returns done.

LeRobot's LIBERO docs recommend a lighter protocol of 10 episodes per task (400 in total) averaged over 3 seeds, and report a 95% Wilson interval with every success rate. At that size, they note, a task scored 9 out of 10 has an interval of roughly 60% to 98%. If you are comparing against a 50-episode result, run 50.

Commands From The Official Repos

Each command below is copied from the current README or docs of the named repository as of October 2026. Flags not shown here are left at the script defaults, which are the values the authors used.

LIBERO itself

The LIBERO README installs into a Python 3.8.13 conda environment:

conda create -n libero python=3.8.13
conda activate libero
git clone https://github.com/Lifelong-Robot-Learning/LIBERO.git
cd LIBERO
pip install -r requirements.txt
pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 --extra-index-url https://download.pytorch.org/whl/cu113
pip install -e .

LIBERO's own libero/lifelong/evaluate.py evaluates the lifelong-learning baselines that ship with the benchmark (BC-RNN, BC-Transformer, BC-ViLT). For VLA checkpoints, use the model's own evaluation script.

OpenVLA

After installing LIBERO, the OpenVLA README installs its extra requirements and runs one suite per fine-tuned checkpoint:

cd openvla
pip install -r experiments/robot/libero/libero_requirements.txt

python experiments/robot/libero/run_libero_eval.py \
  --model_family openvla \
  --pretrained_checkpoint openvla/openvla-7b-finetuned-libero-spatial \
  --task_suite_name libero_spatial \
  --center_crop True

Swap the checkpoint and --task_suite_name for libero_object, libero_goal and libero_10. The README stresses that --center_crop True is required because the model was fine-tuned with random crops covering 90% of the image area. The published numbers used Python 3.10.13, PyTorch 2.2.0, transformers 4.40.1 and flash-attn 2.5.5 on an A100.

OpenVLA-OFT

The OpenVLA-OFT LIBERO guide uses the same script name with fewer flags, because the defaults are already set for the OFT checkpoints (two input images, proprioception, center crop, 8-step action chunks):

python experiments/robot/libero/run_libero_eval.py \
  --pretrained_checkpoint moojink/openvla-7b-oft-finetuned-libero-spatial \
  --task_suite_name libero_spatial

A single checkpoint trained on all four suites, moojink/openvla-7b-oft-finetuned-libero-spatial-object-goal-10, scored 96.8% against 97.1% for the four per-suite checkpoints. The guide asks you to evaluate on the same GPU type you trained on, or to merge the LoRA weights on the evaluation device, because otherwise "performance may drop substantially."

openpi (pi0, pi0-FAST, pi0.5)

The openpi LIBERO example runs the policy as a server and LIBERO as a client. With Docker:

git submodule update --init --recursive
sudo xhost +local:docker
SERVER_ARGS="--env LIBERO" docker compose -f examples/libero/compose.yml up --build

To pick a suite or checkpoint, set CLIENT_ARGS="--args.task-suite-name libero_10" or SERVER_ARGS="--env LIBERO policy:checkpoint --policy.config pi05_libero --policy.dir ./my_custom_checkpoint". Without Docker, run uv run scripts/serve_policy.py --env LIBERO in one terminal and python examples/libero/main.py in a second terminal from the LIBERO virtual environment. If you hit EGL errors, prefix the command with MUJOCO_GL=glx. The client defaults to libero_spatial, so pass the suite explicitly for the other three.

LeRobot

LeRobot wraps LIBERO as a Gym environment and evaluates any LeRobot policy with lerobot-eval. LIBERO in LeRobot requires Linux.

pip install -e ".[libero]"
export MUJOCO_GL=egl

lerobot-eval \
  --policy.path="your-policy-id" \
  --env.type=libero \
  --env.task=libero_spatial,libero_object,libero_goal,libero_10 \
  --eval.batch_size=1 \
  --eval.n_episodes=10 \
  --env.max_parallel_tasks=1

For its pi0.5 reproduction, the docs add --policy.n_action_steps=10. LeRobot preserves LIBERO's hard reset by default and offers --env.init_states=true --env.hard_reset=false as a faster soft reset, but the docs say soft resets are not bit-identical and recommend hard resets when reproducing benchmark results.

Published LIBERO Scores To Check Against

These are success rates (%) as reported by each source. Rows are not all directly comparable: they differ in inputs (third-person camera only, or with wrist camera and proprioception), training data (the filtered OpenVLA datasets or the originals), and episodes per task.

ModelSpatialObjectGoalLongAverageSource
OpenVLA (fine-tuned)84.788.479.253.776.5OpenVLA README, 3 seeds x 500 rollouts
OpenVLA-OFT97.698.497.994.597.1OpenVLA-OFT paper, Table I
pi0 (fine-tuned)96.898.895.885.294.2Compiled in OpenVLA-OFT Table I
pi0-FAST (fine-tuned)96.496.888.660.285.5Compiled in OpenVLA-OFT Table I
pi0.5 @ 30k steps98.898.298.092.496.85openpi LIBERO README
pi0.5 (LeRobot reproduction)97.099.098.096.097.5LeRobot LIBERO docs, 10 episodes per task
SmolVLA (0.45B)9096927187.3SmolVLA paper, Table 2, 10 trials per task
pi0 (3.3B, re-run by SmolVLA authors)9086957386.0SmolVLA paper, Table 2

The last row is the useful one for anyone debugging a reproduction. The same model family that its authors' pipeline reports at 94.2% scored 86.0% when another team trained and evaluated it in their own stack, with LIBERO-Object dropping from 98.8% to 86%. Neither number is wrong, since they measure different pipelines, but the table cannot tell you which pipeline choices caused the gap.

Common Reasons Your Numbers Don't Match

When a reproduction is close to zero, the cause is almost always input preprocessing. When it is a few points to a few tens of points low, it is usually one of the protocol details below.

Image orientation

MuJoCo returns frames through OpenGL bottom-row-first, so raw LIBERO images are upside down. The OpenVLA get_libero_image helper and the openpi client both apply img[::-1, ::-1], a 180 degree rotation, with a comment that it matches training preprocessing. A 180 degree rotation is a vertical flip plus a horizontal mirror, so the standard LIBERO training data shows the scene mirrored left to right. If you train on those datasets, you must apply the same rotation at test time. If you "fix" it with a vertical flip only, the policy sees an unmirrored scene it never trained on. We wrote up the history and the evidence (a milk carton label that reads backwards) in Your Robot May Be Learning The World Flipped.

Action chunking and replan steps

Chunked policies predict several actions per query, and the evaluation script decides how many of them to execute before querying again. The openpi LIBERO client defaults to replan_steps = 5. The LeRobot pi0.5 reproduction sets n_action_steps=10 and describes that as matching the original openpi implementation. OpenVLA-OFT executes 8 actions per query and prints a warning when num_open_loop_steps does not match the chunk size the model was trained with. The effect can be large. In the SmolVLA paper's ablation (Table 13), executing 1, 10, 30 and 50 actions per query gave LIBERO averages of 80.3%, 82.8%, 70.8% and 51.8% for the same ablation model.

Max steps and settling steps

As the suite table shows, LIBERO-Spatial is capped at 220 steps in the OpenVLA and openpi scripts and 280 in LeRobot. Skipping the 10 settling steps, or counting them against the limit, also changes results on the tasks where a policy finishes near the cap.

Episodes, seeds and initial states

Fifty episodes per task and 10 per task are both in use, and the spread at 10 is wide enough to hide a 5 point difference. Use the stored initial states in order, keep the seed fixed when comparing two policies, and report the seed list. LeRobot's docs add that comparing two policies on the same episodes needs the same --seed, --env.init_states=true, and each task run in a single batch.

Simulator and package versions

LIBERO's requirements.txt pins robosuite==1.4.0, while the OpenVLA and openpi LIBERO requirements pin robosuite==1.4.1, and openpi's compiled lockfile pins mujoco==3.2.3. Record the exact versions you ran. The OpenVLA and OpenVLA-OFT authors also ask you to keep their PyTorch and transformers versions, and both note that a different GPU can shift results because of nondeterminism in large models.

Rendering, resize and crop

The OpenVLA and openpi scripts render at 256 by 256 and resize to 224 by 224 for the model (openpi pads to keep the aspect ratio). OpenVLA checkpoints need --center_crop True. LeRobot's soft reset path can change camera observations slightly compared with a hard reset. Any of these, and the choice between EGL and GLX rendering backends, is worth checking against the training pipeline before you look at the model.

Training data revision

The OpenVLA datasets on Hugging Face filter out no-op actions and failed demonstrations, and OpenVLA-OFT reports results separately for filtered and unfiltered data. LeRobot's docs recommend pinning --dataset.revision to a commit hash when you report results, because Hub datasets can be re-uploaded.

Where Manifold Fits

Manifold is Bifrost's platform for evaluating robot policies in simulation, and LIBERO is one of the benchmarks it runs. Rollouts are sharded across GPUs, so a LIBERO sweep that takes 8 hours on a single GPU comes back in about 30 minutes (roughly 16x faster), which makes the full 50-episode protocol across all four suites practical to run every time instead of a 10-episode spot check. Manifold has precomputed baselines for pi0.5, GR00T N1.5, OpenVLA-7B and Octo-base on LIBERO to compare against. Policies and benchmarks declare their image orientation conventions through the Manifold SDK, and a run will not start if the declarations do not line up, which is how it catches the rotation mismatch described above. If you only need one LIBERO number for one checkpoint, the official scripts above are the right tool. Manifold is in early access, and you can join the waitlist on the Manifold page.

For the broader question of which benchmarks to run and how many rollouts a result needs, see How To Evaluate A VLA Policy. For why Manifold exists, see Introducing Manifold.

Sources

Frequently Asked Questions

How many episodes per task does LIBERO evaluation use?

The OpenVLA, OpenVLA-OFT and openpi evaluation scripts default to 50 episodes per task, which is 500 per suite and 2,000 across the four standard suites. LeRobot's docs recommend 10 episodes per task (400 in total) and averaging over 3 seeds. Say which one you used, because the confidence interval at 10 episodes per task is much wider.

What is the difference between LIBERO-10 and LIBERO-Long?

They are the same suite. LIBERO-10 is the 10-task held-out split of LIBERO-100, and most VLA papers label it LIBERO-Long because its tasks are multi-step. The CLI name in every major codebase is libero_10.

Why does my LIBERO success rate drop to near zero?

The most common cause is an image orientation mismatch. The standard LIBERO training datasets store camera frames rotated 180 degrees, and the OpenVLA and openpi evaluation scripts apply the same rotation at test time. If you feed the policy unrotated or vertically flipped frames, a model trained on those datasets sees a different world than it trained on.

Do I need a GPU to run LIBERO evaluation?

LIBERO itself runs on MuJoCo and renders headless with MUJOCO_GL=egl on Linux, but VLA policies such as OpenVLA, pi0 and pi0.5 need a GPU for inference. OpenVLA and OpenVLA-OFT report their published numbers on an NVIDIA A100 and warn that results can shift on a different GPU.

Which LIBERO max episode steps should I use?

The OpenVLA and openpi scripts use 220 steps for LIBERO-Spatial, 280 for LIBERO-Object, 300 for LIBERO-Goal, 520 for LIBERO-10 and 400 for LIBERO-90, plus 10 settling steps. LeRobot uses the same values except 280 for LIBERO-Spatial, so match the limit used by the result you are comparing against.

Get access More Posts