BIFROST
All posts

pi0.5 LIBERO Results, Reproduced On Manifold

We ran the public pi0.5 LIBERO checkpoint for 1,999 episodes across all four suites and averaged 96.7%, against 96.85% reported by openpi. Per-suite results and rerun spread.

The Short Answer

We evaluated the public pi0.5 LIBERO checkpoint from openpi on all four standard LIBERO suites, 1,999 episodes in total, and it averaged 96.7%. openpi reports 96.85% for the same checkpoint. Suite by suite, our numbers land within 2.2 points of theirs: 96.6% on LIBERO-Spatial, 99.8% on LIBERO-Object, 97.6% on LIBERO-Goal and 92.8% on LIBERO-10. Repeated LIBERO-10 runs of the identical setup ranged from 92.0% to 93.8%, which is the run-to-run spread you should expect at 500 episodes.

The rest of this page gives the full setup, the per-suite counts, the rerun spread, and what changed the number when we changed the harness. If you are setting up your own LIBERO evaluation, our guide on how to run LIBERO evaluation covers the protocol and the common reasons numbers drift.

Results By Suite

SuiteOur resultSuccesses / episodesopenpi reportedDifference
LIBERO-Spatial96.6%483 / 50098.8%-2.2
LIBERO-Object99.8%499 / 50098.2%+1.6
LIBERO-Goal97.6%487 / 49998.0%-0.4
LIBERO-1092.8%464 / 50092.4%+0.4
Average96.7%1,933 / 1,99996.85%-0.15

openpi's figures are from the openpi LIBERO example README, which reports per-suite success for the pi05_libero checkpoint. Our runs were completed between July 14 and July 22, 2026. The LIBERO-10 row uses the earliest complete run on that build; the next section lists the others.

The average matches openpi's to within 0.15 points, and so do LIBERO-Goal and LIBERO-10, where the gaps are well inside the noise you get from 500 episodes. Two suites differ by more than that. If openpi also ran 500 episodes per suite, a 95% interval on the difference between two such runs is about plus or minus 1.9 points for Spatial and 1.2 points for Object, so our lower Spatial score (-2.2) and higher Object score (+1.6) are probably real setup differences rather than chance. They point in opposite directions and we have not isolated the cause. For how these intervals are computed and how many episodes you need to separate two policies, see how many trials you need to evaluate a robot policy.

Run-To-Run Spread On LIBERO-10

LIBERO-10 is the long-horizon suite and the one with the most room below 100%, so we ran it several times with the same policy image.

RunBenchmark buildSuccesses / episodesSuccess
1libero-10 0.2.5464 / 50092.8%
2libero-10 0.2.5467 / 49893.8%
3libero-10 0.2.5459 / 49592.7%
4libero-10 0.2.9465 / 50093.0%
5libero-10 0.2.9460 / 50092.0%

The five runs span 1.8 points. LIBERO's initial states are fixed per episode, so the scene is identical each time, and the variation comes from the policy side, most likely pi0.5's sampled action generation and floating-point nondeterminism on the GPU. Two runs recorded fewer than 500 episodes, and the table reports them as recorded rather than padding them.

The practical point is that a single LIBERO-10 number carries about a point of noise even when nothing changes. If two checkpoints differ by less than that on one run each, the comparison does not tell you which is better.

The Harness Version Moved The Number More Than The Reruns Did

An earlier build of our LIBERO-10 benchmark container, with an earlier packaging of the same checkpoint, scored 95.6% (478 of 500) in June 2026, and a second run on that build scored 95.8%. The current builds score 92.0% to 93.8%. The checkpoint weights are the same in both, so a gap of about 3 points came from the evaluation setup alone.

That is the same lesson as the LIBERO orientation bug on a smaller scale. A LIBERO number only means something alongside the exact harness that produced it, which is why we report the build for every run here. We report the current builds, which match openpi's published LIBERO-10 figure more closely, and we have not yet isolated which change between builds accounts for the difference.

Setup

ItemValue
Checkpointgs://openpi-assets/checkpoints/pi05_libero (openpi pi05_libero config)
Policy containerghcr.io/bifrostai/pi05-libero 0.1.7
Tasks per suite10
Episodes per task50, one per LIBERO fixed initial state (indices 0 to 49)
Max steps per episode220 Spatial, 280 Object, 300 Goal, 520 LIBERO-10
Benchmark buildslibero-spatial 0.2.8, libero-object 0.2.5, libero-goal 0.2.5, libero-10 0.2.5 and 0.2.9
SimulatorLIBERO on robosuite and MuJoCo, inside the benchmark containers
Run datesJuly 14 to July 28, 2026

The step limits match those used in the OpenVLA and openpi evaluation scripts. Each run records the policy and benchmark container images by digest.

What This Page Does Not Cover

These are pi0.5 results on LIBERO only. We have not yet published results for other policies or for RoboCasa, and we will add them here as complete runs finish rather than quoting partial ones. Rollout videos exist for every episode but are not yet publicly viewable.

Where Manifold Fits

Manifold is the platform we used to run these evaluations. It runs a policy across simulator benchmarks from one harness, sharded across GPUs, so a LIBERO sweep that takes 8 hours comes back in about 30 minutes (roughly 16x). Failed episodes are clustered by failure mode rather than reduced to a single success rate. Manifold is in early access, and you can join the waitlist on the Manifold page.

Sources

Frequently Asked Questions

What success rate does pi0.5 get on LIBERO?

Physical Intelligence's openpi repository reports 96.85% averaged over the four standard suites for its fine-tuned pi0.5 LIBERO checkpoint (98.8 Spatial, 98.2 Object, 98.0 Goal, 92.4 LIBERO-10). Our run of the same public checkpoint on Manifold averaged 96.7% (96.6, 99.8, 97.6, 92.8) over 1,999 episodes.

How many episodes were these results based on?

Each suite ran 10 tasks with 50 episodes per task, using LIBERO's 50 fixed initial states per task, for 500 episodes per suite. One LIBERO-Goal episode did not record a result, so that suite is reported over 499 episodes and the total is 1,999.

Do reruns of the same pi0.5 checkpoint give the same LIBERO score?

Not exactly. Five LIBERO-10 runs of the same policy image scored between 92.0% and 93.8%. The initial states are fixed, so the spread comes from the policy and serving side rather than the scene. A 95% interval at 500 episodes and about 93% success is roughly plus or minus 2.2 points, so these runs are consistent with each other.

Which pi0.5 checkpoint was evaluated?

The public openpi checkpoint gs://openpi-assets/checkpoints/pi05_libero, which openpi trained with its pi05_libero config. It was packaged in a container and served to the LIBERO benchmark containers by Manifold.

Why might my pi0.5 LIBERO numbers differ from these?

Step limits, how many actions are executed per policy call, image orientation and preprocessing, simulator and benchmark versions, and the number of episodes all move the result. Our own older LIBERO-10 build scored 95.6% with the same checkpoint, which shows how much the harness version alone can matter.

Get access More Posts