BIFROST
All posts

How Many Trials Do You Need To Evaluate A Robot Policy?

About 100 trials gives a 95% interval of roughly ±8 to ±10 points on one success rate. Telling two policies 10 points apart takes 199 to 388 trials each.

The Short Answer

For a single success rate, 100 trials gives a 95% confidence interval of about ±9.6 points at 50% success and ±7.8 points at 80%, while 500 trials narrows that to about ±4.4 and ±3.5 points. Comparing two policies takes more: detecting a 10-point gap (70% vs 80%) with 80% power at a two-sided alpha of 0.05 needs 294 trials per policy, and a 5-point gap needs between 435 and 1,565 per policy depending on the baseline. At 50 trials per policy, even a 20-point gap is missed more than 20% of the time.

Those numbers come from the binomial math below, which you can rerun for your own success rates and budgets. The rest of this post covers how to read them, why per-task numbers need more trials than aggregates, and how paired comparisons on shared initial states reduce the count. For the broader evaluation workflow (benchmark choice, reproducibility, failure analysis), see our guide on how to evaluate a VLA policy.

Why Success Rate Needs An Interval

A success rate is the number of successful episodes divided by the number of episodes. Each episode is a pass or fail draw, so the count of successes follows a binomial distribution, and the observed rate is an estimate of the true rate with sampling error that shrinks only with the square root of the number of trials. Quadrupling the trials halves the interval.

That is why two policies reported as 90% and 92% on 100 trials each are not distinguishable. The 95% Wilson interval for 90 out of 100 runs from 82.6% to 94.5%, and for 92 out of 100 from 85.0% to 95.9%. The intervals overlap almost entirely. Detecting a true 2-point gap at that level would take 3,213 trials per policy.

NVIDIA's RoboLab writeup on evaluating general-purpose robot policies (July 2026) makes the same point with Clopper-Pearson intervals: 70 rollouts at 90% success give an interval from 80.5% to 95.9%, and getting to a ±2-point band (88.0% to 91.8%) takes 1,030 rollouts, about 15 times more. We reproduced both intervals with the code below.

Lookup Table: 95% Interval Half-Widths

The table gives the half-width of the 95% interval in percentage points, computed as (upper bound minus lower bound) divided by 2. Both intervals are asymmetric near 0% and 100%, so the half-width is a summary of width, not a symmetric plus-or-minus.

Trials (n)p = 0.50 Wilsonp = 0.50 Clopper-Pearsonp = 0.80 Wilsonp = 0.80 Clopper-Pearsonp = 0.95 Wilsonp = 0.95 Clopper-Pearson
20±20.1±22.8±16.8±19.0±11.4±12.4
50±13.4±14.5±10.9±11.8±6.6±6.6
100±9.6±10.2±7.8±8.3±4.5±4.8
200±6.9±7.1±5.5±5.8±3.1±3.3
500±4.4±4.5±3.5±3.6±1.9±2.0
1,000±3.1±3.1±2.5±2.5±1.4±1.4

A few readings that come up often:

  • 20 trials at 80% success gives a Wilson interval from 58.4% to 91.9%. That is a 33-point range, which is the precision of many real-robot results.
  • 50 trials at 95% gives 85.1% to 98.4% (Wilson). The policy might be failing one episode in seven.
  • 1,000 trials at 50% still leaves ±3.1 points, so very fine distinctions near the middle of the range are expensive.

Method

Wilson is the score interval for a binomial proportion. Clopper-Pearson inverts the exact binomial test using beta quantiles, so it guarantees at least 95% coverage and comes out slightly wider. For Clopper-Pearson we used k = round(p × n) successes. Here is the code we ran:

from scipy.stats import norm, beta
import math

z = norm.ppf(0.975)  # 1.96 for a 95% interval

def wilson(p, n):
    centre = (p + z*z/(2*n)) / (1 + z*z/n)
    half = z * math.sqrt(p*(1-p)/n + z*z/(4*n*n)) / (1 + z*z/n)
    return centre - half, centre + half

def clopper_pearson(k, n):
    lo = 0.0 if k == 0 else beta.ppf(0.025, k, n - k + 1)
    hi = 1.0 if k == n else beta.ppf(0.975, k + 1, n - k)
    return lo, hi

print(wilson(0.8, 20))         # (0.584, 0.919)
print(clopper_pearson(63, 70)) # (0.805, 0.959), NVIDIA's 70-rollout example

Avoid the textbook normal approximation (p plus or minus 1.96 times the standard error) for small n or success rates near the ends. At 20 out of 20 it returns a zero-width interval, which is wrong.

Trials Needed To Detect A Gap Between Two Policies

An interval answers "how well do I know this one number." Comparing policies asks a different question: how many trials per policy do I need so that a real gap of a given size shows up as statistically significant most of the time. The table below uses the standard two-proportion sample size formula (normal approximation, pooled variance under the null), two-sided alpha of 0.05 and 80% power, for two independent sets of trials.

Baseline successDetect +5 pointsDetect +10 pointsDetect +20 points
50%1,56538893
60%1,47135682
70%1,25129462
80%906199n/a
90%435n/an/a

Each figure is trials per policy, so the total budget is double. Gaps near 50% are the most expensive because the variance of a binomial peaks there. We checked the 70% vs 80% row by simulation: 200,000 simulated comparisons at 294 trials per policy rejected the null 80.4% of the time, matching the 80% target.

The formula, if you want to plug in your own numbers:

def n_per_policy(p1, p2, alpha=0.05, power=0.80):
    za, zb = norm.ppf(1 - alpha/2), norm.ppf(power)
    pbar = (p1 + p2) / 2
    num = (za*math.sqrt(2*pbar*(1-pbar)) + zb*math.sqrt(p1*(1-p1) + p2*(1-p2)))**2
    return math.ceil(num / (p2 - p1)**2)

print(n_per_policy(0.70, 0.80))  # 294

If you plan to analyze with Fisher's exact test or Barnard's test, which are common choices for small samples, budget a little more than this table, since exact tests are more conservative than the normal approximation.

Per-Task Versus Aggregate Numbers

Benchmark suites report an aggregate over many tasks, and the aggregate is much tighter than any single task. A 10-task suite run at 50 rollouts per task gives 500 rollouts. At 90% overall success, the Wilson interval on the pooled rate is 87.1% to 92.3%, about ±2.6 points. Each task on its own has only 50 rollouts, so a task at 80% carries roughly ±10.9 points.

When every task gets the same fixed number of rollouts, the simple pooled binomial interval is slightly conservative, because the variance of a stratified average is never larger than the pooled binomial variance. So the aggregate interval from the lookup table is a safe upper bound on the width.

The catch is that the aggregate answers a narrower question than most people ask it. A policy can gain 3 points on a suite by improving one task by 30 points and staying flat elsewhere, or by improving everything a little. To claim a per-task improvement, apply the gap table to that task's rollout count, and if you test many tasks, correct for multiple comparisons (Holm or Bonferroni) before calling any single task a win.

The LIBERO Protocol And What 50 Rollouts Buys

LIBERO fixes a set of benchmark initial states for each task, and the common VLA evaluation protocol runs 50 rollouts per task, one from each initial state. OpenVLA's LIBERO evaluation script sets num_trials_per_task to 50 and loads the benchmark's fixed initial states, and the OpenVLA paper reports 500 trials per task suite averaged over three random seeds, or 1,500 trials per statistic. The original LIBERO paper used 20 test rollouts per task in its lifelong-learning experiments, so check which protocol a number comes from before comparing.

At 50 rollouts per task and 500 per suite, the lookup table says you can trust a suite-level number to within about ±2 to ±4 points and a task-level number to within about ±7 to ±13 points. Scores near the ceiling make this worse. In the PolaRiS study, nearly all tested VLAs scored between 90% and 95% on LIBERO-90, a spread of the same order as the per-task intervals.

Seeds, Initial States, And Paired Comparisons

Three sources of randomness move a success rate: the initial state of the scene, the simulator's own stochasticity, and the policy's action sampling. They behave differently.

  • Initial states carry most of the task variation. More distinct initial states make the estimate represent the task distribution better.
  • Seeds vary simulator and policy noise. If every seed replays the same 50 initial states, three seeds do not give you 150 independent draws from the task distribution. The scene-level variation is shared, and the interval computed as if n were 150 will be too narrow. Report seeds and initial states separately.
  • Paired comparisons run both policies from the same initial states with the same seeds, then compare outcomes episode by episode. Pairing removes the variance that comes from some initial states being easier than others.

For paired pass/fail outcomes, McNemar's test is the standard choice, and it only uses the discordant pairs, the initial states where one policy succeeded and the other failed. Using a standard normal-approximation sample size formula for McNemar's test, detecting a 10-point gap at alpha 0.05 and 80% power needs 155 paired initial states if 20% of pairs are discordant and 234 if 30% are, compared with 294 trials per policy for the unpaired 70% vs 80% case above. The saving depends on how correlated the two policies are, so estimate the discordance rate from a pilot run.

Paired evaluation is easy in simulation, where initial states are reproducible by construction. On real hardware it means resetting the scene to a recorded configuration for each policy, which is more work but follows the same logic. Kress-Gazit and colleagues make a similar case in Robot Learning as an Empirical Science, arguing that papers should report the number of runs, the initial conditions and the success criteria alongside the score.

Sequential Testing And Stopping Early

The tables above assume you fix the sample size in advance. A common failure is to run 50 trials, look, run 50 more because the result is close, and then test as if 100 had been planned. That inflates the false positive rate.

Sequential tests are built for this. STEP (Snyder et al., RSS 2025) is a sequential test for comparing two policies on success and failure outcomes that lets you decide whether to keep running trials based on intermediate results while keeping its error guarantees. The authors report that it reduces the number of evaluation trials by up to 32% compared with state-of-the-art baselines. If your trials are expensive, especially on hardware, a sequential test is usually the better tool than a fixed-n table.

A Quick Budgeting Rule

  1. Decide the smallest gap that would change a decision, such as 5 or 10 points.
  2. Look up trials per policy in the gap table at your expected baseline.
  3. If you can pair on shared initial states, run a small pilot to estimate the discordance rate and size the run with McNemar's formula instead, which usually needs fewer initial states.
  4. Report n, the interval method and the interval for every number, and per-task counts alongside the aggregate.
  5. For zero-failure results, report the bound, not the perfect score. Zero failures in n trials puts the 95% upper bound on the failure rate at roughly 3 divided by n, so 20 out of 20 is consistent with a 15% failure rate.

Where Manifold Fits

The tables make the cost plain. Distinguishing close policies takes hundreds to thousands of rollouts per policy, per benchmark. Manifold is Bifrost's robot policy evaluation platform: one harness runs your policy on simulator benchmarks including LIBERO, RoboCasa and your own scenarios, with rollouts sharded across GPUs so thousands of rollouts are practical to run. Agents then watch the rollouts, cluster the failures and rank them by impact, which helps with the per-task question that aggregates hide. Manifold is in early access through the waitlist. If you only need a handful of comparisons on one benchmark, a local harness with the code above will do the job. Our Manifold launch post has more on how it works.

Sources

Frequently Asked Questions

How many trials do I need to tell whether policy A is better than policy B?

It depends on the gap you want to detect and the baseline success rate. With a two-sided test at alpha 0.05 and 80% power, a 20-point gap needs 62 to 93 trials per policy, a 10-point gap needs 199 to 388, and a 5-point gap needs 435 to 1,565. Running both policies from the same initial states and using a paired test can cut these numbers substantially.

Is 50 trials per task enough for LIBERO?

Fifty rollouts per task is the protocol most VLA papers follow, as in OpenVLA's evaluation code, giving 500 rollouts per 10-task suite. That is enough to report a suite-level success rate to within about ±3 points near 90% success, but each individual task is only known to within about ±7 to ±13 points, so per-task comparisons between close policies are rarely conclusive at that count.

Should I use the Wilson or Clopper-Pearson interval for a success rate?

Both are fine and both beat the normal approximation, which misbehaves near 0% and 100%. Wilson is slightly narrower and its coverage stays close to 95% on average. Clopper-Pearson is exact and guarantees at least 95% coverage, so it is a little wider, which is why some groups such as NVIDIA's RoboLab writeup use it for conservative reporting.

What does it mean if my policy succeeds in every trial?

Zero failures in n trials does not mean a 0% failure rate. A useful rule of thumb is that the 95% upper bound on the failure rate is about 3 divided by n, so 20 out of 20 is consistent with a failure rate as high as 15% and 50 out of 50 with one as high as 6%. Report the interval, not just the perfect score.

Do more random seeds count as more trials?

Only partly. If every seed replays the same fixed set of initial states, the variation that comes from scene layout is shared across seeds, so three seeds over 50 initial states are not 150 independent samples of the task distribution. Seeds capture simulator and policy-sampling noise, while more distinct initial states capture task variation, and you need both.

Get access More Posts