BIFROST
All posts

How To Evaluate Robot Policies In Simulation At Scale

Eval time goes to policy inference, sim stepping, rendering and resets. Vectorize environments, shard across GPUs, batch inference and serve the policy separately.

The Short Answer

Evaluating a robot policy in simulation at scale means running hundreds to thousands of closed-loop episodes, and the wall-clock goes to four places: policy inference, physics stepping, camera rendering and environment resets. The main levers are to vectorize environments on each GPU, shard episodes across GPUs and nodes, batch policy inference across environments, and run the policy as a separate server (openpi uses a websocket server on port 8000) so simulators and policy replicas scale independently. In Lightwheel's published study, 10 RoboCasa tasks on Isaac Lab-Arena with 4,096 parallel environments finished in 0.76 hours, against 10.2 hours for the MuJoCo RoboCasa baseline, a 13.5x speedup.

The rest of this post breaks down where the time goes, works through the arithmetic for a 500-episode LIBERO run, compares the tools, and covers what breaks reproducibility once you shard.

Where Does The Wall-Clock Go?

A policy evaluation is a closed loop. The simulator renders an observation, the policy turns it into actions, the simulator steps physics with those actions, and the loop repeats until the task succeeds or hits its step limit. Nothing in one episode can run ahead of the policy, so the cost of every stage adds up per step.

StageWhat drives its costMain lever
Policy inferenceForward-pass latency times the number of calls per episodeAction chunking, batching across environments, a dedicated GPU
Physics steppingSteps per episode, solver substeps, contact complexityGPU-parallel simulation or more CPU processes
RenderingCameras per environment, resolution, rasterized or ray-tracedBatched or tiled rendering
Resets and startupScene rebuilds, asset loading, model loading, container startReusing environments, keeping workers warm

For large VLAs, inference can be the largest term. Anyscale's Ray evaluation example reports that one forward pass of GR00T-N1.7-3B takes roughly 1.6 seconds and returns 40 steps of action, of which the harness executes 8 before querying again. openpi's LIBERO evaluation script re-plans every 5 steps (replan_steps = 5), so a 220-step episode means 44 policy calls.

Rendering is the stage teams most often underestimate. openpi renders LIBERO at 256 by 256 pixels and resizes to 224 before inference, with both an agent-view and a wrist camera. At that size rendering is cheap per frame, but it scales with the number of cameras and environments, and it is the reason GPU simulators invest in batched renderers such as Isaac Lab's TiledCamera, ManiSkill3's parallel rendering system and MuJoCo Warp's batch renderer.

Resets and startup are fixed costs that look small in a serial run and come to dominate once everything else is parallel, as the example below shows.

A Worked Example: 500 LIBERO Episodes

This is illustrative arithmetic, not a benchmark result. The per-step timings are assumptions chosen to be plausible for a mid-sized VLA on one GPU, and you should replace them with numbers you measure on your own setup. The structure of the calculation is the useful part.

The job is one LIBERO suite at the openpi defaults: 10 tasks with 50 trials each (num_trials_per_task = 50), 500 episodes in total. openpi caps LIBERO-Spatial episodes at 220 steps, and waits 10 steps at the start of each episode for objects to settle.

AssumptionValue
Simulator steps per episode (10 settle steps plus 180 policy steps on average)190
Physics plus rendering per step20 ms
Policy calls per episode (180 steps, re-planning every 5)36
Inference latency, one observation120 ms
Inference latency, a batch of 8 observations160 ms
Reset per episode2 s
Startup per node (load checkpoint, build environments)60 s

Serial, one environment, one GPU. Each episode costs 190 × 20 ms = 3.8 s of simulation, 36 × 120 ms = 4.32 s of inference, and 2 s of reset, for 10.12 s. Five hundred episodes take 5,060 s, plus 60 s of startup, for 5,120 s, or 85.3 minutes and 1.42 GPU-hours.

Eight environments on one GPU, inference batched. Each episode now waits for a batch of 8, so it costs 3.8 s + 36 × 160 ms + 2 s = 11.56 s. Eight episodes run at once, so the job takes 63 waves (500 ÷ 8, rounded up). That is 63 × 11.56 s = 728.3 s, plus 60 s of startup, for 788.3 s, or 13.1 minutes and 0.22 GPU-hours. The speedup over serial is 6.5x, and the GPU-hours drop because the GPU spends less time idle between single requests.

The same setup sharded across 8 GPUs. Sixty-four episodes run at once, so the job takes 8 waves (500 ÷ 64, rounded up). That is 8 × 11.56 s = 92.5 s, plus 60 s of startup, for 152.5 s, or 2.5 minutes. The speedup over serial is 33.6x, but the run uses 8 × 152.5 s = 20.3 GPU-minutes, or 0.34 GPU-hours, more than the single-GPU batched run.

Three things fall out of this:

  1. Fixed costs take over. In the 8-GPU case, startup is 60 s of a 152.5 s run, 39% of the wall-clock. Leaving startup out of both runs, the speedup would be 54.7x. Keeping workers and policy servers warm between jobs is worth more than adding GPUs at this point.
  2. The slowest episode sets the pace. Failed episodes run to the 220-step cap, which under the same assumptions costs 230 × 20 ms + 44 × 160 ms + 2 s = 13.64 s. If every wave contains one failure, the 8-GPU run takes 8 × 13.64 s + 60 s = 169.1 s. Dynamic scheduling, where a worker picks up the next episode as soon as it finishes, recovers some of this.
  3. Batching is cheaper than sharding. Batching inference raised throughput without adding hardware. Sharding cut wall-clock further, at a higher GPU-hour cost. Which one you want depends on whether researcher time or GPU time is the scarcer resource.

Note that this example keeps LIBERO on CPU MuJoCo, as the official release does, with GPU parallelism only on the policy side. A GPU-native simulator removes the per-process simulation cost as well, which is where the larger published speedups come from.

Approaches And Tools Compared

ApproachHow it scalesTradeoffs
Isaac Lab vectorized environmentsMany environments per GPU, with batched RTX rendering through TiledCameraIsaac Sim needs RTX GPUs with RT cores (A100 and H100 are not supported); benchmarks must exist on Isaac Lab or be ported
Isaac Lab-ArenaParallel environments per GPU, multi-node through OSMO, a policy server for GR00T, pi0.5 or custom policiesAlpha (0.3, September 2026); task suites are ports of RoboCasa and LIBERO, not the originals
ManiSkill3 GPU simulationThousands of environments in one PhysX GPU scene, 30,000+ FPS of RGBD and segmentation on an RTX 4090GPU simulation on Linux with NVIDIA only; assets are CC BY-NC 4.0
CPU MuJoCo with multiprocessingOne process per environment, policy inference batched on GPURuns LIBERO and RoboCasa as published; throughput limited by CPU cores and rendering
RayRay Serve policy replicas plus simulator tasks, each scaled independently across a clusterGeneral-purpose; you still write the harness, episode assignment and result collection
ManifoldHosted platform, rollouts sharded and vectorized across GPUsEarly access through a waitlist

Published figures give a sense of the range. Isaac Lab-Arena's documentation reports a camera-free scaling benchmark of 2,390 environment-steps per second with 1,024 parallel environments on one RTX 5880 Ada GPU, and says multi-node active execution time decreased nearly in proportion to the number of GPUs (docs). Lightwheel's study ran 10 RoboCasa tasks with GR00T N1.5 on a Panda-Omron robot across 8 RTX 6000D GPUs and measured 10.7x, 12.4x and 13.5x speedups over the MuJoCo baseline at 1,024, 2,048 and 4,096 environments. The same study notes that Arena's sequential run (34.9 hours) was slower than the MuJoCo baseline because of its higher-fidelity assets, so the gain comes from parallelism. ManiSkill3's paper reports 10 to 1,000 times faster simulation with rendering than other platforms (Tao et al., 2024).

For a side-by-side of the simulators themselves, see Isaac Lab vs MuJoCo vs ManiSkill vs Genesis for policy evaluation.

Decouple The Policy From The Simulator

Running the policy as its own server is the change that makes the other levers possible. openpi's remote inference design starts a websocket server with scripts/serve_policy.py, and the environment side installs the lightweight openpi-client package, sends resized images and proprioceptive state, and receives an action chunk. The openpi docs give two reasons: it keeps robot and policy environments separate to avoid dependency conflicts, and it lets inference run on a more powerful GPU than the robot or simulator has.

At scale there is a third reason. Anyscale's example points out that a 3B VLA and Isaac Sim's physics and rendering both want a GPU, and putting them on the same one makes them contend for memory and compute. With a server in between, one policy replica can serve many simulators, and you add simulator workers when the job is simulation-bound and policy replicas when it is inference-bound. Isaac Lab-Arena's 0.3 release takes the same approach, with one client contract for GR00T, pi0.5 or a custom policy across processes, GPUs or machines.

The cost is a network hop per policy call. For a policy that returns action chunks and is queried every 5 to 8 steps, that hop is small relative to the forward pass.

Keeping Sharded Runs Reproducible

Sharding introduces ways for the same evaluation to produce different numbers. Most of them come from how episodes and seeds are handed out.

  • Assign episodes from a fixed list, not per worker. Each episode should be identified by its task and trial index, and its initial state and seed should derive from that identity, not from a worker's position or a counter. LIBERO ships fixed initial states per task, which openpi loads with get_task_init_states and applies with set_init_state. openpi also calls env.seed(7) by default, with a comment noting that the seed affects object positions even when a fixed initial state is used. If each worker seeds itself differently, the same trial index can produce a different scene on different shards.
  • Know your simulator's determinism limits. MuJoCo's simulation pipeline is deterministic, but exact reproducibility is only guaranteed within a single version and on the same architecture (MuJoCo docs). Isaac Lab produces identical results for rigid bodies and articulations on the same hardware and Isaac Sim version, can vary across hardware configurations, and PhysX does not guarantee determinism for non-rigid bodies such as cloth (Isaac Lab docs). A cluster with mixed GPU models can return slightly different trajectories for the same episode.
  • Seed the policy, too. Diffusion and flow-matching policies sample noise at every call. If the policy server draws from one shared random stream, the noise an episode receives depends on which other requests arrived first. Pass a per-episode seed with each request or keep a generator per episode.
  • Expect batching to move the bits. GPU kernels are not always batch-invariant, so the same observation can produce slightly different outputs in a batch of 8 than on its own, as Thinking Machines showed for LLM inference. In a closed contact-rich loop, a small numerical difference can change an episode's outcome.
  • Make retries idempotent. Preempted or crashed workers get their episodes rerun. Keyed episode IDs stop a retry from counting an episode twice or skipping it.
  • Record everything per episode. Store the task, trial index, seed, initial state, worker, GPU model, simulator version and checkpoint hash with every result, so a disagreement between two runs can be traced to its cause.

None of this changes how many episodes you need. A success rate on a few hundred episodes still carries a wide confidence interval, as our guide on how to evaluate a VLA policy works through, and parallel evaluation is what makes the larger episode counts needed to narrow it affordable.

Where Manifold Fits

Manifold is Bifrost's hosted platform for this whole loop. It runs thousands of rollouts sharded across GPUs, with simulation optimised at the engine, machine and cluster level, and needs no simulator setup, no cloud setup and no GPUs of your own. On LIBERO, a sweep that took 8 hours lands in about 30 minutes, roughly 16x faster. Every run gets a manifold:// URI that pins the policy checkpoint, simulator version, benchmark version and seeds, which covers the per-run record-keeping described above, and agents cluster failed rollouts by failure mode. We wrote about why we built it in Introducing Manifold.

If you already run one benchmark on hardware you control, the open tools above can get you most of the way. If you want many benchmarks, simulators and checkpoints evaluated without building the harness, Manifold is in early access, and you can join the waitlist.

Sources

Frequently Asked Questions

Can LIBERO run in parallel on a GPU?

The official LIBERO release runs on CPU MuJoCo through robosuite, so you parallelize it by running many environment processes, each with its own copy of the scene, and batching policy inference across them on the GPU. GPU-native versions such as Lightwheel-LIBERO-Tasks on Isaac Lab-Arena exist, but they are ports with different assets and physics, so their numbers are not interchangeable with the original benchmark.

What is a policy server in robot policy evaluation?

A policy server is a separate process that loads the model checkpoint once and answers action requests from simulator clients over the network. openpi uses a websocket server, started with serve_policy.py on port 8000 by default, and Isaac Lab-Arena and Anyscale's Ray example use the same client and server split. It keeps the policy's dependencies and GPU apart from the simulator's.

How much faster is Isaac Lab-Arena than sequential evaluation?

In Lightwheel's study of 10 RoboCasa tasks with GR00T N1.5 on 8 RTX 6000D GPUs, Isaac Lab-Arena with 4,096 parallel environments took 0.76 hours, against 10.2 hours for the MuJoCo RoboCasa baseline and 34.9 hours for Isaac Lab-Arena run sequentially. That is the 13.5x figure Lightwheel reports against MuJoCo.

Does running evaluations in parallel change the success rate?

It should not if episodes are assigned deterministically and the environment is the same, but there are ways it can. Batched GPU inference is not always bit-identical to single-sample inference, GPU physics can vary across hardware, and per-worker random seeds can change which initial states get sampled. Record the seed, initial state and hardware for every episode so differences can be traced.

Do I need Ray to scale robot policy evaluation?

No. Ray is a general-purpose way to run simulator workers and policy replicas across a cluster, and Anyscale has published an Isaac Lab example. Isaac Lab-Arena uses OSMO for multi-node runs, and a single machine with vectorized environments and one policy server covers many evaluation jobs without any cluster scheduler.

Get access More Posts