BIFROST
MANIFOLD · ROBOTICS

Know How Your Robot Policy Will Behave Before It Ships

Manifold runs manipulation, mobile and humanoid policies across thousands of GPU-parallel simulated rollouts, scores every subgoal against rubrics you define, and shows you exactly where each checkpoint breaks.

Evaluation Is the Bottleneck in Robot Learning

Training compute keeps getting cheaper. Knowing whether a new checkpoint is actually better has not. Real-world trials are slow, expensive and impossible to repeat exactly, so most teams fall back on a handful of hand-run episodes and a success rate that hides how the policy failed.

Simulation should fix this, but every benchmark arrives with its own harness, every policy has its own interface, and a full sweep still runs overnight. Manifold is the evaluation layer that removes that work, so the loop from checkpoint to insight takes minutes.

EVALUATING ROBOTS TODAY
  • 01A full simulator sweep runs overnight, so teams test a fraction of their checkpoints
  • 02Every benchmark and every policy needs its own harness, rebuilt by every lab
  • 03Pass and fail scores hide the subgoal where a rollout actually went wrong
  • 04Finding a failure pattern means scrubbing through hundreds of rollout videos by hand

Every Checkpoint, Evaluated Over Lunch

Point Manifold at a checkpoint and pick your benchmarks. Rollouts shard across GPUs automatically, so a sweep that took a night comes back before you have finished the next experiment.

ANY EMBODIMENT
Single-arm, bimanual, mobile manipulators and humanoids on one harness.
NO HARNESS TO WRITE
Skip the simulator setup, GPU orchestration and policy wiring.
MINUTES, NOT NIGHTS
A LIBERO sweep that took 8 hours lands in 30 minutes.

Scores That Match What Your Task Needs

A grasp that succeeds by knocking the object over is still a failure on the factory floor. Break every task into atomic subgoals and grade rollouts against the quality bar your deployment demands.

ATOMIC SUBGOALS
Approach, grasp, transport and place, each scored on its own.
CUSTOM RUBRICS
Encode precision and timing requirements for your task.
ROLLOUT QA
Catch the degraded behavior an aggregate success rate hides.

Failure Modes, Found for You

Agents watch every rollout, group the failures by cause and rank them by how much performance they cost, so you know what to fix in the next training run.

AUTOMATIC CLUSTERING
Failures grouped by object, lighting, pose and trajectory.
RANKED BY IMPACT
See which modes cost you the most, not which happened most often.
ASK IN PLAIN LANGUAGE
Query your failures instead of searching through videos.
TASK COVERAGE

From Academic Benchmarks to Your Own Production Tasks

Run the suites the field reports on, then codify your own real-world work into simulation tasks with success rubrics written for it.

ENVIRONMENT LIBRARY Custom benchmark
ANY BENCHMARK · ANY SIMULATOR · ANY WORLD MODEL
ACADEMIC
LIBERO
LIBERO-Plus
RoboLab
RoboCasa
RoboMimic
CALVIN
SIMPLER
RoboTwin 2.0
RoboMemory
RLBench
TASKS
Order Kitting
Wire Ethernet Cables
Tidy Desk
Wipe Table
Clean Kitchen
Clear Laundry
Operate Microwave
Operate Tap
Tend CNC Machine
Operate Lightswitch
SIMULATORS
NVIDIA Isaac Sim
Unreal Engine
MuJoCo
ManiSkill
Genesis
USE CASES

Where Robotics Teams Use Manifold

VLA policy evaluation

Score vision-language-action models across every major benchmark in one run.

Checkpoint selection

Evaluate every checkpoint from a training run and keep the one that generalizes.

Regression testing

Catch a policy getting worse on old tasks before it reaches hardware.

Industrial manipulation

Kitting, machine tending and assembly, graded against your own spec.

Household and service tasks

Kitchens, laundry and tidying across varied layouts and objects.

Humanoid and mobile manipulation

Whole-body and mobile tasks as embodiments scale.

Sim-to-real gap analysis

Compare simulated and real performance to see where the gap lives.

Model comparison

Compare two architectures on identical tasks, seeds and conditions.

SIMULATORS
NVIDIA Isaac SimIsaac LabMuJoCoManiSkillGenesisUnreal Engine

Manifold evaluates OpenVLA, GR00T, pi-0 and Octo class models, and any policy you can serve behind an inference endpoint.

THE VOCABULARY WE WORK IN
VLAmanipulation policyimitation learningreinforcement learningworld modeldomain randomizationsim-to-realbimanualdexterityhumanoidmobile manipulationLIBERORoboCasaCALVINrolloutsuccess ratesubgoalcheckpoint
FAQ

How do you evaluate a robot policy in simulation?

Serve the policy behind an inference endpoint, choose the benchmarks and tasks, and Manifold runs thousands of rollouts in parallel across GPUs. You get subgoal-level scores, rollout video and clustered failure modes for every task.

Which embodiments does Manifold support?

Single-arm and bimanual manipulators, mobile manipulators and humanoids, on simulators including NVIDIA Isaac Sim, Isaac Lab, MuJoCo, ManiSkill and Genesis.

Can I evaluate on my own tasks, not just academic benchmarks?

Yes. We work with industrial teams to codify real-world work, from kitting to machine tending, into simulation tasks with success rubrics written for that work.

How long does an evaluation take?

Most teams have a scored run within 30 minutes of onboarding. A LIBERO sweep that takes 8 hours lands in about 30 minutes on Manifold.

Does Bifrost also make synthetic training data for robotics?

Yes. Stardust, our synthetic data platform, generates labeled perception data for robotics and autonomy. Manifold is the product for evaluating the policies themselves.

Make Every Checkpoint Count

We onboard robotics teams in small batches. Join the waitlist and we'll reach out as soon as your spot opens.

Get early access HOW MANIFOLD WORKS →