BIFROST
All posts

Isaac Lab vs MuJoCo vs ManiSkill vs Genesis For Policy Evaluation

For policy evaluation, the benchmark usually picks the simulator. LIBERO and RoboCasa run on MuJoCo, SimplerEnv on ManiSkill, Isaac Lab-Arena tasks on Isaac Lab.

The Short Answer

For robot policy evaluation, the simulator is usually chosen by the benchmark you need to report. LIBERO and RoboCasa run on MuJoCo through robosuite, SimplerEnv runs on SAPIEN and has a GPU-parallel port in ManiSkill3, and the RoboCasa, LIBERO and RoboLab-style task suites in Isaac Lab-Arena run on Isaac Lab. Beyond that, pick Isaac Lab when you need RTX ray-traced cameras and NVIDIA's evaluation tooling, MuJoCo when you need published-comparable LIBERO and RoboCasa numbers or CPU-only portability, ManiSkill3 when you want thousands of rendered environments per GPU on Linux, and Genesis when your tasks need deformables, fluids or non-NVIDIA hardware.

Versions move quickly, so everything below is current as of October 2026: Isaac Lab 3.0 Early Access (released September 16, 2026, with general availability targeted for the end of October 2026), MuJoCo 3.14.0 (September 22, 2026), ManiSkill 3.0.1 (April 21, 2026) and Genesis World 1.4.3 (September 30, 2026).

Feature Matrix

Every row below was checked against the project's own repository or documentation.

Isaac LabMuJoCo (with MJX and MuJoCo Warp)ManiSkill3Genesis World
Physics backendPhysX through Isaac Sim; Isaac Lab 3.0 adds OVPhysX and Newton (MuJoCo Warp) backendsMuJoCo C library on CPU; MJX in JAX; MuJoCo Warp on NVIDIA GPUsSAPIEN, using PhysX for CPU and GPU simulationIn-house multi-solver engine: rigid, FEM, MPM, PBD and SPH particles, IPC, SAP
GPU-parallel simulationYes, vectorized environments on each GPU (Isaac Lab-Arena reports 1,024 on one RTX 5880 Ada)Not in the CPU library (parallelize with processes); yes with MJX and MuJoCo WarpYes, on Linux with NVIDIA GPUs, including heterogeneous scenes per environmentYes, parallel and heterogeneous environments
RenderingRTX ray tracing, batched through TiledCamera; 3.0 adds OVRTX and a Newton Warp rendererOpenGL rasterizer; MuJoCo Warp adds a ray-traced GPU batch rendererVulkan rasterizer by default, plus rt, rt-med and rt-fast ray-tracing shader packsNyx (in-house), Luisa ray tracer, Pyrender rasterizer
LicenseBSD-3 (Isaac Lab); Isaac Sim code is Apache 2.0 with additional NVIDIA-licensed componentsApache 2.0 (MuJoCo, MuJoCo Warp, MuJoCo Playground)Apache 2.0 for rigid body environments; assets CC BY-NC 4.0Apache 2.0
HardwareIsaac Sim needs an RTX GPU with RT cores (A100 and H100 not supported), minimum RTX 4080 with 16 GB VRAM, 32 GB RAMAny CPU; MJX on NVIDIA and AMD GPUs, Apple Silicon, TPUs; MuJoCo Warp needs an NVIDIA GPU for fast simulationGPU simulation on Linux with NVIDIA only; Vulkan driver required for renderingCompiles to CUDA, AMD ROCm, Apple Metal, Vulkan, x86 and ARM64
PythonPython 3.12 for Isaac Lab 3.0pip install mujoco, Python 3.10 or laterpip install mani_skill, Gymnasium APIpip install genesis-world, Python 3.10 to 3.13
Evaluation benchmarks on itIsaac Lab-Arena: Lightwheel-RoboCasa-Tasks, Lightwheel-LIBERO-Tasks, a DROID Kitchen Benchmark, RoboLab-inspired tasksLIBERO (robosuite 1.4.0), RoboCasa and RoboCasa365 (robosuite 1.5), MuJoCo PlaygroundSimplerEnv GPU port, ManiSkill3 tasks; RoboTwin 2.0 builds on SAPIEN directlyNone of LIBERO, RoboCasa or SimplerEnv run on it natively

Why The Benchmark Usually Decides

A success rate means something only next to other success rates measured the same way. That requires the same task logic, the same assets, the same physics engine and version, and the same camera pipeline. If you want to say your policy beats a published LIBERO number, you need to run LIBERO where LIBERO runs, which is MuJoCo through robosuite.

Ports are useful, but they are new benchmarks. Lightwheel's RoboCasa tasks on Isaac Lab-Arena use higher-fidelity assets with refined collision geometry than the original MuJoCo version, according to Lightwheel's own performance study. That is an improvement for many purposes, and it also means a score on the port and a score on the original are not the same measurement.

Rendering conventions matter as much as physics. We found that the most widely used LIBERO training datasets store images mirrored relative to the simulated scene, a side effect of how MuJoCo returns OpenGL buffers and how conversion scripts corrected for it (the full write-up is here). Moving a policy between simulators multiplies the number of places a convention like that can drift.

Isaac Lab

Isaac Lab is NVIDIA's robot learning framework, built on Isaac Sim. Its full-featured workflows use PhysX for physics and RTX for rendering, and its TiledCamera sensor collects camera data from every parallel environment through a single render product, which is what makes rendered evaluation across thousands of environments practical (camera docs). The framework is described in Mittal et al., 2025.

Isaac Lab 3.0 Early Access changes the shape of the stack. It introduces one task API across multiple physics, rendering and visualization backends, selected at launch (physics=isaacsim_physx, physics=ovphysx or physics=newton_mjwarp, and renderer=isaacsim_rtx, renderer=ovrtx or renderer=newton_renderer), plus kit-less execution that does not install or launch Isaac Sim (release notes). Supported backend combinations remain task-specific.

The hardware constraint is the one that catches teams. Isaac Sim's requirements state that GPUs without RT cores, including the A100 and H100, are not supported, with a GeForce RTX 4080 as the minimum and an RTX 5080 recommended. If your cluster is H100s, RTX-rendered Isaac Lab evaluation needs separate hardware.

For evaluation specifically, Isaac Lab-Arena (Apache 2.0) is the layer to look at. NVIDIA co-developed it with Lightwheel, and it compiles environments from independent object, scene, embodiment and task blocks (NVIDIA Technical Blog). Lightwheel has built more than 250 tasks for it across the Lightwheel-RoboCasa-Tasks and Lightwheel-LIBERO-Tasks suites. The 0.3 alpha release on September 10, 2026 added 31 Kitchen Benchmark environments for DROID manipulation, 38 RoboLab-inspired tasks generated from natural-language prompts, environment variation sweeps, multi-node evaluation through OSMO, and a policy server for GR00T, pi0.5 or your own policy (release announcement).

Use Isaac Lab when you need photorealistic ray-traced cameras, you are evaluating on Isaac Lab-Arena tasks, your policy targets NVIDIA's GR00T stack, or you need large numbers of parallel rendered environments and have RTX hardware to run them.

MuJoCo, MJX And MuJoCo Warp

MuJoCo is the physics engine underneath most of the manipulation benchmarks VLA papers report. LIBERO pins robosuite 1.4.0, and RoboCasa uses robosuite 1.5. RoboCasa365, released February 18, 2026, has 365 tasks and more than 2,500 kitchen scenes. Both benchmarks step the standard MuJoCo library on CPU, so evaluation throughput comes from running many processes rather than from GPU parallelism.

MuJoCo's documentation states that its simulation pipeline is entirely deterministic, with the caveat that exact reproducibility is only guaranteed within a single version and on the same architecture (computation docs). For evaluation that is a useful property, because a run on a fixed version and machine type can be repeated exactly.

The GPU side of MuJoCo has grown into three related projects:

  • MJX provides a JAX API for MuJoCo. Its JAX implementation runs on NVIDIA and AMD GPUs, Apple Silicon and Google Cloud TPUs, and it can also use MuJoCo Warp as a backend.
  • MuJoCo Warp is a GPU implementation for NVIDIA hardware, maintained by Google DeepMind and NVIDIA as part of the Newton project. It includes a GPU batch renderer that ray-traces MuJoCo scenes across many parallel worlds. It is also the physics behind Isaac Lab 3.0's Newton backend.
  • MuJoCo Playground is a suite of GPU-accelerated environments built with MJX, covering dm_control classics, quadruped and biped locomotion, and non-prehensile and dexterous manipulation, with vision support through the MuJoCo Warp batch renderer.

mjlab sits between the two worlds. It pairs Isaac Lab's manager-based API with MuJoCo Warp under Apache 2.0, and needs an NVIDIA GPU for training.

The catch for policy evaluation is that LIBERO and RoboCasa have not moved to these GPU backends in their official releases. MJX and MuJoCo Warp make new GPU-native tasks fast, but they do not speed up the existing benchmarks unless someone ports them.

Use MuJoCo when you need numbers comparable with published LIBERO or RoboCasa results, you want bit-exact repeatability on a fixed version, or you need to run on CPUs, Macs or non-NVIDIA accelerators. Use MJX or MuJoCo Warp when you are writing new GPU-parallel tasks and want to stay in the MuJoCo model format.

ManiSkill3

ManiSkill3 is built on SAPIEN, which runs PhysX on the GPU by placing every parallel environment in its own sub-scene of one shared physics scene. Its paper reports that simulation with rendering runs 10 to 1,000 times faster with 2 to 3 times less GPU memory than other platforms, reaching more than 30,000 FPS in benchmarked environments (Tao et al., 2024), and the README cites RGBD plus segmentation collection at more than 30,000 FPS on an RTX 4090. It also supports heterogeneous simulation, where every parallel environment has a different scene.

The platform limits are specific. GPU simulation works only on Linux with NVIDIA GPUs; Windows and macOS get CPU simulation and rendering, and WSL gets neither GPU simulation nor rendering. Rendering needs a Vulkan driver. The default shader is a rasterizer, and ray-traced shader packs (rt, rt-med, rt-fast) are available when you need more realistic images.

For evaluation, ManiSkill3's main draw is SimplerEnv, the real-to-sim benchmark for Google Robot and WidowX Bridge policies. The original SimplerEnv runs on SAPIEN and the CPU-based ManiSkill2, and its Bridge environments have been integrated into ManiSkill3 with GPU parallelization. ManiSkill describes its real-to-sim environments as evaluating real-world policies 100 times faster through GPU simulation. RoboTwin 2.0, a bimanual benchmark, also builds on SAPIEN directly.

Read the license before you build on it. Rigid body environments are Apache 2.0, but ManiSkill's assets are CC BY-NC 4.0, which rules out commercial use of those assets.

Use ManiSkill3 when you are evaluating on SimplerEnv, you want the highest rendered-environment throughput per GPU on a Linux NVIDIA machine, or you need every parallel environment to hold a different scene.

Genesis

Genesis World (Apache 2.0) is the broadest physics engine of the four. One scene can combine rigid, FEM, MPM, PBD and SPH particle, IPC and SAP solvers sharing one state, which matters for tasks involving cloth, food, liquids or soft objects. Its Quadrants compiler lowers Python kernels to CUDA, AMD ROCm, Apple Metal, Vulkan, x86 and ARM64, so the same scene code runs on NVIDIA, AMD and Apple hardware. Rendering comes through three camera sensor paths: Nyx, Genesis's in-house renderer for robotics, the Luisa ray tracer, and the Pyrender rasterizer.

The gap is benchmark coverage. None of LIBERO, RoboCasa or SimplerEnv run on Genesis natively, so evaluating on it today mostly means authoring your own tasks and success criteria. That is fine for internal evaluation and harder for results you want to compare with the literature.

Use Genesis when your tasks need deformable or fluid physics, you are not on NVIDIA hardware, or you are building a custom evaluation suite and do not need parity with published benchmark numbers.

Should You Evaluate On More Than One?

A policy that holds up across two physics engines and two renderers is a stronger result than one that holds up on one. Contact models, friction, solver tolerances and image pipelines all differ, and a policy that has overfit to one simulator's quirks tends to show it when moved. NVIDIA's own benchmarking research makes a related argument for measuring how well simulated performance predicts real performance rather than assuming it does (Yang et al., 2025).

The cost is integration work. Each simulator has its own observation format, action space conventions, camera conventions and reset semantics, and each benchmark defines success differently. Our guide on how to evaluate a VLA policy covers the reproducibility checklist that applies once you run more than one, and the companion post on evaluating robot policies in simulation at scale covers how to make the rollout counts affordable.

Where Manifold Fits

Manifold is Bifrost's robot policy evaluation platform. It runs one harness across NVIDIA Isaac Sim, Unreal Engine, MuJoCo, ManiSkill and Genesis, with native support for NVIDIA's Isaac Lab Arena, and covers academic benchmarks including LIBERO, LIBERO-Plus, RoboLab, RoboCasa, RoboMimic, CALVIN, SIMPLER, RoboTwin 2.0, RoboMemory and RLBench. Rollouts are sharded across GPUs, and every run gets a manifold:// URI that pins the policy checkpoint, simulator version, benchmark version and seeds, so a result on one simulator can be rerun exactly.

If you only need one simulator and already have a working harness, the open tooling above is enough. If you want the same policy scored on several of them without maintaining a harness per simulator, Manifold is in early access, and you can join the waitlist on the Manifold page.

Sources

Frequently Asked Questions

Is Isaac Lab better than MuJoCo for robot policy evaluation?

Neither is better in general. Isaac Lab gives you RTX ray-traced cameras and thousands of GPU-parallel environments, while MuJoCo is the simulator behind LIBERO and RoboCasa, so it is the one you need if you want numbers that compare with published results on those benchmarks. Many teams end up running both.

Can Isaac Lab run on an A100 or H100?

Not for workflows that use Isaac Sim. NVIDIA's Isaac Sim requirements state that GPUs without RT cores, including the A100 and H100, are not supported, and list a GeForce RTX 4080 with 16 GB of VRAM as the minimum. Isaac Lab 3.0 adds a kit-less Newton path that does not install Isaac Sim, but as of October 2026 its docs do not publish a separate hardware list for it, so test before you plan around data-center GPUs.

Does LIBERO run on Isaac Lab?

The original LIBERO benchmark runs on MuJoCo through robosuite 1.4.0. Lightwheel has built Lightwheel-LIBERO-Tasks for Isaac Lab-Arena, but a port with different assets, physics and rendering is a different benchmark, so its success rates should not be compared directly with numbers reported on the original LIBERO.

What is the difference between MJX and MuJoCo Warp?

MJX is a JAX API for MuJoCo, and its JAX implementation runs on NVIDIA and AMD GPUs, Apple Silicon and Google TPUs. MuJoCo Warp is a GPU implementation written in NVIDIA Warp, maintained by Google DeepMind and NVIDIA as part of the Newton project, and tuned for NVIDIA GPUs. MJX can use MuJoCo Warp as a backend.

Can I use ManiSkill commercially?

ManiSkill's rigid body environments are released under permissive licenses such as Apache 2.0, but its assets are licensed under CC BY-NC 4.0, which prohibits commercial use. Check which assets your tasks load before using ManiSkill in a commercial evaluation pipeline.

Get access More Posts