Why VLA Policies Fail: Common Failure Modes In Simulation
VLA policies fail by grabbing the wrong object, missing grasps, closing on air, stalling and not recovering, and they get worse when cameras, layouts or wording change.
Why Do VLA Policies Fail?
Vision-language-action (VLA) policies fail in a small set of recurring ways. They grasp the wrong object or ignore the instruction, misjudge where things are, miss a grasp or close the gripper on empty air, collide with the scene, stall or loop until the episode times out, and fail to recover once something goes wrong. Most of these get much worse when the scene moves away from the training data. On LIBERO-Plus, OpenVLA-OFT falls from 97.1% to 37.2% when the robot's initial state changes, and on NVIDIA's RoboLab-120, pi0.5 drops from 28.0% to 15.3% when instructions are made vague. The common cause is that these policies are trained by imitation on successful demonstrations, so they learn a tight mapping from familiar scenes to familiar trajectories and see almost no examples of what to do after a mistake.
The rest of this post sets out a taxonomy of failure modes with how each one shows up in rollouts and how to detect it automatically, summarizes the published evidence on how far perturbations push success rates down, and describes a failure analysis workflow you can run on your own policy. The examples are mostly from simulation, where failures are cheap to reproduce and the simulator state makes them easy to label.
A Taxonomy Of VLA Failure Modes
The table groups failures by what goes wrong in the episode, not by the root cause, since the same root cause (for example, a shifted camera) can produce several different symptoms. Each row lists a published source that documents the failure.
| Failure mode | How it shows up in rollouts | How to detect it automatically | Published source or example |
|---|---|---|---|
| Language grounding, wrong object | The policy grasps a distractor or the object it usually picks in this scene, regardless of the instruction | Compare the identity of the first grasped object with the instruction's target; a wrong-object event | RoboLab tracks wrong-object grasps automatically (Yang et al., 2026). LIBERO-PRO replaced a salad dressing bottle with alphabet soup and the model grasped the soup anyway |
| Spatial reasoning | Reaches to where the object was in training, or places the object in the wrong receptacle | Distance from the end effector to the target at first grasp; final pose predicate on the goal region | On LIBERO-PRO, OpenVLA and pi0 fall to 0% once objects move more than 0.2 units from their usual positions |
| Grasp failure | Repeated grasp attempts that never lift the object, or an object dropped in transit | Contact between both fingers and the object, object height after grasp, drop events | FailureSpot reports pi0 commonly makes "repeated but unsuccessful grasp attempts" on LIBERO-10 (Ma et al., 2026). RoboLab records dropped objects |
| Ghost grasp (closing on air) | The gripper closes with nothing in it and then never reopens, so the episode stalls | Gripper closed with no finger-object contact for more than a few steps | The 2025 BEHAVIOR Challenge winners call this one of the most common failures across all tasks, and a rule that reopens the gripper roughly doubled success on affected tasks (Larchenko et al., 2025) |
| Collision | The arm or gripper strikes furniture, fixtures or other objects, sometimes knocking the target away | Contact forces between robot links and non-target bodies above a threshold | RoboLab records gripper collisions as a failure event |
| Timeout or stalling | The arm hovers in a near-fixed pose, oscillates, or repeats the same motion until the step budget runs out | End-effector speed below a threshold for N steps, repeated action chunks, no subgoal progress | FailureSpot finds OpenVLA "frequently enters stagnation" and pi0-FAST produces excessive swinging. VLA-FAIL describes "infinite loop" retries |
| Recovery failure | One mistake (a nudged object, a missed grasp) and the policy never gets back on track | First failed subgoal followed by no further progress; compare against episodes with the same early error that recovered | FailSafe notes that robot datasets provide successful trajectories, and the few with failures offer only text explanations. NVIDIA's RoboLab writeup shows pi0.5 grasping an orange by mistake and then recovering |
| Perturbation sensitivity, lighting | Success drops when light intensity, color or shadows change | Paired runs from the same initial states under nominal and changed lighting | LIBERO-Plus: OpenVLA falls from 76.5% to 4.4% under lighting changes. RoboLab: 90 to 100% success held across its lighting variations |
| Perturbation sensitivity, camera pose | Small camera shifts cause misreaches or freezing | Paired runs with the camera pose offset by known amounts | LIBERO-Plus: pi0 falls from 94.2% to 15.8% under camera changes. RoboLab: success depends on the wrist camera staying near its nominal pose |
| Perturbation sensitivity, distractors | Extra objects pull the policy to the wrong target or slow it down | Success as a function of object count in the scene | RoboLab: success dropped sharply as the number of objects increased |
| Perturbation sensitivity, instruction paraphrase | A reworded or vague instruction changes behavior, or a nonsense instruction does not | Paired runs with the original, paraphrased and blank instruction | RoboLab: pi0.5 drops from 28.0% to 15.3% with vague instructions. LIBERO-PRO: nonsense strings such as "xxx" produce nearly the same trajectory |
Two rows deserve a closer look because they say something about how these policies are trained.
The ghost grasp comes from the training data. Larchenko, Zarin and Karnatak explain that there are almost no training examples in which the robot opens the gripper after closing it, so once a grasp misses, the policy has no learned behavior for trying again. Their fix was a rule outside the model, which treats a gripper that closes at a stage where it was never closed in training as a failed grasp and opens it fully. They report that this rule alone roughly doubled success on the tasks where grasping was a common failure, and measured a 2.2x increase in q-score on a subset of 13 tasks with 3 episodes each.
The paraphrase row cuts both ways. A policy that fails under paraphrase is brittle to wording, but a policy that succeeds under nonsense input is not reading the instruction at all. The LIBERO-Plus authors found that OpenVLA-OFT's performance on LIBERO-Object stayed largely unchanged with no valid language input and dropped substantially only on LIBERO-10, and concluded that it behaves more like a vision-action policy. A high score under paraphrase is only evidence of language grounding if a blank or wrong instruction lowers it.
How Much Do Perturbations Hurt?
Perturbation benchmarks give the clearest numbers on how narrow a policy's competence is. A few rows from LIBERO-Plus, Table 1 (success rate, %):
| Model | Original | Camera | Robot initial state | Language | Light | Layout |
|---|---|---|---|---|---|---|
| OpenVLA | 76.5 | 1.1 | 4.1 | 26.8 | 4.4 | 31.6 |
| OpenVLA-OFT | 97.1 | 59.7 | 37.2 | 81.5 | 85.8 | 77.1 |
| pi0 | 94.2 | 15.8 | 6.6 | 61.0 | 79.6 | 70.4 |
| pi0-FAST | 85.5 | 66.4 | 24.8 | 63.3 | 73.0 | 70.3 |
The LIBERO-Plus authors found camera viewpoint and robot initial state the hardest factors, since both need an understanding of spatial geometry and proprioception. Rankings also change under perturbation, with pi0-FAST below pi0 on the original tasks and well above it under camera changes. RoboLab adds a different angle. It fine-tunes policies only on real DROID data and evaluates them on 120 new tasks, where the best of five policies, pi0.5, reached 28.0%, and the sensitivity analysis found policies tolerant of lighting, background and table textures but dependent on the wrist camera pose, the number of objects and the distance to the object.
We go through LIBERO-PRO and LIBERO-Plus in more detail, including why standard LIBERO scores hide these drops, in Is LIBERO Saturated?.
Not every collapse is the policy's fault. When we tested models on LIBERO, some that score above 90% in their authors' own setups scored close to zero in others, and the cause was an image orientation mismatch between the training data and the simulator. We wrote that up in Your Robot May Be Learning The World Flipped. Before you attribute a failure mode to the policy, rule out a mismatch in camera orientation, image resolution, action space or control frequency between training and evaluation.
How To Run Failure Analysis On A VLA Policy
Failure analysis turns a success rate into a list of things to fix. The workflow below works for any simulated benchmark, and most of it works on real robots with more manual labeling.
- Record everything for every episode. Save the camera video the policy saw, the actions, the gripper state, contacts and object poses from the simulator, the seed and the initial state. You cannot cluster failures you did not record, and you will want to replay successful episodes too, to check that they did not succeed by accident.
- Score subgoals, not only the final outcome. Break each task into atomic steps (reach, grasp, lift, move, place, release) and write a predicate for each against simulator state. RoboLab's graded task score gives partial credit for completed subtasks for the same reason. Per-subgoal scores tell you where an episode stopped making progress, which is usually where the failure started.
- Tag events with rules first. Wrong-object grasps, drops, collisions, ghost grasps and stalls can all be detected from simulator state with the checks in the table above. These rules are cheap and reproducible, and they label the bulk of failures before you look at any video.
- Cluster what the rules miss. Group the remaining failures by where they stopped (first failed subgoal), by scene factors (objects, layout, camera, lighting) and by trajectory shape. A vision-language model reading the episode video and event log can propose labels for clusters, but check a sample by hand before trusting them.
- Rank modes by impact, not frequency. A mode that appears in many failed episodes may not be the one costing you the most success, if those episodes would have failed later anyway. Attribute each failed episode to its first failure event, then estimate how many episodes would succeed if that mode were removed, using the episodes where the same early error was recovered from as a guide.
- Run enough episodes to see the rare modes. A mode behind 5% of failures may not appear at all in a 50-episode run. Our guide to how many trials you need has the numbers for success rates, and the same reasoning applies to the share of each failure mode.
- Confirm with paired perturbations. Once you suspect a cause (say, camera pose), rerun the same initial states with only that factor changed. A paired comparison isolates the factor in a way that an observational cluster cannot.
For a broader checklist that covers benchmark choice and what to publish alongside the score, see How To Evaluate A VLA Policy.
Learned Failure Detectors
Rule-based checks need simulator state. On a real robot, or when you want a warning before a failure happens rather than a label afterwards, a family of recent papers trains detectors on the policy's own signals.
| Method | Signal | Evaluated on |
|---|---|---|
| SAFE (2025) | Detector on the VLA's internal features that outputs a failure likelihood, tested on unseen tasks | OpenVLA, pi0 and pi0-FAST in simulation and the real world |
| SAFECAST (2026) | Hidden-state probe trained with contrast sets of visual and language perturbations, with conformal calibration of the failure threshold | LIBERO-Spatial and LIBERO-Plus in simulation, DROID Franka in the real world |
| FailureSpot (2026) | Timestamp-level detector trained with action-derived weak labels on 15% of trajectories | LIBERO-10 with pi0, pi0-FAST and OpenVLA |
| VLA-FAIL (2026) | Last-layer Mahalanobis distance plus consistency between overlapping action chunks, with no failure data needed | LIBERO-Plus and six real-world tasks with pi0.5 and X-VLA |
| ReconVLA (2026) | Conformal prediction on action tokens and a Mahalanobis distance on robot state | LIBERO-Object and a UR5 arm with pi0 and OpenVLA-OFT |
These methods are complementary to failure analysis rather than a replacement. A detector tells you an episode is going wrong. Failure analysis tells you which kind of wrong is costing you the most across the whole run. FailureSpot's timestamp-level AUROC on unseen tasks reached 90.5% for pi0 against 76.6% for SAFE, but 66.5% for OpenVLA, so detection quality still varies a lot by policy.
Where Manifold Fits
Manifold is Bifrost's platform for evaluating robot policies in simulation, and failure analysis is built into it. Every task decomposes into atomic subgoals that are scored one by one, and you can define custom rubrics to grade rollout quality. Agents watch the rollouts, cluster the failures and rank them by impact, grouping them by scenario, objects, sensor, lighting and trajectory, so you can query failures instead of scrubbing through hundreds of videos, and replay any rollout in 3D. Rollouts are sharded across GPUs, so a LIBERO sweep that takes 8 hours comes back in about 30 minutes (roughly 16x), which makes running enough episodes to see the rarer failure modes practical. Policy and benchmark also declare their conventions, such as camera orientation, and Manifold checks they line up before a run starts. Manifold is in early access, and you can join the waitlist from the Manifold page.
Sources
- Fei et al., LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models, 2025
- Zhou et al., LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization, 2025
- Yang et al. (NVIDIA), RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies, RSS 2026, and How to Evaluate General-Purpose Robot Policies for Real-World Deployment, NVIDIA Technical Blog, July 2026
- Ma, Liu and Zhu, FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models, 2026
- Rajaprakash et al., SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration, 2026
- Seligmann et al., VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models, 2026
- Chen, Lyu and Beksi, ReconVLA: An Uncertainty-Guided and Failure-Aware Vision-Language-Action Framework for Robotic Control, 2026
- Gu et al., SAFE: Multitask Failure Detection for Vision-Language-Action Models, 2025
- Lin et al., FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models, 2025
- Larchenko, Zarin and Karnatak, Task Adaptation of Vision-Language-Action Model: 1st Place Solution for the 2025 BEHAVIOR Challenge, 2025
Frequently Asked Questions
What is the most common failure mode of VLA policies?
It depends on the policy and the task, so measure it rather than assume it. On LIBERO-10, the FailureSpot authors found that pi0 most often makes repeated but unsuccessful grasp attempts, pi0-FAST more often makes excessive swinging motions, and OpenVLA more often stalls in a nearly fixed pose. The first-place team in the 2025 BEHAVIOR Challenge reported that a failed grasp that closes the gripper on empty air was one of the most common failures across all their tasks.
Why do VLA policies ignore the language instruction?
Benchmarks like LIBERO let a policy succeed by recognizing the scene rather than reading the instruction, because each scene is tied to a small set of tasks. LIBERO-PRO found that models execute nearly the same trajectory when the instruction is replaced with a meaningless string, and LIBERO-Plus found that OpenVLA-OFT's performance on LIBERO-Object stayed largely unchanged with no valid language input at all.
Which perturbation hurts VLA policies the most?
In LIBERO-Plus, camera viewpoint and robot initial state were the hardest factors, with pi0 falling from 94.2% to 15.8% under camera changes and to 6.6% under robot initial state changes. In NVIDIA's RoboLab, policies were most sensitive to the wrist camera staying close to its nominal pose, while lighting changes had little effect.
How do you detect robot policy failures automatically?
In simulation, the most reliable method is to check the simulator state directly, with predicates for each subgoal, contact checks for collisions and grasps, and motion thresholds for stalling. Learned detectors such as SAFE, SAFECAST, FailureSpot and VLA-FAIL go further and predict failure from the policy's internal features or action chunks, which is useful when you need a warning before the failure happens.
How many failed episodes do I need for failure analysis?
Enough that each failure mode you care about appears several times, which usually means hundreds of rollouts per task suite rather than tens. A mode that causes 5% of failures in a 50-episode run may appear only once or not at all, so small runs tend to show the most frequent mode and miss the rest.