BIFROST
All posts

How To Close The Sim-To-Real Gap For Perception Models

Measure the gap on held-out real data, model your sensor, randomize appearance, match the real content distribution, then fine-tune on a small real set.

The Short Answer

You close the sim-to-real gap for a perception model in stages, measuring as you go. First measure the gap by training on synthetic data and testing on held-out real data. Then fix the appearance gap (sensor noise, blur, exposure and randomized textures and lighting) and the content gap (the objects, poses and layouts your real data contains), and finally fine-tune on a small real set. In published results, sensor-effect augmentation alone added 7.28 AP on KITTI cars, structured scene randomization more than doubled AP on hard KITTI objects over plain randomization, and fine-tuning on about 200 real images lifted a synthetic-pretrained detector from 52.5 to 69.5 AP.

This guide works through those steps as a checklist, says which part of the gap each one addresses, and ends with a table of techniques and their relative cost. If you first want to know whether synthetic data helps your detector at all, start with our review of whether synthetic data works for object detection.

The Gap Has Two Parts

The most useful framing comes from Meta-Sim (Kar et al., 2019), which separates the synthetic-to-real domain gap into two kinds. The appearance gap is the difference in how things look: textures, materials, lighting, and the noise and optics of the camera. The content gap is the difference in what the scene contains: which object types appear, how many, at what scale and pose, and how they are laid out relative to each other. The authors note that most domain adaptation work, such as GAN-based image translation, targets the appearance gap, and that the two gaps are orthogonal.

The distinction matters because the fixes do not overlap. A perfect photoreal render of the wrong scene still has a content gap, and a well-matched scene distribution rendered with a clean pinhole camera still has an appearance gap. Before spending effort on either, find out which one is hurting you.

Measure The Gap Before You Try To Close It

The basic measurement is two training runs evaluated on one real test split:

  • Train on synthetic, test on real. This is the number you are trying to raise.
  • Train on real, test on real. This is the in-domain ceiling, if you have enough real labels to train a reference model.
  • Train on real data from another domain, test on your real data. This is the baseline synthetic data has to beat to be worth more than an existing public dataset.

Prakash et al. (2019) report all three on KITTI hard cars with Faster R-CNN. Structured domain randomization scored 52.5 AP, 6,000 real KITTI images scored 88.8, and 70,000 real BDD100K images scored 45.6. So the synthetic data was 36.3 points short of in-domain real data but already ahead of out-of-domain real data. Kiefer et al. (2021) show how large the gap can be in maritime work. On the SeaDronesSee benchmark, YOLOv5 trained on synthetic data alone scored 10.5 mAP@50 against 55.8 for real data, partly because the synthetic set had no life-jacket class.

Break the result down by class, object size and condition. A gap concentrated in one missing class is a content problem. A gap spread evenly across classes but worse in glare or at night usually points to appearance or sensor modeling.

A Checklist For Closing The Sim-To-Real Gap

  1. Hold out real data that never touches training. This step fixes nothing directly, but every later step depends on it. Keep a real test split that is never used for checkpoint selection or early stopping, and a separate real validation split for those decisions. Tremblay et al. (2018) note they stopped training when performance on the test set saturated, which is common in papers and inflates results. Gaidon et al. (2016) found that validating the stopping point of real fine-tuning was critical to avoid overfitting a small real set.

  2. Model your sensor, not a perfect camera. This addresses the low-level appearance gap. Renderers produce clean images, and real cameras add blur, noise, chromatic aberration, exposure shifts and color casts. Carlson et al. (2018) applied a physically based pipeline for those five effects to synthetic training images and raised Faster R-CNN car AP on KITTI from 54.60 to 61.88 with 2,975 Virtual KITTI images. The augmented 2,975-image set beat the full 21,000-image unaugmented set (58.25). Match resolution, field of view, lens distortion and mounting position to your real rig as well.

  3. Randomize appearance more than you think is reasonable. This also targets the appearance gap, by making the model indifferent to it rather than by matching it. Domain randomization (Tobin et al., 2017) renders scenes with random textures, lighting and camera positions so that the real world looks like one more variation. Tobin et al. trained an object localizer on non-realistic simulated images only and reached 1.5 cm accuracy on real images. Tremblay et al. extended the idea to car detection, where 100,000 domain-randomized images (78.1 AP) came close to a photorealistic KITTI replica (79.7 AP).

  4. Keep scene structure while you randomize. This addresses the content gap that plain randomization creates. Tremblay et al. traced their remaining errors to ignoring context, such as how parked cars sit along a road. Structured domain randomization keeps realistic layouts (vehicles in lanes, roads with sidewalks) and randomizes everything else. On full KITTI it scored 65.6 AP on moderate and 52.2 on hard cars, against 38.8 and 24.0 for plain randomization and 52.6 and 42.1 for 200,000 GTA V frames (Prakash et al., 2019). In their ablation, removing context cut AP from 52.5 to 45.3, the second-largest drop after removing random saturation (43.9).

  5. Match the real content distribution. This is the content gap proper. List the classes, sizes, poses, densities and conditions in your real target data and check that the synthetic set covers them. Meta-Sim learns scene attributes to match real data automatically, but a manual audit catches most problems, such as Kiefer et al.'s missing life jackets and tricycles. Nowruzi et al. (2019) concluded that diversity mattered more than photorealism, after the game-derived Playing for Benchmarks data trained better detectors than the more photorealistic Synscapes.

  6. Use neural rendering or image translation for the last stretch of appearance. This targets the appearance gap with a learned model. CyCADA (Hoffman et al., 2018) translated GTA5 images toward Cityscapes and raised segmentation from 21.7 to 39.5 mIoU with a DRN-26 backbone, against 67.4 for training on real Cityscapes labels. Richter et al. (2021) feed the renderer's intermediate buffers into the enhancement network and change how training patches are sampled, to reduce artifacts such as the hallucinated objects seen in earlier methods, and Cosmos-Transfer1 from NVIDIA conditions a diffusion world model on segmentation, depth and edge inputs for sim-to-real transfer. The risk with any generative pass is geometry drift. If the output moves an object, its label is wrong, so condition on the scene's structure and spot-check labels after generation.

  7. Target the rare cases explicitly. This is a content fix for class imbalance. Applied Intuition's nuImages case study raised cyclists from 0.3% of instances in the real data to 27.4% in the synthetic set, and cyclist AP rose from 0.342 to 0.386 after fine-tuning. The synthetic budget goes furthest on the classes and conditions the model currently gets wrong.

  8. Fine-tune on a small real set. This closes what is left of both gaps. Pretraining on synthetic data and fine-tuning on real data beat mixed training in Nowruzi et al. and Kiefer et al. The gain is largest when real labels are scarce. In Prakash et al., about 200 real KITTI images gave 69.5 AP after structured-randomization pretraining against 55.8 alone, and with all 6,000 the difference was 0.2 AP.

  9. Re-measure after every change, then repeat. Rerun the synthetic-to-real and real-to-real numbers on the same held-out split after each fix. Closing the gap is an iterative loop of finding the largest failure, generating data that covers it, and retraining.

Techniques Compared

The cost column is a rough relative guide to engineering effort, not a price.

TechniqueGap it addressesPublished effectRough cost
Sensor-effect modelingAppearance (low-level)+7.28 AP on KITTI cars from 2,975 Virtual KITTI images (Carlson et al.)Low
Domain randomizationAppearance, by invariance78.1 AP vs 79.7 for a photoreal KITTI replica (Tremblay et al.)Low to medium
Structured domain randomizationContent (context and layout) plus appearance52.2 vs 24.0 AP on hard KITTI cars over plain randomization (Prakash et al.)Medium
Photorealistic renderingAppearance (high-level)Did not beat a more diverse, less realistic set in Nowruzi et al.High
Learned content distributionContentMeta-Sim improved downstream task performance over a hand-built scene grammar (Kar et al.)Medium to high
Image translation or neural renderingAppearance21.7 to 39.5 mIoU, GTA5 to Cityscapes (CyCADA)Medium
Targeted rare-class generationContent (imbalance)Cyclist AP 0.342 to 0.386 (Applied Intuition)Low to medium
Fine-tuning on a small real setBoth55.8 to 69.5 AP with about 200 real images (Prakash et al.)Low, plus labeling

If you have plenty of unlabeled real images but no simulator, image translation alone may be the cheapest place to start. If you have a few hundred labeled real images, fine-tuning a synthetic-pretrained model is almost always worth doing. If your real data is missing whole classes or conditions, no amount of appearance work will help until the content is there.

Where Stardust Fits

Stardust is Bifrost's synthetic data platform, and several of the steps above are built into it. Sensor perturbation (noise, intrinsics and perspective shifts) and conditions such as fog, mist, rain, snow, high winds and night operations are scenario parameters. Scenes draw on a library of more than 1,000 objects and render RGB, IR, depth and segmentation in registration, with ground-truth labels generated automatically for every frame from the 3D scene. For content and rare cases, the Real2Sim workflow takes captured footage, rebuilds failures as 3D assets and measures the fix after retraining.

For the appearance gap, our neural rendering preview adds a generative pass on top of the 3D engine that varies lighting, weather, backgrounds and surface detail while the scene geometry, and therefore the labels, stays fixed. Our maritime detection case study shows the measure-and-patch loop from step 9 at small scale: a first model trained on 2,000 synthetic images missed oblong buoys on real MODS footage, and 500 targeted synthetic images lifted F1 from just over 70% to 82%.

If you would rather start from open models, NVIDIA has released Cosmos-Transfer1's weights and code, which cover much of the appearance side of this checklist. The evaluation discipline in steps 1 and 9 applies whichever tool you use.

Sources

Frequently Asked Questions

What is the sim-to-real gap in computer vision?

It is the drop in accuracy when a model trained on simulated images is tested on real ones. Kar et al. (Meta-Sim) split it into an appearance gap, meaning differences in texture, lighting and sensor characteristics, and a content gap, meaning differences in which objects appear, where, and in what layouts. The two need different fixes.

How do you measure the sim-to-real gap?

Train the same model once on synthetic data and once on real data, then test both on the same held-out real split. The difference is the gap. Prakash et al. report 52.5 AP for structured domain randomization against 88.8 AP for 6,000 real KITTI images on hard cars, and it is worth also testing a model trained on real data from another domain, which scored 45.6 in that study.

Is domain randomization better than photorealistic rendering?

Neither wins on its own. Tremblay et al. matched a photorealistic KITTI replica with deliberately unrealistic randomized images (78.1 vs 79.7 AP), but unstructured randomization did poorly on small, context-dependent objects. Structured domain randomization, which keeps realistic scene layout while randomizing appearance, beat both plain randomization and 200,000 photorealistic game frames on KITTI.

Can neural rendering or image translation close the gap without fine-tuning?

It can close part of the appearance gap. CyCADA raised GTA5 to Cityscapes segmentation from 21.7 to 39.5 mIoU with a DRN-26 backbone, but a model trained on real Cityscapes labels scored 67.4. Generative passes do not fix a content gap, and if they alter geometry the labels stop matching the pixels, so most teams still fine-tune on a small real set.

How much real data do I need to close the gap?

Often a few hundred labeled images. In Prakash et al., fine-tuning a synthetic-pretrained detector on about 200 real KITTI images reached 69.5 AP against 55.8 for the same 200 images alone. With all 6,000 real images the benefit shrank to 0.2 AP, so synthetic pretraining matters most when real labels are scarce.

Explore Stardust More Posts