BIFROST
All posts

Does Synthetic Data Work For Object Detection? What The Evidence Says

Yes, mostly as a complement. Pretraining on synthetic images and fine-tuning on a small real set beats real-only training in most published results.

The Short Answer

Yes, synthetic data works for object detection, but in most published results it works best as a complement to real data rather than a replacement. Pretraining on synthetic images and fine-tuning on a small real set beat real-only training on KITTI cars (69.5 vs 55.8 AP with about 200 real images), on nuImages cyclists (0.386 vs 0.342 AP), and on the SeaDronesSee maritime benchmark (60.3 vs 55.8 mAP@50). Synthetic-only training usually trails in-domain real data, sometimes by a wide margin, except on narrow, well-specified tasks where it can match or beat it.

The rest of this post collects the published numbers behind that answer, explains how much real data you need alongside synthetic, and covers the cases where synthetic data does not help. If your question is how to make the synthetic data itself transfer better, see our companion guide on how to close the sim-to-real gap for perception models.

What The Published Results Show

The table below lists results where the same detector was trained on real data, synthetic data, and a combination, then tested on real images. Metrics are not comparable across rows because each study uses its own benchmark, threshold and split. Read each row left to right.

StudyTask and metricReal onlySynthetic onlySynthetic plus real
Tremblay et al., 2018KITTI cars, Faster R-CNN, AP@0.52.1% below the fine-tuned model (6,000 images)78.1 (100K domain-randomized), 79.7 (2.5K Virtual KITTI)98.5 (domain randomization, then 6,000 real)
Johnson-Roberson et al., 2017KITTI cars, Faster R-CNN, mAP@0.7, moderate0.4274 (2,975 Cityscapes images)0.3828 (10K GTA V), 0.5257 (200K GTA V)Not tested
Prakash et al., 2019KITTI cars, Faster R-CNN, AP@0.7, hard55.8 (about 200 images), 88.8 (6,000)52.5 (25K structured randomization)69.5 (about 200 real), 89.0 (6,000 real)
Gaidon et al., 2016KITTI multi-object tracking, Fast R-CNN detections, MOTA71.9%64.3% (Virtual KITTI)76.7% (virtual pretraining, then real)
Ros et al., 2016CamVid segmentation, FCN, per-class accuracy52.862.5 (SYNTHIA-Rand)72.1 (mixed batches)
Hinterstoisser et al., 201964 retail objects, Faster R-CNN, mAP / mAP@500.54 / 0.76 (1,158 labeled images)0.67 / 0.89Not tested
Applied Intuition, 2021nuImages cyclists, Cascade Mask R-CNN, AP@0.5:0.950.342Not reported0.386 (pretrain, then 100% real), 0.376 to 0.378 (mixed)
Kiefer et al., 2021SeaDronesSee maritime search and rescue from UAV, YOLOv5, mAP@5055.810.560.3 (synthetic pretraining, then real)
Kiefer et al., 2021VisDrone aerial traffic, YOLOv5, mAP@5043.910.245.0
Kiefer et al., 2021Cattle from UAV, YOLOv5, mAP@5088.864.286.9

Three patterns come out of this table.

Synthetic plus real beats real alone in almost every row. The exception is the last one. On the simple cattle dataset, synthetic pretraining slightly hurt YOLOv5 (86.9 vs 88.8), although the same study found it helped EfficientDet-D0 and Faster R-CNN on that dataset.

Synthetic only usually trails in-domain real data. Structured domain randomization reached 52.5 AP on KITTI against 88.8 for 6,000 real KITTI images. On SeaDronesSee the gap was 10.5 against 55.8, partly because the synthetic set lacked the life-jacket class entirely. The exceptions are informative. Hinterstoisser et al. beat real data with a carefully controlled curriculum on 64 known objects. SYNTHIA's 13,400 synthetic images beat CamVid's 300-image real training split on CamVid itself. And in Driving in the Matrix and Prakash et al., the synthetic model beat real data from a different city or fleet (Cityscapes and BDD100K, tested on KITTI). Synthetic data can beat a small in-domain real set, and it is often a better bet than someone else's real dataset.

Volume matters for synthetic-only training. In Driving in the Matrix, 10,000 GTA V images scored below 2,975 real Cityscapes images on KITTI, and it took 50,000 to 200,000 to pull ahead. The authors concluded that a single simulated image carried less training value than a single real one, which is the trade you make for labels that cost nothing.

Mixing Or Fine-Tuning: Which Works Better?

There are two common ways to combine the data. Mixed training puts synthetic and real images in the same batches. Pretrain-then-fine-tune trains on synthetic data to convergence, then continues training on real data, usually at a lower learning rate.

The studies that compared both mostly favor fine-tuning:

  • Nowruzi et al. (2019) trained SSD-MobileNet on three synthetic datasets combined with 2.5%, 5% or 10% of three real datasets and concluded that fine-tuning on limited real data gives better results than mixed training, with a large increase in recall for the person class.
  • Kiefer et al. reported that synthetic pretraining followed by transfer training was superior to combined training in preliminary experiments, and used it for all their main results.
  • In Applied Intuition's case study, pretraining on synthetic data and fine-tuning on all the real nuImages data gave the best cyclist AP (0.386), ahead of mixed training at 0.5 to 1 (0.376) and 1 to 1 (0.378) synthetic to real.

Mixing still helps. SYNTHIA's mixed batches of 6 real and 4 synthetic images raised per-class accuracy by up to 18.3 points, and both of Applied Intuition's mixing ratios beat the real-only baseline on cyclists. Fine-tuning has a practical advantage on top of accuracy, which Nowruzi et al. point out. The expensive synthetic training run happens once, and the same base model can be fine-tuned cheaply for each new deployment region or camera.

How Much Real Data Do You Need?

The benefit of synthetic pretraining is largest when real data is scarce and shrinks as real data grows. Prakash et al. show this cleanly. With about 200 real KITTI images (3.3% of the 6,000 available), structured domain randomization plus real data scored 69.5 AP against 55.8 for real alone, a 13.7-point gain. With all 6,000 real images the gap closed to 89.0 against 88.8.

The other way to read the same curve is as a real-data saving. In Applied Intuition's study, synthetic pretraining plus 70% of the real data scored 0.496 overall bounding-box mAP against 0.487 for the model trained on 100% of the real data. With 50% of the real data, synthetic pretraining recovered most of the loss (0.482, against 0.447 for 50% real alone) without quite matching the full baseline.

Some working guidance from these results:

  1. Keep some real data. Every row in the table above with a strong final number used real images in training. A few hundred well-labeled real images go a long way after synthetic pretraining.
  2. Expect diminishing returns. Tremblay et al. found performance saturated after about 10,000 domain-randomized images with pretrained weights and about 50,000 without, and Prakash et al. saw structured randomization saturate around 10,000.
  3. Target the classes you are short on. Applied Intuition's synthetic set raised cyclists from 0.3% of instances in nuImages to 27.4%, and the gains landed mostly on cyclists. Synthetic data does the most work on rare classes and rare conditions.
  4. Validate early stopping on real data. Gaidon et al. found that choosing when to stop the fine-tuning stage on a validation set was critical to avoid overfitting the small real training set.

When Synthetic Data Does Not Help

The published failures point to five common causes.

  • The class is missing or looks wrong. Kiefer et al. traced low synthetic-only scores to classes absent from the synthetic set (life jackets, tricycles) and to default object appearances that differed from the real ones.
  • The content distribution is off. Tremblay et al. found their domain-randomized model lost precision at high recall and attributed it to ignoring scene context, such as how parked cars line up. Prakash et al. added that structure back and more than doubled AP on the hard KITTI split (52.2 vs 24.0).
  • The simulator lacks variety. Nowruzi et al. found CARLA-derived data consistently weaker than the other synthetic sets, and found the game-derived Playing for Benchmarks data trained better detectors than the more photorealistic Synscapes. Diversity outweighed photorealism.
  • Real data is already plentiful and the task is easy. With 6,000 real KITTI images the synthetic gain was 0.2 AP, and on the cattle dataset synthetic pretraining slightly lowered YOLOv5's score.
  • The synthetic set is too small to matter. Ten thousand game frames did not beat three thousand real ones in Driving in the Matrix.

A useful habit is to report both numbers every time: the synthetic-trained model on held-out real data, and the real-trained model on the same split. The difference tells you how much work the synthetic data still has to do.

Where Stardust Fits

Stardust is Bifrost's synthetic data platform. It generates labeled RGB, IR, depth and segmentation data from controllable 3D scenes, with sensor profiles, weather and lighting set per scenario. Our published case study on bootstrapping maritime detection with synthetic data in 2 days is a small example of the targeted-class pattern above. One engineer retrained a pretrained YOLOv11n on 2,000 synthetic images of sailboats and buoys, reached just over 70% F1 on a real sequence from the public MODS dataset, found the model missed horizontal oblong buoys, added 500 synthetic images of them, and reached 82% F1 on the second run, all on one desktop GPU.

That result used synthetic training data only, on a narrow two-class task, which is where the literature says synthetic-only training is most likely to hold up. For broader class sets, the evidence above points to synthetic pretraining plus a real fine-tuning set. If you are choosing real datasets to pair with synthetic data, our list of maritime datasets for perception scores the public options on ontology, diversity and label quality.

Sources

Frequently Asked Questions

Can you train an object detector on synthetic data only?

Sometimes. Hinterstoisser et al. trained a Faster R-CNN on purely synthetic renders of 64 retail objects and beat a model trained on 1,158 hand-labeled real images (0.67 vs 0.54 mAP). For harder open-world scenes the gap is usually large, for example 10.5 vs 55.8 mAP@50 on the SeaDronesSee maritime benchmark, so most teams use synthetic data for pretraining and finish on real images.

What is the best ratio of synthetic to real data?

No single ratio works everywhere. Published studies have tested 0.5 to 1 and 1 to 1 synthetic to real (Applied Intuition), 6 real to 4 synthetic per batch (SYNTHIA), and 90 to 97.5 percent synthetic (Nowruzi et al.). The more consistent finding is about order and scarcity, since pretraining on synthetic then fine-tuning on real tends to beat mixing, and the gain is largest when real data is scarce.

Is it better to mix synthetic and real data or to pretrain and fine-tune?

In the studies that compared both, pretraining on synthetic data and fine-tuning on real data came out ahead. Nowruzi et al. found fine-tuning gave better results than mixed training with SSD-MobileNet, Kiefer et al. found it superior to combined training for UAV detection, and Applied Intuition's fine-tuned model scored 0.386 cyclist AP against 0.376 to 0.378 for mixed training.

How many synthetic images do you need for object detection?

For single-class car detection, published saturation points sit around 10,000 images for domain randomization with pretrained weights and for structured domain randomization, and about 50,000 for plain domain randomization without pretraining. In Driving in the Matrix, 10,000 game-engine images scored below 2,975 real images, and it took 50,000 to pull ahead. Narrow tasks need far less; Bifrost's maritime case used 2,500 images for two classes.

Does photorealism matter for synthetic training data?

Less than diversity and content, according to several studies. Tremblay et al. matched a photorealistic KITTI replica with deliberately unrealistic domain-randomized images (78.1 vs 79.7 AP), and Nowruzi et al. found the game-derived Playing for Benchmarks data trained better detectors than the more photorealistic Synscapes. Sensor effects such as noise, blur and exposure tend to matter more than rendering quality alone.

Explore Stardust More Posts