Synthetic Data For Drone Detection: What The Research Shows
Synthetic data improves drone detectors most when mixed with real images. Published hybrids beat real-only training, while synthetic-only models trail on hard test sets.
The Short Answer
Synthetic data works for drone detection mainly as a supplement to real imagery, not a replacement for it. In SynDroneVision (WACV 2025), training YOLOv9e on synthetic Unreal Engine images together with the real DUT Anti-UAV set raised mAP@0.5 on the DUT Anti-UAV test set from 0.915 to 0.955, while the synthetic-only model scored 0.733. In thermal imagery, a 2026 LWIR study found that synthetic pre-training followed by 100 real frames beat those 100 real frames alone for every detector it tested, but synthetic-only models generalized poorly. Pure synthetic training can come within a point of real data on easy test sets (0.970 against 0.978 AP50 on MAV-VID) and fall far behind on harder ones (0.678 against 0.977 on Anti-UAV RGB).
As of October 2026, the pages that rank for this question are individual papers, sensor vendor guides and data-labeling vendors, and each covers one piece of the question. This post puts the published numbers side by side, lists the public datasets you would mix with synthetic data, and explains which variables in a synthetic pipeline the evidence says matter.
What Published Studies Report
The table collects results where the same detector was trained on real data, on synthetic data, and in most cases on both, then tested on real imagery. All numbers come from the papers' own tables. Metrics differ between studies, so compare within a row, not across rows.
| Study | Sensor and setup | Detector, metric | Real only | Synthetic only | Synthetic + real |
|---|---|---|---|---|---|
| SynDroneVision (Lenhard et al., WACV 2025) | RGB surveillance, Unreal Engine renders; tested on DUT Anti-UAV | YOLOv9e, mAP@0.5 | 0.915 | 0.733 | 0.955 (joint training) |
| SynDroneVision, out-of-distribution test | RGB; tested on UAV-Eagle | YOLOv9e, mAP@0.5 | 0.819 | 0.810 | 0.879 |
| Wisniewski et al., 2024 | RGB, 5,000 Blender renders with structured domain randomization; tested on MAV-VID | Faster R-CNN, AP50 | 0.978 | 0.970 | not tested |
| Wisniewski et al., 2024 | Same model; tested on Drone-vs-Bird | Faster R-CNN, AP50 | 0.632 | 0.498 | not tested |
| Wisniewski et al., 2024 | Same model; tested on Anti-UAV RGB | Faster R-CNN, AP50 | 0.977 | 0.678 | not tested |
| Sim2Air (Barisic et al., RA-L 2022) | RGB air-to-air, 52,500 Blender renders with texture randomization; tested on UAV-Eagle | YOLOv4-tiny, mAP@0.5 | 90.38% | 73.80% | 95.24% (fine-tuned) |
| Sim2Air | Same; tested on a harder small-object set (T2) | YOLOv4-tiny, mAP@0.5 | 65.73% | 68.83% | 71.54% (fine-tuned) |
| Liiv et al., 2026 | LWIR thermal, ground-to-air, 20,000 synthetic IR-like scenes plus 100 real frames | RF-DETR-L, mAP@50:95 | 0.698 | 0.463 | 0.713 (fine-tuned) |
| Liiv et al., 2026 | Same | Faster R-CNN, mAP@50:95 | 0.012 | not tabulated | 0.352 (fine-tuned) |
In the Wisniewski rows, the real-only baselines are Faster R-CNN models from the Isaac-Medina et al. benchmark, each trained on the dataset it was tested on, while the synthetic model never saw any of the three test sets. Sim2Air's real-only baseline is YOLOv4-tiny trained on R-UAV, a set of about 11,700 real images.
One more study belongs here even though it does not fit the table. Dieter et al. (Electronics, 2023), from DLR, built a 17,412-image synthetic set in Unreal Engine with Microsoft AirSim and trained YOLOv5l on its training split. On four real test sets recorded with their own ground cameras, the synthetic-only model scored between 0.432 and 0.656 mAP@0.5, and on a fifth it scored 0.136. Fine-tuning that model with real-data shares of 5 percent or less raised mAP by 40 to 50 percentage points on several of the real sets.
What The Numbers Say Together
Mixing beats real-only in every hybrid result above. The gains range from 1.5 points (RF-DETR-L in thermal) to 6 points (YOLOv9e on UAV-Eagle) when the real baseline is already strong, and they are much larger when it is weak. The Faster R-CNN thermal baseline trained on 100 real frames barely worked at 0.012, and synthetic pre-training lifted it to 0.352.
Synthetic-only results depend on how hard the real test set is. The same Wisniewski model came within 0.8 points of real data on MAV-VID, trailed by 13.4 points on Drone-vs-Bird, and trailed by 29.9 points on Anti-UAV. The authors put the Anti-UAV drop down to on-screen overlays in that footage (characters, numbers and crosshairs that caused false positives) and to night sequences their renders never included.
Synthetic data can beat real data from a different domain. On Sim2Air's T2 set, the synthetic-only detector (68.83%) outperformed the detector trained on real R-UAV images (65.73%), because T2's small, custom drones looked unlike the commercial drones in R-UAV. The same pattern shows up in other object detection work, which we cover in Does Synthetic Data Work For Object Detection?.
Public Drone Detection Datasets
Every result above relies on a real dataset for testing or mixing. These are the public ones most often used, with sizes and terms taken from the papers and repositories.
| Dataset | Modality | Size | Task | Access and license |
|---|---|---|---|---|
| Anti-UAV (Jiang et al., 2021) | Paired RGB and thermal IR video | More than 300 video pairs, over 580,000 manually annotated boxes | Tracking | Repository under MIT License; downloads via Google Drive and Baidu |
| Anti-UAV410 (Huang et al., TPAMI 2023) | Thermal IR video | 410 videos, over 438,000 manually annotated boxes | Tracking | Free download; no license stated in the repository |
| DUT Anti-UAV (Zhao et al., IEEE T-ITS 2022) | RGB | 10,000 detection images (5,200 train, 2,600 val, 2,200 test) and 20 tracking videos; 35 drone models | Detection and tracking | Repository under Apache 2.0 |
| Drone-vs-Bird (Coluccia et al.) | RGB video | 77 annotated training videos averaging 1,384 frames, plus 14 unannotated test videos (2021 edition); resolution from 300×168 to 4K | Detection with birds present but not labeled | Free on request after signing a data usage agreement; research use |
| Halmstad drone dataset (Svanström et al., 2021) | Thermal IR, visible and audio | 650 videos (365 IR, 285 visible) totaling 203,328 annotated frames, plus 90 audio clips | Detection and classification of drone, bird, airplane and helicopter | CC0 1.0 |
| Det-Fly (Zheng et al., 2021) | RGB | More than 13,000 images of a target drone filmed from a DJI Mavic 2 | Air-to-air detection, front, top and bottom views | Repository under MIT License |
| SynDroneVision (Lenhard et al., 2025) | Synthetic RGB | 131,307 annotated images from 72 sequences at 2560×1489, about 900 GB | Detection | CC BY 4.0 on Zenodo |
| S-UAV-T (Barisic et al., 2022) | Synthetic RGB | 52,500 images, 10 UAV models | Air-to-air detection | Free download; no license stated in the repository |
Licenses on repositories describe the repository. Some dataset pages say nothing about commercial use, so read the terms before training a product model. If your sensor is thermal, the Synthetic Thermal Infrared Training Data post covers IR-specific sets and the routes to labeled IR imagery in more depth.
Why Drone Detection Is Hard
The targets are tiny. The Drone-vs-Bird organizers note that drones can be 10 to 20 pixels across in a full high-resolution frame. In DUT Anti-UAV the average drone covers 1.3 percent of the image area (an object scale of 0.013, per the SynDroneVision paper), and in Det-Fly nearly half the targets are smaller than 5 percent of the image. Teledyne FLIR's counter-UAS sensor guide puts detection at more than 10×10 pixels on target, recognition (quadcopter versus fixed wing) at about 20×10, and identification of a specific type at about 30×20.
Birds look like drones. At 15 pixels, a gull and a quadcopter share a silhouette, and both move against the sky. Drone-vs-Bird was built around this confusion, and it is one of the reasons the synthetic-only model above lost 13.4 points there. The Halmstad dataset labels birds, airplanes and helicopters explicitly, which makes it useful for measuring false alarms.
Sky and ground clutter. Clouds, sun glare and fast changes in exposure are common in real footage, and backgrounds below the horizon are worse. In Dieter et al., almost 99.95 percent of the false positives on two real test sets traced back to a single background object (a traffic sign in one, a car in the other) that their simulation did not contain, and textured backgrounds such as trees were the main cause of missed drones.
Thermal and RGB fail differently. FLIR's guide notes that IR cameras can detect drone motion in a pixel cluster as small as 2×2 against a cold sky, while EO cameras provide 2x to 8x more pixels on target. LWIR imagery has little texture, weak thermal contrast and sensor noise, and the Liiv et al. analysis found that real frames carried histogram artifacts from automatic gain control that their synthetic frames lacked. A model or a synthetic pipeline tuned for one band does not transfer to the other.
Labels on small targets are noisy. Wisniewski et al. show ground-truth boxes in public datasets that are less accurate than the model's prediction, and FLIR warns that boxes at extreme range may miss the target entirely. That noise lowers measured mAP and limits what a model can learn at low signal-to-noise ratios. Rendered data has exact boxes, which is one of its practical advantages for this problem.
What A Synthetic Pipeline Should Vary
The studies above ran ablations on several of these variables, which makes it possible to say which ones changed results.
- Range and target size. Wisniewski et al. randomized camera position within bounds from 20 m to 320 m. Bounds of 20 to 80 m gave similar results, Drone-vs-Bird performance peaked between 80 and 160 m (more small drones), and 320 m hurt every test set because the model could no longer learn the drone's shape. Match the size distribution to the ranges you need to cover, and include targets below the 10×10 pixel detection threshold.
- Drone models and textures. Sim2Air's texture randomization pushed the detector toward shape, raising mAP on an unfamiliar real set by 17 points for YOLO and 20 points for Faster R-CNN. A wide library of airframes (quadcopters, fixed wing, VTOL) matters for the same reason.
- Backgrounds and clutter. Model the backgrounds the camera will see, including what sits below the horizon. Dieter et al. traced most false positives to background objects their simulation left out, and Wisniewski et al., whose backgrounds were mostly sky HDRIs, did worst on the cluttered Drone-vs-Bird scenes.
- Distractors. Results are mixed. Synthetic birds lowered Drone-vs-Bird AP50 from 0.497 to 0.403 in Wisniewski et al., while real RGB bird images used as hard negatives improved pre-training for every architecture in Liiv et al. Treat distractors as an experiment and measure them on a held-out real set.
- Lighting, weather and time of day. Sim2Air's texture-randomized detector held up better under challenging illumination, by 16 points on T2 and 4 points on UAV-Eagle. The SynDroneVision authors have since added SynDroneVision-Weather, 55,187 images with rain, snow and fog at several severities across three seasons. Night matters too, since Anti-UAV includes night sequences.
- The sensor itself. Adding Gaussian noise gave a small gain and JPEG compression none in Wisniewski et al. In thermal, Liiv et al. found synthetic drones rendered as uniform silhouettes and lacked the gain-control artifacts of real frames. If the real camera burns an overlay into its video, as Anti-UAV's does, either render it or crop it out of both.
- Sequences, not just frames. Counter-UAS systems track, so synthetic data with consistent object identities across frames supports trackers as well as detectors.
Whatever the pipeline, keep a real held-out test set from your own sensor and judge every change against it. How To Close The Sim-To-Real Gap For Perception covers how to measure that gap and the methods for narrowing it.
Where Stardust Fits
Stardust is Bifrost's synthetic data platform, and counter-UAS is one of the use cases on its aerial page. Scenarios can include several drones in overlapping airspace, birds as the distractor that drives false alarms, and different UAV types classified by role and payload configuration. The asset library covers quadcopters, fixed-wing UAVs, VTOL systems and tactical unmanned platforms, and altitude, camera perspective, distance and sensor configuration are set per scenario.
Sensors include EO/RGB, IR, thermal and SWIR, with depth and segmentation output and radar in development. Labels (bounding boxes, pixel-level segmentation and scenario metadata) come from the 3D scene, and sequences keep consistent object identities across frames for tracking. Stardust is one option among several. If you need a free, self-hosted toolkit and have the engineering time, the pipelines behind SynDroneVision (Unreal Engine with Colosseum) and Sim2Air (Blender) are described in detail in their papers. Synthetic Data Platforms Compared covers the commercial alternatives.
Sources
- Lenhard, Weinmann, Franke, and Koch, SynDroneVision: A Synthetic Dataset for Image-Based Drone Detection, WACV 2025
- Lenhard, Weinmann, and Koch, Beyond Clear Skies: Synthetic Seasonal and Weather Variations for Real-World Drone Detection, 2026
- Wisniewski, Rana, Petrunin, Holt, and Harman, Drone Detection using Deep Neural Networks Trained on Pure Synthetic Data, 2024
- Barisic, Petric, and Bogdan, Sim2Air: Synthetic Aerial Dataset for UAV Monitoring, IEEE RA-L 2022
- Dieter, Weinmann, Jäger, and Brucherseifer, Quantifying the Simulation–Reality Gap for Deep Learning-Based Drone Detection, Electronics 2023
- Liiv et al., Training with Synthetic Data for Drone Detection in Thermal Imagery, 2026
- Jiang et al., Anti-UAV: A Large Multi-Modal Benchmark for UAV Tracking, 2021
- Zhao, Zhang, Li, and Wang, Vision-based Anti-UAV Detection and Tracking, IEEE T-ITS 2022
- Coluccia et al., Drone vs. Bird Detection: Deep Learning Algorithms and Results from a Grand Challenge, Sensors 2021
- Teledyne FLIR, Thermal Infrared Sensor Design Considerations for Counter-UAS Defense
Frequently Asked Questions
Can you train a drone detector on synthetic data alone?
You can, and on simple test sets it can come close to real data. A Faster R-CNN trained only on 5,000 Blender renders reached 0.970 AP50 on MAV-VID against 0.978 for a model trained on MAV-VID itself. On harder sets the gap is large, for example 0.678 against 0.977 on Anti-UAV RGB, so most teams use synthetic data together with some real imagery.
What is the best public dataset for drone detection?
It depends on the sensor and the task. DUT Anti-UAV (10,000 RGB detection images) is a common detection benchmark, Anti-UAV and Anti-UAV410 cover thermal tracking, Drone-vs-Bird is the standard test for bird confusion, and the Halmstad dataset has IR, visible and audio with birds, airplanes and helicopters labeled. None of them covers every range, weather and drone type you are likely to deploy against.
Does adding birds to drone training data reduce false alarms?
The evidence is mixed. In one 2024 study, adding synthetic birds as distractors lowered AP50 on the Drone-vs-Bird test from 0.497 to 0.403. In a 2026 thermal study, adding real RGB bird images to synthetic pre-training improved mAP for every architecture tested before fine-tuning. Test bird distractors against your own validation set rather than assuming they help.
How much real data do you need alongside synthetic drone data?
Less than you might expect, but not zero. In a 2023 DLR study, fine-tuning a synthetic-trained YOLOv5 model with real-data shares of 5 percent or less raised mAP by 40 to 50 percentage points on several real test sets. In a 2026 LWIR study, 100 real thermal frames on top of synthetic pre-training was enough to beat 100 real frames alone.
Is thermal or RGB better for counter-UAS detection?
Each has a role. Teledyne FLIR notes that IR cameras can pick up drone motion in a pixel cluster as small as 2x2 against a cold sky, while EO cameras give 2x to 8x more pixels on target. Thermal works at night and in haze, RGB gives the detail needed for classification, and many systems fuse both.