Synthetic Data For Defense ISR: EO/IR Training Data For Perception Teams
Synthetic data gives ISR perception teams labeled EO/IR and SAR imagery of rare targets they cannot collect or share. Train on it with some real data, then test on real.
The Short Answer
Synthetic data for defense ISR is rendered imagery of targets, backgrounds and sensors, labeled automatically, that perception teams use to train and test detectors where real labeled data is rare, restricted or too costly to collect. It matters most for electro-optical and infrared (EO/IR) and SAR tasks such as vessel detection, vehicle detection in overhead imagery, counter-UAS and automatic target recognition, because the targets that matter are the ones least present in public data. The evidence favors mixing it with real imagery. In RarePlanes, 50,000 synthetic satellite images with about 630,000 aircraft annotations were released next to 14,700 hand-annotated real aircraft, and the authors report that models trained on the synthetic set and fine-tuned on about 10% of the real training data reached 91 to 96% of the mAP of models trained on all of the real data.
Buyers are asking for it directly. In June 2026, DefenseScoop reported that SOCOM was seeking a self-service synthetic data platform that produces EO/IR training imagery for its unmanned systems autonomy program. As of October 2026, the pages that rank for this question are vendor product pages and whitepapers, most of them short, without public datasets, evidence or a FAQ. This post covers why ISR perception lacks labeled data, what synthetic data has to model for each task and sensor, which public datasets to use alongside it, the deployment questions that decide whether a tool can be used at all, and how the vendors compare.
Why ISR Perception Lacks Labeled Data
The targets are rare. Commercial and academic datasets are full of cars, buildings and cargo ships. A specific class of patrol craft, a mobile air defense system, or a semi-submersible appears in few public images, often at a handful of aspect angles, and seldom in the weather and lighting where a model will be used. Even xView, one of the largest public overhead datasets with over 1 million objects across 60 classes, was built around general object categories rather than any program's target list.
The real data is controlled. The imagery a program collects of the targets it cares about is often classified or controlled unclassified information. It cannot be sent to a commercial labeling vendor, moved to a public cloud, or shared with a subcontractor without approvals, so the labeled pool stays small and slow to grow.
Sensors differ. A detector trained on 0.3 m WorldView-3 imagery does not transfer to an MWIR gimbal camera on a UAV, and neither transfers to SAR. Every band (visible, SWIR, MWIR, LWIR, SAR), resolution, look angle and platform altitude needs its own training data, and many programs fly sensors that have no public dataset at all.
Adversaries vary their appearance. Camouflage, decoys, paint schemes, netting and new equipment change what a target looks like faster than a collection campaign can capture it. A model needs examples of variants that have not been photographed yet.
Small targets are hard to label. A 4.5 m car in 0.3 m satellite imagery is about 15 pixels long, and the Drone-vs-Bird challenge organizers report that most annotated drones in their training videos are smaller than 32 by 32 pixels, many under 16 by 16, in frames up to 3840 by 2160. Human labels at that size are inconsistent, and in thermal imagery they are worse. Rendered data has exact boxes and masks at any size, which is one of its practical advantages.
The SOCOM solicitation described by DefenseScoop puts the problem in similar terms. It says CV models need large volumes of labeled, operationally relevant EO/IR imagery, and manual collection is constrained by cost, access and the inability to capture rare events and contested environments.
ISR Tasks, Sensors, And What Synthetic Data Has To Model
Each ISR task pairs with a small set of sensors, and each pairing sets the things a synthetic pipeline must get right. The table lists the common ones.
| ISR task | Typical sensors | What synthetic data has to model |
|---|---|---|
| Vessel detection and classification | EO and SWIR from UAVs and shore; MWIR/LWIR at night; SAR from satellites | Sea state and wakes, sun glint, haze near the horizon, vessel classes by type and aspect, small craft at long range, thermal signature of engines and exhausts, SAR backscatter and speckle |
| Vehicle detection in overhead imagery | Satellite EO at 0.3 to 0.5 m; aerial EO; SAR | Ground sample distance and off-nadir angle, sun elevation and shadows, cloud cover, parking and convoy layouts, camouflage and netting, partial occlusion by trees and buildings |
| UAS detection (counter-UAS) | EO, MWIR/LWIR, sometimes SWIR, plus radar | Targets of a few pixels against sky and ground clutter, birds as distractors, airframe variety (quadcopter, fixed wing, VTOL), range and motion blur, thermal contrast against cold sky |
| Change detection and site monitoring | Revisit satellite EO; SAR | Paired images of the same site at different times, seasonal and lighting change that should be ignored, new construction or equipment that should be flagged, registration error |
| Tracking (ground, maritime, air) | Full-motion video in EO and MWIR/LWIR | Consistent object identities across frames, occlusion and re-acquisition, platform motion and gimbal jitter, frame-to-frame sensor noise |
| Automatic target recognition | MWIR/LWIR, EO, SAR | Fine-grained classes at many aspect angles and ranges, diurnal thermal cycles, target articulation (turrets, hatches), decoys and camouflaged variants |
Two things cut across every row. The first is the sensor model: spectral band, optics, noise, gain control, compression and any overlay the real system burns into its video. The second is metadata. Range, aspect, time of day and weather logged for every frame let a team find which conditions a model fails in, and that is where much of the value of synthetic data lies after the first model. Our post on synthetic thermal infrared training data covers IR sensor modeling in more detail, and Synthetic Data For Drone Detection covers the counter-UAS row.
Public Defense-Relevant Datasets
Synthetic data works best next to real data, for mixing and for testing. These are the public datasets most often used for ISR-style tasks, with figures from their papers and project pages.
| Dataset | Sensor | Contents | Useful for | Terms and gaps |
|---|---|---|---|---|
| xView (DIUx and NGA, 2018) | WorldView-3 EO, 0.3 m | Over 1 million objects, 60 classes, more than 1,400 km² | Overhead vehicle, vessel, aircraft and building detection | CC BY-NC-SA 4.0; general object classes, not a target list; single sensor |
| DOTA (v2.0) | Google Earth, GF-2, JL-1 and aerial imagery | 1,793,658 instances in 18 categories across 11,268 images, oriented boxes | Oriented detection of ships, planes and vehicles | Academic use only; commercial use prohibited |
| FAIR1M | High-resolution satellite EO | More than 1 million instances in more than 15,000 images, 5 categories and 37 sub-categories | Fine-grained classification of aircraft, ships and vehicles | Check terms on the dataset page; EO only |
| xBD (xView2) | Satellite EO, before and after events | 850,736 building annotations across 45,362 km² with damage labels | Change detection and damage assessment | Natural disasters, not military activity |
| SpaceNet 6 | SAR and optical, under 1 m GSD | Over 48,000 building footprints across 120 km² of Rotterdam | SAR segmentation and SAR-optical fusion | CC BY-SA 4.0 on AWS; one city, buildings only |
| xView3-SAR | Sentinel-1 SAR | 991 scenes and 243,018 verified maritime objects over 43.2 million km² | Vessel detection, including dark vessels without AIS | Download requires an account; Sentinel-1 imagery here has about 20 m resolution, and labels cover vessel versus fixed structure, fishing activity and vessel length rather than fine-grained vessel types |
| MSTAR (DARPA and AFRL) | X-band SAR, 1 ft resolution | Ten vehicle classes in the standard set, imaged through 360° of aspect | SAR automatic target recognition | Strong background correlation between training and test images, a flaw a 2020 survey calls out |
| DSIAC ATR database | MWIR and visible video | People, foreign military vehicles and civilian vehicles collected by the U.S. Army Night Vision and Electronic Sensors Directorate at 1,000 to 5,000 m, MWIR also at night | EO/IR automatic target recognition at range | Distribution A, requested through DSIAC (206 GB); highly correlated frames make naive splits too easy (Long-range thermal study) |
| RarePlanes | WorldView-3 EO plus synthetic | 14,700 real aircraft in 253 scenes; 50,000 synthetic images with about 630,000 annotations | Aircraft detection and fine-grained attributes | One of the few public sets that pairs real and synthetic data for overhead imagery |
None of these covers a program's exact sensor, targets and operating area. Their realistic role is pretraining, benchmarking and a first real test set while a program builds its own. The SpaceNet 6 paper gives a sense of how much prior data helps in a hard modality. Segmentation models pretrained on optical imagery and then trained on SAR reached an F1 of 0.21, compared with 0.135 for models trained on SAR alone.
What The Evidence Says About Synthetic Data For These Tasks
Results specific to military targets are mostly unpublished, so the public evidence comes from neighboring tasks. RarePlanes is the clearest overhead example, with real and synthetic aircraft that share 10 fine-grained attributes such as wingspan and wing shape. In counter-UAS work, the published results we collected in Synthetic Data For Drone Detection show mixed real and synthetic training beating real-only training in every hybrid comparison, while synthetic-only models fall well behind on hard test sets. Our broader review, Does Synthetic Data Work For Object Detection?, reaches the same conclusion across domains, which is to pretrain on synthetic data, fine-tune on a smaller real set, and measure on held-out real imagery.
The practical takeaway for an ISR team is to start with a real test set from the operational sensor, even a small one, before generating any synthetic data. Without it there is no way to tell whether the synthetic data is closing the gap or adding a new one. How To Close The Sim-To-Real Gap For Perception covers how to set that test up.
Deployment, Export Control, And Data Handling
For defense programs, where a tool runs and what it touches often decide the choice before image quality does. These are general considerations, not legal advice, and a program's security and export compliance staff have the final say.
On-prem and air-gapped operation. Programs working with controlled or classified data usually need the generator, the asset library and the training pipeline inside their own enclave. A cloud-only tool can still be useful for unclassified pretraining data, but any step that touches real program imagery has to run where that imagery is allowed to live. Ask vendors whether they support on-prem or air-gapped installs, how licensing works offline, and how updates and new assets reach a disconnected network.
Export control. In the United States, defense articles and technical data fall under ITAR, administered by the State Department, and dual-use items fall under the EAR, administered by the Commerce Department. Detailed 3D models of military equipment, sensor performance models and the trained models themselves can carry controls, separate from any imagery. Ask how a vendor's asset library is classified for export, and treat your own custom assets and trained weights as items that need a determination.
Data handling. Synthetic imagery does not contain collected intelligence, which is one reason teams use it, but metadata, target lists and scenario descriptions can still reveal what a program is looking for. Keep scenario definitions and any real backplates under the same handling rules as the program data they describe.
Self-service versus managed service. SOCOM's request, as reported, asked for a platform personnel can operate organically without vendor assistance and that plugs into an existing MLOps pipeline. Many defense teams want the same thing, because a managed service means sending requirements, and sometimes data, outside the program for every iteration.
Vendors And Tools For Defense Synthetic Data
Several companies sell synthetic data into defense, and they differ mainly in sensor coverage and delivery model. The descriptions below come from each company's own pages as of October 2026.
| Vendor or tool | What it offers for ISR | Notes |
|---|---|---|
| Anyverse DEFENCE | Camera, lidar and radar simulation, with thermal and NIR; UAV and ISR use cases | Described as a cloud-based platform |
| Duality AI Falcon | Digital twin simulation with a configurable sensor library; counter-UAS listed as a capability area | 2025 contract from the U.S. Army XM30 program office for counter-drone training data (announcement) |
| Rendered.ai | SAR, infrared, multispectral, hyperspectral, EO and RGB | Platform subscription or synthetic data as a managed service |
| L3Harris | Physics-based remote sensing simulation for panchromatic, multispectral, thermal, hyperspectral and SAR | Cites a 40-year legacy of radiometrically correct remote-sensing simulation |
| ThermoAnalytics MuSES | Physics-based infrared signature prediction for vehicles, aircraft, buildings, satellites and unmanned systems | ThermoAnalytics markets it for generating synthetic images for ATR training, and Xitadel describes MuSES-based EO/IR training data |
| NVIDIA Isaac Sim and Omniverse Replicator | Free, self-hosted rendering with programmatic labels | Best when a team has graphics engineers and wants full control |
| Bifrost Stardust | Maritime, aerial and geospatial scenes in RGB, SWIR, MWIR and thermal, with satellite emulation | Described below |
If your sensor is SAR today, Rendered.ai and L3Harris list it as available now. If you need physics-accurate thermal signatures of specific vehicles, MuSES-based pipelines are built for that. If you have the engineering staff and want everything on your own hardware, NVIDIA's tools are the strongest free option. Synthetic Data Platforms Compared covers these vendors and others in more depth.
Where Stardust Fits
Stardust is Bifrost's synthetic data platform for perception and autonomy. The U.S. Air Force, Shield AI and ST Engineering appear on its aerial page under "Trusted by airborne ISR and autonomy programs," and Saronic, Havoc, Seadronix, ClassNK and ST Engineering appear on its maritime page.
For ISR, Stardust's published use cases include maritime vessel detection from UAVs (including narcotic semi-submersibles), terrestrial detection of ground vehicles up to military vehicles, counter-UAS classification by UAV type and role, port defense patrols with virtual USVs, and maritime threat classification that distinguishes civilian from military vessels. Its maritime asset library includes aircraft carriers, destroyers, frigates, submarines, patrol vessels and minesweepers alongside commercial ships. Sensors include RGB, stereo, SWIR, MWIR and thermal, with depth and segmentation output and radar in development. On the geospatial side, it emulates WorldView-3, Legion 06 and SkySat, composites synthetic objects onto real overhead backplates, and is developing coherent SAR with NTT Data. Every frame comes labeled with boxes, pixel-level masks and scenario metadata.
Stardust's site does not publish deployment options, so ask about on-prem and air-gapped requirements directly. For SAR work today, or for an entirely self-hosted toolchain, the alternatives above are the better fit.
Sources
- Harper, SOCOM seeks 'self-service' synthetic data generation platform to boost drones' computer vision, DefenseScoop, June 9, 2026
- Lam et al., xView: Objects in Context in Overhead Imagery, 2018
- Ding et al., Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges, 2021, and the DOTA dataset page
- Sun et al., FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery, 2021
- Gupta et al., xBD: A Dataset for Assessing Building Damage from Satellite Imagery, 2019
- Shermeyer et al., SpaceNet 6: Multi-Sensor All Weather Mapping Dataset, CVPRW 2020
- Paolo et al., xView3-SAR: Detecting Dark Fishing Activity Using Synthetic Aperture Radar Imagery, NeurIPS 2022 Datasets and Benchmarks
- Kechagias-Stamatis and Aouf, Automatic Target Recognition on Synthetic Aperture Radar Imagery: A Survey, 2020
- Shermeyer et al., RarePlanes: Synthetic Data Takes Flight, WACV 2021
- DSIAC, ATR Algorithm Development Image Database
- Coluccia et al., Drone vs. Bird Detection: Deep Learning Algorithms and Results from a Grand Challenge, Sensors, 2021
Frequently Asked Questions
What is synthetic data used for in defense AI?
Mostly for training and testing perception models where real labeled imagery is scarce, restricted or expensive. Typical tasks are vehicle and vessel detection in overhead imagery, counter-UAS detection, automatic target recognition in EO/IR video, and change detection. Synthetic data is rendered from 3D scenes, so every frame comes with exact labels and metadata.
Can a defense perception model be trained on synthetic data alone?
Sometimes, but most published results use synthetic data together with a smaller real set. The RarePlanes authors found that aircraft detectors trained on synthetic overhead imagery and fine-tuned on about 10% of the real data reached 91 to 96% of the mAP of real-only models, and counter-UAS studies show mixed training beating real-only training. Always measure on held-out real imagery from your own sensor.
Which public datasets exist for defense ISR computer vision?
The most used are xView (over 1 million objects in 60 classes from WorldView-3), DOTA, FAIR1M, xBD for damage assessment, SpaceNet 6 for SAR and optical, xView3-SAR for vessel detection in Sentinel-1 imagery, MSTAR for SAR vehicle recognition, and the DSIAC ATR database for MWIR and visible vehicle imagery. Several are licensed for non-commercial or academic use only.
Does synthetic data avoid export control and classification issues?
It can reduce them, because rendered imagery does not contain collected intelligence, but it does not remove them. 3D models of military equipment, sensor models and trained weights can themselves be controlled, and programs often require on-prem or air-gapped tools. Check with your program's security and export compliance staff before moving any data or models.
What does SOCOM want from a synthetic data platform?
According to DefenseScoop in June 2026, SOCOM's program executive office for SOF digital applications was seeking a self-service synthetic data generation platform that produces EO/IR training imagery for the Unmanned Systems Autonomy and Interoperability program. The command wanted personnel to run it without vendor assistance and to integrate it into the program's existing MLOps pipeline.