Synthetic Thermal Infrared Training Data: How To Get Labeled IR Imagery
Labeled thermal IR training data comes from three routes: hand-labeling real frames, translating RGB images to thermal, or rendering scenes with physics-based simulation.
The Short Answer
There are three ways to get labeled thermal infrared training data. You can collect real thermal footage and label it by hand, which gives you the true sensor signal but costs the most per labeled frame. You can translate existing labeled RGB images into thermal-looking images with GANs or diffusion models, which is cheap but only as good as the model's guess about temperatures it never saw. Or you can render scenes with physics-based simulation, where materials carry emissivity and temperature, and every frame arrives with exact labels, depth and segmentation. Most teams that ship thermal detectors end up mixing routes: a public dataset such as the 26,442-frame Teledyne FLIR ADAS set for real signal, synthetic data for coverage, and a held-out set of their own real frames to check that the mix transfers.
Route 1: Collect And Label Real Thermal Imagery
Real thermal data is the reference everything else is measured against. It contains the actual noise, optics blur, non-uniformity and gain behavior of the camera you will deploy, which no synthetic method fully reproduces.
The costs are the familiar ones, made worse by the modality. Thermal cameras usually cost more than RGB cameras, the scenes you most need (night, fog, smoke, rare objects at long range) are the hardest to stage, and small, low-contrast thermal targets are hard to label consistently. Teledyne FLIR's counter-UAS sensor guide warns that manual annotation of small targets introduces label noise that can confuse the detector.
If you go this route, record raw radiometric frames as well as the 8-bit display output. The FLIR ADAS dataset ships thermal images both as 14-bit TIFFs with no automatic gain control (AGC) and as 8-bit JPEGs with AGC applied, so you can train on either and compare. Most other public datasets only provide the processed 8-bit images, and that choice is baked into every model trained on them.
Route 2: Translate RGB Images Into Thermal
Image-to-image translation takes a labeled RGB dataset and converts each image into a thermal-looking image while keeping the labels. The early methods are general-purpose translation networks: pix2pix for paired data and CycleGAN for unpaired data. Later work specializes them for thermal, for example InfraGAN (Özkanoğlu and Ozer, Pattern Recognition Letters, 2022) and the edge-guided multi-domain model of Lee, Jeon, Cho, and Kim (ICRA 2023). More recent work moves to diffusion models: DiffV2IR combines a diffusion translator with vision-language understanding and introduces IR-500K, a dataset of 500,000 infrared images.
The Lee et al. paper shows how much the choice of method matters. They generated four synthetic thermal detection datasets from the same synthetic RGB source using four translators, trained a VFNet detector on each, and evaluated all four on the real FLIR-ADAS validation set.
| Translator used to make the training set | mAP | mAP50 |
|---|---|---|
| Edge-guided (Lee et al., proposed) | 0.239 | 0.408 |
| UNIT | 0.225 | 0.384 |
| MUNIT | 0.069 | 0.137 |
| CycleGAN | 0.007 | 0.010 |
Source: Lee et al., ICRA 2023, Table III. Detector trained only on translated images, tested on real FLIR-ADAS validation frames.
The same labeled source images produced a usable detector with one translator and a near-useless one with another. Even the best result, 0.239 mAP, is a starting point for fine-tuning rather than a deployable model.
There is also a physical limit. An RGB image records reflected visible light, and a thermal image records emitted long-wave radiation, which depends on surface temperature and emissivity. A parked car and a car that has been driving for an hour look the same in RGB and very different in LWIR. A translator has to guess which one it is looking at from the statistics of its training set, so it works best when your target scenes match its paired training data and worst for the out-of-distribution cases you wanted synthetic data for.
Route 3: Physics-Based Rendering
Physics-based rendering builds a 3D scene, assigns materials their thermal properties, and simulates what a given sensor would record. The IR signature of an object depends on its temperature, its emissivity, reflections of its surroundings and the sky, solar loading, atmospheric attenuation and the sensor's own optics and electronics. Engines such as RIT's DIRSIG, which simulates imagery from the visible through the thermal infrared, model this chain for remote sensing, and a 2024 survey by Upadhyay et al. covers both these model-based methods and newer deep learning approaches.
Rendering has two structural advantages. Labels are exact, because they come from the scene graph rather than from an annotator reading a blurry hot spot. And you control the variables that matter for thermal: time of day, ambient temperature, which engines are running, how warm the water is relative to the air, and the target's range and aspect. That makes it the practical way to cover thermal crossover (the hours when target and background reach the same apparent temperature and contrast disappears), which real collections capture only by luck.
The failure mode is the reverse of translation's. A renderer is physically consistent, but a detector can still tell its output apart from real frames if the sensor model is too clean. A 2026 study of synthetic-first drone detection in LWIR (Liiv et al.) measured this gap. Its 20,000 synthetic samples combine generative IR-like backgrounds with drones rendered in Blender, which the authors say is not a fully radiometric LWIR simulator. Real frames had higher histogram total variation, which the authors describe as consistent with AGC-like stretching of high-bit-depth data into 8 bits, and the synthetic pipeline did not reproduce it. Inside the target boxes, Sobel gradient variance was over four times higher in real detections than in synthetic ones. Synthetic pre-training without fine-tuning did not beat a baseline trained on 100 real LWIR images (Faster R-CNN was the exception), but synthetic pre-training followed by fine-tuning on those 100 images beat the real-only baseline for every detector family tested. All the detectors started from COCO weights, so the synthetic data served as a second pre-training stage.
Hybrid approaches render only the objects and composite them into real thermal backgrounds, such as Bongini et al. on FLIR ADAS and Barisic Kulas et al. (ECMR 2025), who added a drone class to the HIT-UAV aerial thermal dataset. These keep real sensor statistics in the background and control the targets, at the cost of lighting and thermal interactions between object and scene.
How The Three Routes Compare
| Real capture and labeling | RGB-to-thermal translation | Physics-based rendering | |
|---|---|---|---|
| Label quality | Limited by annotator skill on low-contrast, small targets | Inherited from the RGB labels; can drift if the translator moves or erases objects | Exact, from the scene graph, including depth and segmentation |
| Physical fidelity | Ground truth for the camera used | Plausible appearance, no real temperature or emissivity information | Consistent with the thermal model; depends on material data and sensor model |
| Cost per labeled frame | Highest (hardware, field time, annotation) | Lowest once a translator is trained | Moderate; mostly scene and asset setup, then cheap to scale |
| Rare and dangerous cases | Captured only if they happen during collection | Limited to scenes that exist in RGB | Can be composed on demand |
| Main failure mode | Coverage gaps and label noise | Hallucinated thermal states; large variance between methods | Sensor model too clean (missing AGC, noise, blur) |
| Best use | Validation set and fine-tuning data | Quick bootstrapping when good paired RGB data exists | Coverage of conditions, ranges and objects real data lacks |
Whatever mix you choose, keep a held-out set of real frames from your own sensor that no training route touches, because results on that set tell you whether the synthetic data helped. Our posts on whether synthetic data works for object detection and closing the sim-to-real gap for perception cover mixing and measurement in more detail.
Public Thermal Datasets Worth Starting With
| Dataset | Content | Size | Registration with RGB | Terms |
|---|---|---|---|---|
| Teledyne FLIR ADAS v2 | Driving scenes, 15 classes, thermal (Tau 2, 640×512) and RGB | 26,442 annotated frames, 520,000 boxes; 9,711 thermal and 9,233 RGB train/val images plus 7,498 video frames | Separate thermal and RGB cameras; an earlier release had enough misaligned pairs that Zhang et al. hand-filtered a 4,129 train / 1,013 test aligned subset | Free with registration |
| KAIST Multispectral Pedestrian | Driving scenes, day and night, person / people / cyclist | About 95,000 color-thermal pairs at 640×480, 103,128 annotations, 1,182 unique pedestrians | Aligned in hardware with a beam splitter | CC BY-NC-SA 4.0 by default; some bundled code is BSD 2-Clause |
| LLVIP | Mostly very dark scenes, pedestrians | 15,488 visible-infrared pairs (30,976 images) | Strictly aligned in time and space; raw unregistered data also released | Non-commercial use only |
| M3FD | Campus, coastal and urban scenes, 6 classes | 4,200 pairs, mostly 1024×768 | Registered via calibration and homography | Repository is GPL-3.0; no dataset-specific license stated |
| HIT-UAV | Aerial thermal from 60 to 130 m altitude, persons and vehicles | 2,898 images, 24,899 labeled objects (including a DontCare category) | Thermal only | CC BY 4.0 |
Licenses vary a lot, and two of the most cited datasets (KAIST and LLVIP) restrict the data to non-commercial use. Almost all of the data is also street-level driving or surveillance of people and cars, which leaves maritime, aerial, defense and industrial thermal applications with little public data to start from. For maritime LWIR, our roundup of maritime datasets lists what exists.
Sensor Details That Decide Whether Synthetic IR Transfers
LWIR And MWIR Are Different Training Domains
Long-wave infrared (roughly 8 to 14 micrometers) is where objects near room temperature emit most strongly; by Wien's displacement law, a 300 K surface peaks at about 9.7 micrometers. Most affordable thermal cameras are uncooled LWIR microbolometers. Mid-wave infrared (roughly 3 to 5 micrometers) usually needs a cooled detector, and Teledyne FLIR's counter-UAS guide describes cooled MWIR cameras with resolutions up to 1280×1024 and focal lengths beyond 1000 mm for long-range work. Hot sources such as exhaust plumes, and reflected sunlight, weigh more heavily in MWIR. Capture or render training data in the band of the sensor you deploy, and ask which band a dataset labeled only "IR" or "thermal" was captured in.
Registration With RGB And Depth
Fusion models need thermal, RGB and often depth to line up pixel for pixel, which is hard when the cameras do not share a lens. KAIST solved it in hardware with a beam splitter. M3FD calibrates the cameras and warps the infrared image with a homography, which is exact only for a single plane, so nearby objects with parallax still shift. FLIR ADAS uses two separate cameras, which is why researchers hand-filtered aligned subsets of its first release for fusion work. Rendered data avoids the problem, since every modality comes from the same virtual camera.
Noise, Non-Uniformity Correction And Gain
Uncooled microbolometers drift with temperature, so the pixels need periodic re-normalization against a uniform source, usually a shutter, to refresh the per-pixel non-uniformity correction (NUC). Teledyne FLIR calls this step flat field correction and triggers it by elapsed time and temperature change. Between corrections you see fixed-pattern noise and slow drift, and after AGC remaps 14-bit data into 8 bits, the histogram of each frame depends on whatever else is in the scene, so a hot object entering the frame changes how every other pixel is displayed. Synthetic thermal data that skips these effects looks cleaner than any real camera output, which is the gap Liiv et al. measured. At a minimum, add fixed-pattern noise, temporal noise, optics blur and a gain stage that mimics your camera's AGC, or train on raw radiometric values end to end if your deployment pipeline allows it.
Where Stardust Fits
Stardust is Bifrost's synthetic data platform, and it takes the rendering route. One scene is rendered across modalities in registration, with RGB, IR and depth from the same 3D scene and ground-truth labels generated automatically for every frame. The Thermal with Depth render output produces thermal imagery and a per-pixel depth map from a single render, pixel-aligned by construction, so there is no cross-sensor registration step. For maritime work, Stardust simulates RGB, stereo, SWIR, MWIR and thermal sensors matched to your hardware (maritime), and the aerial page covers EO, IR and thermal for UAV perception. Scenes can be varied across fog, rain, snow and night operations, sensors can be perturbed with noise, intrinsics and perspective shifts, and a neural rendering pass adds the glare, grain and noise of real sensors on top of the 3D render.
The same advice applies to our data as to anyone's. Synthetic thermal data earns its place by covering the conditions and objects your real data lacks, and it should be judged on a held-out set of real frames from your own camera.
Sources
- Teledyne FLIR, Free ADAS Thermal Dataset v2
- Hwang et al., KAIST Multispectral Pedestrian Dataset, CVPR 2015
- Jia et al., LLVIP: A Visible-infrared Paired Dataset for Low-light Vision, ICCV Workshops 2021
- Liu et al., TarDAL and the M3FD dataset, CVPR 2022
- Suo et al., HIT-UAV: A High-altitude Infrared Thermal Dataset, Scientific Data 2023
- Zhang et al., Multispectral Fusion for Object Detection with Cyclic Fuse-and-Refine Blocks, 2020
- Lee, Jeon, Cho, and Kim, Edge-guided Multi-domain RGB-to-TIR Image Translation for Training Vision Tasks with Challenging Labels, ICRA 2023
- Ran et al., DiffV2IR: Visible-to-Infrared Diffusion Model via Vision-Language Understanding, 2025
- Upadhyay et al., A Comprehensive Survey on Synthetic Infrared Image Synthesis, 2024
- Liiv et al., Training with Synthetic Data for Drone Detection in Thermal Imagery, 2026
- Bongini et al., Partially Fake It Till You Make It: Mixing Real and Fake Thermal Images, 2021
- Barisic Kulas, Jurasovic, and Bogdan, Unlocking Thermal Aerial Imaging: Synthetic Enhancement of UAV Datasets, ECMR 2025
- Rochester Institute of Technology, DIRSIG
Frequently Asked Questions
Can I train a thermal object detector on RGB images alone?
You can start from RGB-pretrained weights, and thermal studies routinely begin from COCO-pretrained detectors, but the model still needs thermal images to learn from because thermal contrast comes from temperature and emissivity rather than reflected color. In one 2026 LWIR drone study, synthetic pre-training plus only 100 real thermal images beat training on those 100 real images alone for every detector family tested.
Is the FLIR ADAS thermal dataset free to use?
The Teledyne FLIR ADAS dataset is a free download after registration, and version 2 contains 26,442 annotated frames across 15 classes. Check the conditions shown at registration before using the data in a commercial product.
What is the difference between LWIR and MWIR training data?
LWIR covers roughly 8 to 14 micrometers and is what most uncooled microbolometer cameras see, while MWIR covers roughly 3 to 5 micrometers and usually needs a cooled detector. Hot exhaust, engines and solar glint look different in the two bands, so a dataset captured or rendered in one band is not a drop-in substitute for the other.
Do GAN-translated thermal images work for training detectors?
Sometimes, with large differences between methods. In one ICRA 2023 study, a detector trained on edge-guided translated thermal images reached 0.239 mAP on the real FLIR validation set, while the same detector trained on CycleGAN translations reached 0.007. Translation also cannot recover temperature information that the RGB image never contained.
How do I align thermal images with RGB and depth?
Real rigs either share an optical path through a beam splitter, as the KAIST dataset did, or calibrate separate cameras and warp one image onto the other with a homography, which only holds for a single scene depth. Rendered data avoids the problem because every modality comes from the same virtual camera.