BIFROST
All posts

Synthetic Thermal Infrared Training Data: How To Get Labeled IR Imagery

Labeled thermal IR training data comes from three routes: hand-labeling real frames, translating RGB images to thermal, or rendering scenes with physics-based simulation.

The Short Answer

There are three ways to get labeled thermal infrared training data. You can collect real thermal footage and label it by hand, which gives you the true sensor signal but costs the most per labeled frame. You can translate existing labeled RGB images into thermal-looking images with GANs or diffusion models, which is cheap but only as good as the model's guess about temperatures it never saw. Or you can render scenes with physics-based simulation, where materials carry emissivity and temperature, and every frame arrives with exact labels, depth and segmentation. Most teams that ship thermal detectors end up mixing routes: a public dataset such as the 26,442-frame Teledyne FLIR ADAS set for real signal, synthetic data for coverage, and a held-out set of their own real frames to check that the mix transfers.

Route 1: Collect And Label Real Thermal Imagery

Real thermal data is the reference everything else is measured against. It contains the actual noise, optics blur, non-uniformity and gain behavior of the camera you will deploy, which no synthetic method fully reproduces.

The costs are the familiar ones, made worse by the modality. Thermal cameras usually cost more than RGB cameras, the scenes you most need (night, fog, smoke, rare objects at long range) are the hardest to stage, and small, low-contrast thermal targets are hard to label consistently. Teledyne FLIR's counter-UAS sensor guide warns that manual annotation of small targets introduces label noise that can confuse the detector.

If you go this route, record raw radiometric frames as well as the 8-bit display output. The FLIR ADAS dataset ships thermal images both as 14-bit TIFFs with no automatic gain control (AGC) and as 8-bit JPEGs with AGC applied, so you can train on either and compare. Most other public datasets only provide the processed 8-bit images, and that choice is baked into every model trained on them.

Route 2: Translate RGB Images Into Thermal

Image-to-image translation takes a labeled RGB dataset and converts each image into a thermal-looking image while keeping the labels. The early methods are general-purpose translation networks: pix2pix for paired data and CycleGAN for unpaired data. Later work specializes them for thermal, for example InfraGAN (Özkanoğlu and Ozer, Pattern Recognition Letters, 2022) and the edge-guided multi-domain model of Lee, Jeon, Cho, and Kim (ICRA 2023). More recent work moves to diffusion models: DiffV2IR combines a diffusion translator with vision-language understanding and introduces IR-500K, a dataset of 500,000 infrared images.

The Lee et al. paper shows how much the choice of method matters. They generated four synthetic thermal detection datasets from the same synthetic RGB source using four translators, trained a VFNet detector on each, and evaluated all four on the real FLIR-ADAS validation set.

Translator used to make the training setmAPmAP50
Edge-guided (Lee et al., proposed)0.2390.408
UNIT0.2250.384
MUNIT0.0690.137
CycleGAN0.0070.010

Source: Lee et al., ICRA 2023, Table III. Detector trained only on translated images, tested on real FLIR-ADAS validation frames.

The same labeled source images produced a usable detector with one translator and a near-useless one with another. Even the best result, 0.239 mAP, is a starting point for fine-tuning rather than a deployable model.

There is also a physical limit. An RGB image records reflected visible light, and a thermal image records emitted long-wave radiation, which depends on surface temperature and emissivity. A parked car and a car that has been driving for an hour look the same in RGB and very different in LWIR. A translator has to guess which one it is looking at from the statistics of its training set, so it works best when your target scenes match its paired training data and worst for the out-of-distribution cases you wanted synthetic data for.

Route 3: Physics-Based Rendering

Physics-based rendering builds a 3D scene, assigns materials their thermal properties, and simulates what a given sensor would record. The IR signature of an object depends on its temperature, its emissivity, reflections of its surroundings and the sky, solar loading, atmospheric attenuation and the sensor's own optics and electronics. Engines such as RIT's DIRSIG, which simulates imagery from the visible through the thermal infrared, model this chain for remote sensing, and a 2024 survey by Upadhyay et al. covers both these model-based methods and newer deep learning approaches.

Rendering has two structural advantages. Labels are exact, because they come from the scene graph rather than from an annotator reading a blurry hot spot. And you control the variables that matter for thermal: time of day, ambient temperature, which engines are running, how warm the water is relative to the air, and the target's range and aspect. That makes it the practical way to cover thermal crossover (the hours when target and background reach the same apparent temperature and contrast disappears), which real collections capture only by luck.

The failure mode is the reverse of translation's. A renderer is physically consistent, but a detector can still tell its output apart from real frames if the sensor model is too clean. A 2026 study of synthetic-first drone detection in LWIR (Liiv et al.) measured this gap. Its 20,000 synthetic samples combine generative IR-like backgrounds with drones rendered in Blender, which the authors say is not a fully radiometric LWIR simulator. Real frames had higher histogram total variation, which the authors describe as consistent with AGC-like stretching of high-bit-depth data into 8 bits, and the synthetic pipeline did not reproduce it. Inside the target boxes, Sobel gradient variance was over four times higher in real detections than in synthetic ones. Synthetic pre-training without fine-tuning did not beat a baseline trained on 100 real LWIR images (Faster R-CNN was the exception), but synthetic pre-training followed by fine-tuning on those 100 images beat the real-only baseline for every detector family tested. All the detectors started from COCO weights, so the synthetic data served as a second pre-training stage.

Hybrid approaches render only the objects and composite them into real thermal backgrounds, such as Bongini et al. on FLIR ADAS and Barisic Kulas et al. (ECMR 2025), who added a drone class to the HIT-UAV aerial thermal dataset. These keep real sensor statistics in the background and control the targets, at the cost of lighting and thermal interactions between object and scene.

How The Three Routes Compare

Real capture and labelingRGB-to-thermal translationPhysics-based rendering
Label qualityLimited by annotator skill on low-contrast, small targetsInherited from the RGB labels; can drift if the translator moves or erases objectsExact, from the scene graph, including depth and segmentation
Physical fidelityGround truth for the camera usedPlausible appearance, no real temperature or emissivity informationConsistent with the thermal model; depends on material data and sensor model
Cost per labeled frameHighest (hardware, field time, annotation)Lowest once a translator is trainedModerate; mostly scene and asset setup, then cheap to scale
Rare and dangerous casesCaptured only if they happen during collectionLimited to scenes that exist in RGBCan be composed on demand
Main failure modeCoverage gaps and label noiseHallucinated thermal states; large variance between methodsSensor model too clean (missing AGC, noise, blur)
Best useValidation set and fine-tuning dataQuick bootstrapping when good paired RGB data existsCoverage of conditions, ranges and objects real data lacks

Whatever mix you choose, keep a held-out set of real frames from your own sensor that no training route touches, because results on that set tell you whether the synthetic data helped. Our posts on whether synthetic data works for object detection and closing the sim-to-real gap for perception cover mixing and measurement in more detail.

Public Thermal Datasets Worth Starting With

DatasetContentSizeRegistration with RGBTerms
Teledyne FLIR ADAS v2Driving scenes, 15 classes, thermal (Tau 2, 640×512) and RGB26,442 annotated frames, 520,000 boxes; 9,711 thermal and 9,233 RGB train/val images plus 7,498 video framesSeparate thermal and RGB cameras; an earlier release had enough misaligned pairs that Zhang et al. hand-filtered a 4,129 train / 1,013 test aligned subsetFree with registration
KAIST Multispectral PedestrianDriving scenes, day and night, person / people / cyclistAbout 95,000 color-thermal pairs at 640×480, 103,128 annotations, 1,182 unique pedestriansAligned in hardware with a beam splitterCC BY-NC-SA 4.0 by default; some bundled code is BSD 2-Clause
LLVIPMostly very dark scenes, pedestrians15,488 visible-infrared pairs (30,976 images)Strictly aligned in time and space; raw unregistered data also releasedNon-commercial use only
M3FDCampus, coastal and urban scenes, 6 classes4,200 pairs, mostly 1024×768Registered via calibration and homographyRepository is GPL-3.0; no dataset-specific license stated
HIT-UAVAerial thermal from 60 to 130 m altitude, persons and vehicles2,898 images, 24,899 labeled objects (including a DontCare category)Thermal onlyCC BY 4.0

Licenses vary a lot, and two of the most cited datasets (KAIST and LLVIP) restrict the data to non-commercial use. Almost all of the data is also street-level driving or surveillance of people and cars, which leaves maritime, aerial, defense and industrial thermal applications with little public data to start from. For maritime LWIR, our roundup of maritime datasets lists what exists.

Sensor Details That Decide Whether Synthetic IR Transfers

LWIR And MWIR Are Different Training Domains

Long-wave infrared (roughly 8 to 14 micrometers) is where objects near room temperature emit most strongly; by Wien's displacement law, a 300 K surface peaks at about 9.7 micrometers. Most affordable thermal cameras are uncooled LWIR microbolometers. Mid-wave infrared (roughly 3 to 5 micrometers) usually needs a cooled detector, and Teledyne FLIR's counter-UAS guide describes cooled MWIR cameras with resolutions up to 1280×1024 and focal lengths beyond 1000 mm for long-range work. Hot sources such as exhaust plumes, and reflected sunlight, weigh more heavily in MWIR. Capture or render training data in the band of the sensor you deploy, and ask which band a dataset labeled only "IR" or "thermal" was captured in.

Registration With RGB And Depth

Fusion models need thermal, RGB and often depth to line up pixel for pixel, which is hard when the cameras do not share a lens. KAIST solved it in hardware with a beam splitter. M3FD calibrates the cameras and warps the infrared image with a homography, which is exact only for a single plane, so nearby objects with parallax still shift. FLIR ADAS uses two separate cameras, which is why researchers hand-filtered aligned subsets of its first release for fusion work. Rendered data avoids the problem, since every modality comes from the same virtual camera.

Noise, Non-Uniformity Correction And Gain

Uncooled microbolometers drift with temperature, so the pixels need periodic re-normalization against a uniform source, usually a shutter, to refresh the per-pixel non-uniformity correction (NUC). Teledyne FLIR calls this step flat field correction and triggers it by elapsed time and temperature change. Between corrections you see fixed-pattern noise and slow drift, and after AGC remaps 14-bit data into 8 bits, the histogram of each frame depends on whatever else is in the scene, so a hot object entering the frame changes how every other pixel is displayed. Synthetic thermal data that skips these effects looks cleaner than any real camera output, which is the gap Liiv et al. measured. At a minimum, add fixed-pattern noise, temporal noise, optics blur and a gain stage that mimics your camera's AGC, or train on raw radiometric values end to end if your deployment pipeline allows it.

Where Stardust Fits

Stardust is Bifrost's synthetic data platform, and it takes the rendering route. One scene is rendered across modalities in registration, with RGB, IR and depth from the same 3D scene and ground-truth labels generated automatically for every frame. The Thermal with Depth render output produces thermal imagery and a per-pixel depth map from a single render, pixel-aligned by construction, so there is no cross-sensor registration step. For maritime work, Stardust simulates RGB, stereo, SWIR, MWIR and thermal sensors matched to your hardware (maritime), and the aerial page covers EO, IR and thermal for UAV perception. Scenes can be varied across fog, rain, snow and night operations, sensors can be perturbed with noise, intrinsics and perspective shifts, and a neural rendering pass adds the glare, grain and noise of real sensors on top of the 3D render.

The same advice applies to our data as to anyone's. Synthetic thermal data earns its place by covering the conditions and objects your real data lacks, and it should be judged on a held-out set of real frames from your own camera.

Sources

Frequently Asked Questions

Can I train a thermal object detector on RGB images alone?

You can start from RGB-pretrained weights, and thermal studies routinely begin from COCO-pretrained detectors, but the model still needs thermal images to learn from because thermal contrast comes from temperature and emissivity rather than reflected color. In one 2026 LWIR drone study, synthetic pre-training plus only 100 real thermal images beat training on those 100 real images alone for every detector family tested.

Is the FLIR ADAS thermal dataset free to use?

The Teledyne FLIR ADAS dataset is a free download after registration, and version 2 contains 26,442 annotated frames across 15 classes. Check the conditions shown at registration before using the data in a commercial product.

What is the difference between LWIR and MWIR training data?

LWIR covers roughly 8 to 14 micrometers and is what most uncooled microbolometer cameras see, while MWIR covers roughly 3 to 5 micrometers and usually needs a cooled detector. Hot exhaust, engines and solar glint look different in the two bands, so a dataset captured or rendered in one band is not a drop-in substitute for the other.

Do GAN-translated thermal images work for training detectors?

Sometimes, with large differences between methods. In one ICRA 2023 study, a detector trained on edge-guided translated thermal images reached 0.239 mAP on the real FLIR validation set, while the same detector trained on CycleGAN translations reached 0.007. Translation also cannot recover temperature information that the RGB image never contained.

How do I align thermal images with RGB and depth?

Real rigs either share an optical path through a beam splitter, as the KAIST dataset did, or calibrate separate cameras and warp one image onto the other with a homography, which only holds for a single scene depth. Rendered data avoids the problem because every modality comes from the same virtual camera.

Explore Stardust More Posts