Synthetic Data For Lunar Landing And Planetary Hazard Detection
Lunar hazard detection models train on synthetic terrain because labeled lander-scale imagery barely exists. Teams build it from LOLA DEMs, procedural rocks and low sun.
The Short Answer
Teams training lunar landing hazard detection rely heavily on synthetic data, because labeled imagery at the scale a lander cares about barely exists. NASA's ALHAT hazard detection system was designed to detect 30 cm roughness and 1 degree slope hazards, while LRO's Narrow Angle Camera images the surface at 0.5 m per pixel from 50 km and the Robbins global crater catalog is approximately complete only above about 1 to 2 km. The usual pipeline starts from a LOLA digital elevation model (DEM), adds procedural craters and boulders below the DEM's resolution, renders them under physically based low-sun illumination, applies the lander's camera or lidar model, and exports per-pixel hazard, depth and pose labels.
Why Real Labeled Data Is Scarce
Descent imagery exists from a limited number of missions, and each used a different camera, altitude profile and sun angle. Almost everything we know about lander-scale terrain comes from orbit, and orbital data stops being useful at about the size that decides whether a landing leg is safe.
| Source | Resolution or completeness | What it can label | What it misses |
|---|---|---|---|
| LROC Narrow Angle Camera | 0.5 m per pixel from 50 km altitude | Craters and boulders a few meters across | A 30 cm rock is smaller than one pixel |
| LOLA high-resolution south pole DEMs | 5 m per pixel around high-priority landing sites (four sites in the first release) | Regional slopes and large crater geometry | Rocks, small craters and local roughness |
| Robbins crater database | About 1.3 million craters, complete above about 1 to 2 km | Landmarks for orbital and high-altitude navigation | Every crater a lander would avoid |
| Bickel et al. rockfall map | 136,610 rockfalls found in more than 2 million NAC images | Where boulders are mobile and common | Rock shapes and heights as a descent camera sees them |
Viewing geometry adds a second gap. Orbital images are close to nadir, taken from 50 km, at whatever sun angle the orbit allowed. A descent camera looks at the surface obliquely from a few kilometers down to a few meters, in its own spectral band, with its own exposure and motion blur. Even if you labeled every boulder in the NAC archive, you would not have descent-camera training data.
Recent missions show why this matters. ispace's HAKUTO-R Mission 1 crashed in April 2023, and the company's own analysis traced it to how the software handled the terrain. As the lander passed a cliff about 3 km high, which turned out to be a crater rim, its measured altitude jumped far from the estimated value. The onboard software treated the measurement as a sensor fault and rejected it, estimated its altitude as zero while it was about 5 km above the surface, and descended until its propellant ran out. ispace also noted that the landing site was changed after the February 2021 critical design review, and the earlier landing simulations did not adequately include the lunar environment along the new route. JAXA's SLIM, by contrast, landed on January 20, 2024 using image-matching navigation that compared its descent images against onboard maps. JAXA's project review evaluates its landing accuracy at around 10 m or better, based on images taken at about 50 m altitude, reports that obstacle detection functioned correctly, and notes that a propulsion problem made it drift east during the final descent and touch down about 60 m east of the target. Both outcomes depend on how terrain looks to the sensor during descent, which is the condition that is hardest to collect real data for.
Why Polar Illumination Makes It Harder
NASA's Artemis program and several commercial landers are aiming for the lunar south pole, and polar lighting breaks assumptions that equatorial imagery bakes in. Seen from the south pole, the Sun never moves more than 1.5 degrees above or below the horizon. Instead of rising and setting, it circles the horizon, so shadows are long and sweep across the terrain as the Moon rotates. NASA notes that long shadows and periods without sunlight can complicate navigation, and that in permanently shadowed regions sunlight may never reach the surface at all.
For a vision model this creates four specific problems:
- Extreme dynamic range. With no atmosphere to scatter light, shadows are close to black next to sunlit slopes. A camera exposed for one loses the other.
- Shadows that look like hazards. A long cast shadow can hide a boulder or look like a crater. Models trained at equatorial sun angles learn the wrong cues.
- Appearance that changes with sun azimuth. The same site looks different as the Sun circles the horizon, so a terrain-relative navigation map built from imagery taken at one sun angle may not match a descent days later.
- Regolith reflectance effects. Lunar soil brightens sharply when the Sun is directly behind the viewer (the opposition effect), so brightness depends on viewing geometry as well as on albedo.
NASA Ames built the POLAR stereo dataset to recreate these conditions on Earth: analog terrains lit by oblique sun sources and imaged in stereo, with over 2,500 HDR stereo pairs and lidar ground truth at sub-2 mm accuracy. It is a good real-world test set, but it covers 13 terrains built in a lab which is why it is usually paired with simulation.
Public Datasets And Tools
| Name | Type | What it provides |
|---|---|---|
| NASA POLAR | Real analog imagery | 13 analog terrains under polar-style lighting, over 2,500 HDR stereo pairs, lidar ground truth |
| POLAR-Sim (Chen et al.) | Labels plus digital twins | About 23,000 labels and semantic segmentation annotations for rocks, shadows and craters on POLAR, plus 3D reconstructions of all 13 terrains |
| Artificial Lunar Landscape Dataset (Pessia and Ishigami, Keio University) | Synthetic renders | 9,766 Terragen renders with masks for sky, smaller rocks and larger rocks, plus bounding boxes for larger rocks; no depth or camera pose |
| LuSNAR (Liu et al., 2024) | Synthetic multi-sensor | 9 Unreal Engine lunar scenes with stereo pairs, semantic labels, dense depth, lidar and rover pose |
| Synthetic Lunar Terrain (Märtens et al., 2024) | Real analog, multimodal | Event camera, RGB and laser scans of an analog site with synthetic craters under high-contrast lighting |
| DeepMoon (Silburt et al., Icarus) | Crater detection model and data | A CNN trained on lunar DEMs that recovered 92% of test-set craters and nearly doubled the number of detections |
| PANGU | Simulator (University of Dundee, ESA-funded) | Rendered camera views of planets and asteroids, from orbit down to the surface, for testing vision-based navigation for landers and rovers |
| OmniLRS | Open-source simulator on Isaac Sim | Procedural lunar terrain, a large-scale environment in later releases that refines a 5 m DEM procedurally to 2.5 cm, and a synthetic data pipeline; the paper reports a YOLOv8 rock segmentation model trained on its synthetic data came within 5% of one trained on real data |
| NASA DUST | Simulator on Unreal Engine 5 | South pole terrain from LRO DEMs with Sun lighting, for early landing and traverse site analysis |
None of these is a ready-made training set for a specific lander. The synthetic sets use generic cameras and sun angles, and the real analog sets are small. Most teams use them for benchmarking and build their own synthetic data matched to their sensor and landing site.
A Synthetic Data Pipeline For Hazard Detection
A hazard detection pipeline has to produce imagery that matches the descent sensor and labels that match the hazard definition. The stages below are the ones that matter for both.
- Terrain base from a DEM. Start from the best DEM for the target site, such as the 5 m per pixel LOLA products for the south pole. This fixes the large-scale slopes, crater rims and horizon that drive shadows and navigation landmarks.
- Procedural craters and rocks below DEM resolution. Add small craters and boulders that the DEM cannot resolve, using size-frequency distributions calibrated against counts from NAC imagery of the same terrain type. Rock shape, burial depth and clustering around fresh craters all change how hazards look.
- Regolith materials. Use a reflectance model built for lunar soil, such as a Hapke-style BRDF, so brightness behaves correctly with phase angle, including the opposition surge.
- Physically based illumination. Place the Sun from ephemeris for the site and time window (NASA's SPICE system provides the ephemerides), with low elevations for polar sites, sharp shadows and no atmospheric fill. Add lander-mounted lights if the mission will descend into shadow.
- Sensor model. Match the descent camera's intrinsics, spectral response, exposure, bit depth, noise, blur and rolling shutter. For lidar hazard detection, simulate range noise, beam pattern and dropouts on dark or steep surfaces.
- Descent trajectories. Render along realistic approach paths, from high-altitude navigation frames to the final hazard-scan altitudes, so the model sees each object at every scale it will meet.
- Labels. Export per-pixel hazard classes, rock instance masks, depth, surface normals or slope maps, crater rims and the exact camera pose for every frame. Log sun elevation and azimuth per frame so you can slice results by lighting.
- Validation on real data. Hold out real imagery, such as NAC frames of comparable terrain or POLAR stereo pairs, and measure the gap before trusting a synthetic-only model.
Synthetic data helps most at steps 2, 4 and 7. You control rock distributions and sun angles that no orbital pass will give you, and labels come from the scene geometry instead of from a human tracing shadows. Step 8 matters as much, and our post on closing the sim-to-real gap for perception covers how to measure that gap before you try to close it.
Tasks And The Data Each Needs
| Task | When it runs | Typical sensor | Labels needed | What synthetic data adds |
|---|---|---|---|---|
| Hazard detection and avoidance | Final descent, roughly the last kilometer | Lidar, camera | Rock and crater masks, slope and roughness maps, safe-site maps | Rock fields and slopes at sites no probe has imaged at that scale, across sun angles |
| Terrain-relative navigation | High and mid-altitude descent | Camera, matched to an onboard map | Camera pose, landmark (crater) positions | Many sun angles and viewpoints of the same mapped terrain, with exact pose |
| Crater detection | Orbit and descent, also science mapping | Camera or DEM | Crater rims and diameters | Small craters below catalog completeness, with exact rims |
| Rock segmentation | Descent and surface operations | Camera, stereo | Per-pixel rock instances, depth | Exact masks for small and partly shadowed rocks that humans label poorly |
| Rover traversability | Surface operations | Stereo, depth | Terrain classes, depth, slope | Long traverses with consistent labels under changing light |
Terrain-relative navigation is already flight-proven. Mars 2020's terrain-relative navigation matched descent images to preloaded maps to guide Perseverance into Jezero Crater, and NASA's SPLICE project packages terrain-relative navigation with a navigation Doppler lidar, a hazard detection lidar and a descent and landing computer for future landers. Each of these systems needs to be tested against far more terrain and lighting combinations than real data can supply.
Where Stardust Fits
Stardust, our synthetic data platform, has a space vertical built for lunar and Martian verification and validation. You can upload a custom DEM, select surface materials and set crater and rock distributions to recreate a specific site. Lighting gives full control over sun elevation, artificial illumination and lunar-specific effects such as the opposition effect. Every frame ships with per-pixel segmentation of terrain classes (bedrock, sand, rocks and hazards), camera pose, per-pixel depth, and sun elevation and illumination metadata. Sensors include RGB and EO, grayscale, depth and segmentation, and you can set camera traversal paths through the terrain to produce time-series datasets. Data can be generated in real time at 24 FPS on off-the-shelf GPUs. NASA appears among the organizations shown on our space page.
Stardust does not replace the public datasets above. Use POLAR and real NAC imagery to measure the sim-to-real gap, and use synthetic data to cover the rock fields, slopes and sun angles they do not. For related reading, see our posts on neural rendering and on comparing synthetic data platforms.
Sources
- LROC, camera specifications
- NASA Goddard PGDA, High-Resolution LOLA Topography for Lunar South Pole Sites
- USGS Astropedia, Moon Crater Database v1 (Robbins)
- Bickel et al., Impacts drive lunar rockfalls over billions of years, Nature Communications 2020
- NASA Langley, Autonomous Landing Hazard Avoidance Technology (ALHAT), April 2014
- ispace, Results of the HAKUTO-R Mission 1 lunar landing, May 2023
- JAXA, SLIM project review press briefing, December 2024
- NASA, Earth and Sun from the Moon's south pole and Moonbase environment
- NASA Ames, POLAR Stereo Dataset
- Chen et al., POLAR-Sim; Liu et al., LuSNAR; Märtens et al., Synthetic Lunar Terrain; Silburt et al., DeepMoon; OmniLRS and its repository
- ESA, PANGU; NASA, DUST, SPLICE, terrain-relative navigation
Frequently Asked Questions
Is there a public dataset for lunar rock and boulder detection?
The Artificial Lunar Landscape Dataset from Keio University has 9,766 rendered lunar scenes with masks for sky, smaller rocks and larger rocks. POLAR-Sim adds about 23,000 rock, shadow and crater labels to NASA's real POLAR analog images, and LuSNAR provides semantic labels, depth and lidar across 9 Unreal Engine lunar scenes. For orbital boulder tracks, Bickel and colleagues mapped 136,610 lunar rockfalls in LRO imagery.
How many craters are in the Robbins lunar crater catalog?
About 1.3 million, according to the USGS Astropedia listing. The catalog is approximately complete for craters larger than about 1 to 2 km in diameter, which makes it useful for navigation from orbit but far too coarse for lander-scale hazards.
What size of hazard does a lunar lander need to detect?
NASA's ALHAT project designed its hazard detection lidar to resolve local slopes to about 1 degree and roughness hazards of 30 cm, over a hazard map roughly 200 by 200 m. The exact threshold for any lander depends on its leg geometry and ground clearance.
Why is lunar south pole lighting hard for computer vision?
At the south pole the Sun never moves more than 1.5 degrees above or below the horizon, so shadows are long and sweep across the terrain as the Moon rotates. With no atmosphere to scatter light, shadowed areas are close to black, so cameras face extreme dynamic range and the same terrain can look very different a few days apart.
What software can generate synthetic lunar images?
Common options include PANGU, developed by the University of Dundee with ESA funding, OmniLRS, an open-source lunar simulator built on NVIDIA Isaac Sim, NASA's DUST tool built on Unreal Engine 5, and commercial platforms such as Bifrost Stardust, which accepts custom DEMs and controls sun elevation and lunar lighting effects.