NASA‑IBM Lunar Foundation Model: a multimodal lunar AI backbone for research, not navigation

NASA and IBM turned 17 years of lunar data into a reusable AI backbone, usable for research, not for navigation

NASA and IBM Research, with academic partners, released the NASA‑IBM Lunar Foundation Model (LFM). It’s a publicly available, pretrained multimodal foundation model built on SomBench, which the teams call the largest co‑registered lunar corpus to date. The model combines imagery, gravity, composition and other layers into a single backbone researchers and companies can fine‑tune for mapping and detection, but it has clear limits.

Below is a concise technical summary, where to get the model, what the team reports about performance (and why to treat those numbers cautiously), practical use cases with validation notes, and a prioritized checklist procurement and mission teams should insist on before adoption.

“NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job.”, Kevin Murphy, NASA’s chief science data officer.

SomBench: the dataset behind LFM (the hard numbers)

NASA/IBM say SomBench holds nearly 2 million spatially aligned “tile bundles” (~1.96M total), assembled from 17 years of Lunar Reconnaissance Orbiter (LRO) data plus other missions. Key counts the teams provide:

  • ~1, 000, 000 Narrow Angle Camera (NAC) images at roughly 1 m/pixel (high resolution).
  • Just under 964, 000 Wide Angle Camera (WAC) multispectral images at ~100 m/pixel (coarse scale).
  • More than 30 spatially aligned data layers from nine instruments across four missions, including LRO, GRAIL (gravity), Lunar Prospector (hydrogen), and JAXA’s Kaguya/SELENE (mineralogy).
  • Corpus splits were done by geographic map‑zones to reduce spatial leakage between training, validation and test sets.

NASA/IBM describe SomBench as “the largest co‑registered multimodal lunar corpus to date.” The teams also publish the ML‑ready pretraining datasets and benchmark collections alongside the model artifacts (see availability below).

What LFM actually is, architecture and training choices that matter

LFM is a pretrained multimodal backbone meant to be adapted to downstream lunar tasks. Important engineering choices reported by the authors:

  • The model used an architecture based on TerraMind (an Earth observation multimodal model), but LFM was trained from scratch on SomBench rather than fine‑tuned from TerraMind weights.
  • Each training tile includes explicit imaging geometry metadata (solar illumination angles, sun position, tile extent). Feeding lighting and viewing metadata alongside pixels reduces the need for the model to implicitly learn complex shadow and phase effects that dominate lunar imagery.
  • Multi‑scale training was enabled by a FlexiViT adapter: the same model can accept different patch sizes (coarse 100 m and fine 1 m) without retraining separate networks.
  • For downstream adaptation, the team evaluated both full fine‑tuning and LoRA (Low‑Rank Adaptation) lightweight adapters; LoRA kept pace on most tasks and outperformed on crater detection, while full fine‑tuning won only on the smallest tasks.

Model parameter count and the exact training compute (GPU/TPU types × hours) are not published in the materials the team released; procurement teams should request those figures before committing to a deployment or internal reproduction effort.

Where to find the model and code

The authors report the following releases and integrations:

  • Model weights and a technical report on Hugging Face (model card and artifacts).
  • Code and training scripts on GitHub: the repository listed is NASA‑IMPACT/NASA‑IBM‑Lunar‑Foundation‑Model.
  • Integration into the TerraTorch open‑source toolkit for geospatial ML.

The team’s materials do not fully document license details, parameter counts, or training compute in the public artifact set; verify the Hugging Face model card and GitHub README for licensing and provenance metadata before reuse.

Reported performance, what IBM/NASA say, and what to verify

The team evaluated LFM on several tasks: crater detection at 100 m and 1 m scales, polar permanently‑shadowed region ice prediction, and segmentation of Irregular Mare Patches (IMPs). IBM reports the pretrained LFM “matches or outperforms” common baselines (SwinV2‑B) on these tasks, with headline deltas that the authors publish:

  • Polar ice prediction: up to a 22% reduction in prediction error versus SwinV2‑B (IBM report).
  • Coarse‑scale (100 m) crater detection: nearly a 19% improvement over SwinV2‑B; that comparison reportedly used only half the training data for the baseline in the cited run.
  • IMP segmentation: approximately a 3% lead over SwinV2‑B (differences appear small and may be within variance).

Two important caveats:

  • Absolute metric values, the exact metric definitions (e.g., RMSE, IoU, F1), test set sizes and confidence intervals are not always shown in the summarized press materials. Percentage deltas without absolute scores or error bars can be misleading; ask for the full evaluation tables before drawing conclusions.
  • The authors report that an architecturally identical, randomly initialized control model (no lunar pretraining) also outperformed five of seven baselines on the ice prediction task. That counterintuitive result can indicate (a) baseline selection or training parity issues, (b) task fragility with small test sets, or (c) evaluation mismatches. Request the control model training logs and baseline training regimes to understand what happened.

Hard limits and operational cautions

The project team is explicit about what LFM is not. Notably:

  • It is not a substitute for in‑situ measurements. For mission‑critical needs (landing site certification, resource confirmation, navigation), physical measurements and rigorous mission assurance remain mandatory.
  • Generation tests reported by the authors produced substantial geodetic errors, described as tens of degrees of latitude/longitude error in some trials, and absolute elevation offsets in generated outputs. That makes any autonomous geodetic generation or coordinate production from the model unsuitable for operational navigation or precision landing without additional calibration and validation.
  • Controlled ablation studies that isolate the contribution of multimodality, explicit imaging geometry, FlexiViT, and the overall pretraining regime are still pending. Until those appear, it’s unclear which component(s) drive the reported gains.

Three near‑term use cases (with validation notes)

  • Polar prospecting prioritization

    Use LFM outputs to rank candidate zones in permanently shadowed regions for follow‑up remote sensing and targeted lander scouting. ROI: reduce expensive reconnaissance time by focusing follow‑ups on top N ranked targets. Validation: require an observational follow‑up plan (orbiter or lander) that confirms >X% of top K predictions before trusting the model for mission decisions.

  • Scale‑bridging crater catalogs

    Automate crater detection at regional (100 m) and meter scales to speed chronology and hazard mapping. ROI: accelerate catalog updates across scales and save manual annotation hours. Validation: compare automated detections to a vetted human‑annotated subset and report precision/recall by crater diameter bins.

  • Transfer learning template for other bodies

    Adopt the LFM approach (co‑registration + multimodal pretraining + imaging geometry inputs) as a pipeline for Mars or asteroid datasets. ROI: reduce labeled‑data needs for morphology and composition tasks on other bodies. Validation: run small‑scale transfer experiments and report delta in labeled samples required to reach target accuracy.

What procurement and mission teams should demand, three highest‑priority asks

If you can only request three items before allocating budget or greenlighting operational integration, make them these:

  1. Full dataset provenance and licensing. Dataset CRSs, instrument footprints, preprocessing pipelines, and a clear license for model weights and datasets.
  2. Training and model compute accounting. Model parameter count, hardware used (GPU/TPU types), and total wall‑clock training hours or FLOPs so engineering teams can cost reproduction or fine‑tuning.
  3. Complete evaluation tables and ablations. Absolute metrics, test set sizes, confidence intervals, and controlled ablation experiments that isolate imaging geometry, FlexiViT, and multimodality effects.

Beyond those three, insist on reproducible training artifacts (containerized scripts), a public timeline for planned ablations, and an operational validation plan tied to in‑situ or future lander datasets for any output you plan to use in mission design or resource claims.

Governance, risk, and a short checklist for C‑suite decision makers

  • Confirm license permits your intended commercial or government use and redistribution.
  • Require dataset provenance metadata attached to every layer (instrument, processing level, CRS).
  • Ask for independent validation plans and acceptance tests before operational deployment.
  • Document known failure modes (e.g., geodetic generation errors, low confidence regions) and publish a “performance envelope” that lists where the model is known to fail.
  • Plan for a staged adoption: internal research → targeted field validation → cautious operational pilot with human oversight.

Big picture, why this matters to business and science

LFM is part of a broader industry trend: groups are compressing long, multimodal geospatial archives into pretrained backbones that reduce labeled‑data friction and let cross‑instrument patterns emerge. That matters commercially because firms offering lunar services must quickly triage terrain, hazards and resource prospects across instruments and scales, a reusable backbone can cut weeks of manual annotation and harmonization work.

But the distinction between research readiness and operational readiness is critical. LFM looks valuable as a research and prioritization tool: it can speed mapping, reduce labeling, and provide transferable embeddings. Turning it into a certified operational tool for landing, navigation or resource confirmation will require the ablation results, provenance, and in‑situ validation the teams themselves flag as the next steps.

Key questions and short answers

  • What data trained the model?

    SomBench: nearly 2 million co‑registered tile bundles (~1.96M), spanning 11 modalities and two spatial scales (NAC ~1 m/pixel and WAC ~100 m/pixel), assembled from 17 years of LRO data plus contributions from GRAIL, Lunar Prospector, and Kaguya/SELENE, as reported by NASA/IBM.

  • Where can I get the model and code?

    The team reports model weights and a technical report on Hugging Face, code on GitHub under NASA‑IMPACT/NASA‑IBM‑Lunar‑Foundation‑Model, and an integration into the TerraTorch toolkit. Verify the model card and GitHub README for licensing and provenance details before reuse.

  • Does it outperform existing baselines?

    IBM reports up to a 22% reduction in error on polar ice prediction and nearly a 19% improvement on coarse crater detection versus SwinV2‑B. These are team‑reported relative improvements; request absolute metric tables, test set sizes and confidence intervals to assess statistical significance and operational relevance.

  • Can I use it for landing site selection or navigation?

    No. The authors explicitly warn the model is not a substitute for in‑situ measurements. They report generation tests with tens of degrees of latitude/longitude error in some trials and elevation offsets, so treat LFM as a prioritization and research asset, not an operational navigation source.

Practical next step for technical leads: pull the Hugging Face model card and GitHub README, confirm license and provenance, and request the three priority items listed above (dataset provenance, training compute/parameter counts, and full evaluation/ablation tables) from the project team before any procurement or operational testing.