Closing Cloud Detection Gaps in Optical Satellite Imagery

The Challenge: Clouds render satellite imagery useless 

Cloud detection is the critical first step in almost every optical Earth Observation (EO) task. Before any AI model can extract valuable insights from an image, it must first distinguish between clear pixels and those obscured by clouds. While thick clouds fully block the ground and are easy to detect and discard, thin clouds present a much more complex bottleneck.

Thin clouds only partially obscure the ground, meaning the data underneath might still be usable. If a model cannot make this subtle distinction, it will either discard valuable imagery or process clouded pixels as if they were clear, ruining the analysis. Unfortunately, public datasets lack the high volume and variety of thin clouds needed to train models effectively, leaving them undertrained and prone to costly errors.

The Solution: On-Demand Synthetic Cloud Generation Another Earth resolves this training gap with precisely engineered synthetic data. We render specific cloud types, altitudes, and densities over any chosen landscape, generating pixel-perfect labels automatically at the point of creation. Because we control the environmental parameters, our synthetic scenes contain 2.3x more thin clouds (21% of pixels) compared to real archives (9%), directly targeting the models’ weakest point.

Proven Results: Boosting Accuracy & Reducing Cost. We rigorously validate our models on a held-out test set of ~1,000 real-world scenes with varying cloud coverage—never on synthetic data. The results demonstrate massive efficiency gains:

  • Drastic Data Reduction: We can reduce manual ground truth (GT) requirements by up to 75% while keeping model performance stable.
  • Scarce Data Acceleration: Starting with just 400 real images, adding 3x synthetic data boosted thin-cloud IoU from 29% to 36%—matching the performance of a fully real dataset four times its size.
  • Enhancing Large Datasets: Even with a robust baseline of ~1,700 real images, adding synthetic data improved thin cloud accuracy from 36% to 44%.

The Next Frontier: Cloud Removal We are now applying this technology to cloud removal. Reconstructing what sits beneath a cloud has traditionally lacked good training data, as perfect “before and after” satellite images of the exact same moment do not exist. By rendering identical synthetic scenes twice—once clear and once clouded—we produce the exact matched pairs required to solve this problem.