Four synthetic satellite image tiles generated by Another Earth: two cloud scenes and their corresponding cloud masks.

Synthetic satellite data, answered: The 5 questions we heard most

June was a month of conversations. We spent it across three events that each brought together a different crowd but all circled the same future: Web Summit Rio, MundoGEO Connect in São Paulo, and EO Summit in London. The energy in each room was its own thing, fast and founder-heavy in Rio, deeply technical and geospatial in São Paulo, and focused squarely on Earth observation in London. Three very interesting weeks, full of listening, sharing, learning and many, many questions. From us and towards us.

Synthetic satellite data is still new to many, including some who work with Earth observation every day, and there is genuine curiosity about what it is and where it fits. Rather than leave those answers in passing booth chats, we have gathered the five questions we were asked most often and answered each one properly below.

What is synthetic satellite data?

This was the first question almost everywhere, and usually the most important to get right. The word ‘synthetic’ is often met with a little mistrust as it is associated with fake, doctored or random. But actually, the opposite is true. Synthetic satellite data is an image created in a way that makes it physically and statistically faithful to an image a real satellite sensor would record over a given scene. It is an artificially created satellite image that is not only visually identical to a real-world image but made in a way, down to the finest details, that even a highly sophisticated machine learning model will treat and learn from it like it would from a satellite-captured one. The bonus: we choose the conditions, which is the exact point where the value lies. With this we have flexibility and freedom to create an equal dataset, but without the dependencies on real-world conditions.

Instead of waiting for a satellite to pass over the right place at the right moment, we generate the scene ourselves. The difference from a generic AI picture is just as important. We are not producing something that merely looks plausible to the human eye. We are reproducing the way light, terrain, materials, and a sensor actually interact, so the result behaves like real data when a model learns from it.

The advantage of generating a scene is that we know everything about it. Every object, boundary, and surface type is known because we placed it there, which means the labels come for free and they are exact. That is rarely true of real imagery. So synthetic data is less a substitute for reality and more a way to produce reality on demand, with perfect ground truth attached.

Why use synthetic satellite data if real imagery already exists?

A reasonable challenge, and we heard it often from people who live close to real archives. The honest answer is that real imagery is abundant in the wrong proportions. There is an enormous amount of ordinary, cloud-free, mid-latitude farmland, and very little of the rare conditions that models most need to learn.

Think about what actually matters operationally: a particular crop disease in its early stage, a specific vessel type at sea, a flood at its peak, a wildfire front, an oil spill, or structural damage after an earthquake. These events are scarce in the historical record. When they do appear, someone still has to find them and label them by hand, which is slow and expensive, and human labelling introduces its own errors.

Synthetic data lets us produce balanced datasets that include the rare and awkward cases on demand, already labelled with precision. We can vary season, illumination, sensor, and geography deliberately, rather than hoping the archive happens to contain the spread we need. None of this replaces real imagery. The strongest results almost always come from combining the two, using real data for grounding and synthetic data to fill the gaps it cannot cover.

How does synthetic satellite data handle cloud cover?

We handle clouds by generating them ourselves rather than only removing them, a small reversal that gives us something real archives almost never offer. Clouds are a constant problem, and this question came up at every event. At any given moment, roughly two-thirds of the planet sits under cloud, so a large share of optical satellite imagery is partly or fully obscured. For anyone relying on a clear view of the ground, that is a persistent and costly problem.

This is one of the areas we focus on most. We generate synthetic cloud data built specifically to master cloud detection and to train cloud-removal models. Because we create every scene ourselves, we know exactly where each cloud and its shadow sit, down to thin cirrus and faint haze that are notoriously hard for models to recognise. That gives perfectly accurate masks for training detection models.

Just as usefully, we can render the same scene twice, once clouded and once clear, as a matched pair. That pairing is almost impossible to obtain from real archives, where you rarely capture an identical scene under both conditions with reliable ground truth. With paired data, a cloud-removal model has a clean target to learn towards rather than an approximation. The result is models that can flag, mask, and reconstruct obscured areas far more reliably than imagery alone would allow.

Can models trained on synthetic data work in the real world?

Yes, but never on faith: a model trained on synthetic imagery only earns trust once it is measured against real data, and there is a specific reason that test can be passed. This was the sharpest question we received, and rightly so. This may be the question that matters most, because synthetic data is only worth anything if it transfers to reality.

The gap between synthetic and real data has a name, the domain gap, and closing it is the core of the work. We do not assume a model trained on generated imagery will simply perform in the field. We measure it. Performance is benchmarked on held-out real imagery, and the synthesis is tuned so that the statistical and physical properties of our scenes match what real sensors produce, from noise characteristics to spectral response.

We are also honest about what synthetic data is. It is a tool, not a shortcut around reality, and its value depends entirely on its fidelity. Low-quality synthetic data trains low-quality models. That is why we treat validation against real ground truth as part of the process rather than an afterthought, and why the strongest pipelines usually blend synthetic and real imagery. The aim is never to convince anyone that synthetic data is magic. It is to show, with measurement, that a model has actually learned what it needs to.

Where is this actually useful?

By the end of each conversation, the questions usually turned practical. The short answer is that synthetic Earth observation data helps anywhere you need a robust model but lack enough labelled real imagery for the conditions that matter most, which turns out to span across far more fields than people expect.

In agriculture, that means training models to spot stress, disease, or specific crop types earlier and across more varied conditions than the archive offers. In insurance and finance, it supports risk and damage assessment for events that are rare by nature. In the maritime domain, it helps detect and classify vessels, including the unusual cases that real datasets barely contain. For disaster response, it lets teams prepare models for floods, fires, and structural damage before those events happen, rather than scrambling for data during them.

The same logic runs through defence and security, infrastructure monitoring, and climate work. Wherever the limiting factor is data rather than ambition, synthetic imagery widens what is possible: more edge cases, more balanced training sets, and reliable testing for situations that are dangerous, rare, or simply absent from what has already been captured. That is the thread connecting every use we discussed in June.

If you asked us one of these questions over the past month, thank you for the conversation. If you have a sixth that we did not cover here, that is the one we would most like to hear next.

Leave a Reply

Your email address will not be published. Required fields are marked *