What happened
Weather forecasting is the oldest serious prediction problem there is. The system to beat is ENS, run by the European Centre for Medium-Range Weather Forecasts — the world's leading forecaster [1][2].
In December 2024, Google DeepMind published GenCast in Nature. Against ENS it was more accurate on 97.2% of 1,320 tested forecast targets, and on 99.8% of them for forecasts more than 36 hours ahead [1].
The interesting part is how it answers. An earlier model, GraphCast, produced a single best guess at future weather. GenCast instead produces an ensemble — a bundle of 50 or more forecasts, each one a way the weather might go. So the output is not "14°C and raining" but a spread of what might happen [1].
That spread is what makes it useful for the things people actually need forecasts for: extreme heat, high winds, the path of a tropical cyclone. A decision about whether to evacuate depends on how likely the bad case is, not on the single most likely case [1].
It is also fast. A 15-day global forecast takes about eight minutes on a single specialised chip, against hours on a supercomputer for the traditional physics-based systems. It was trained on four decades of ECMWF's own weather records up to 2018, then tested on 2019 [1].
Its authors and outside experts are clear about the limits. It underpredicts the strength of cyclones, struggles high in the atmosphere, and is trained on a past climate that is not the one coming. One meteorologist's point is worth keeping: humans still turn all this into the forecast anyone acts on [2].
Where it helped
This is the shape of AI worth wanting. It beat a system that decades of physics and public investment had produced, and it did so while giving MORE information about its own uncertainty rather than less. The fast, cheap ensemble is a genuine public good — a national weather service without a supercomputer budget can now run one. And it stayed honest about where it is weak, which is the part that makes the rest usable [1][2].
Where it can still burn
The danger is downstream, in what happens as the forecast travels. A spread of 50 outcomes gets squashed into a single number by the time it reaches an app, a headline or a meeting. The range is the first thing dropped, because ranges are awkward to display and a single number sounds more authoritative. The model did the hard part and said how sure it was; the chain of people passing it on deleted that. Every wrong-sounding forecast you have complained about was probably a 30% chance that happened, presented as a promise that it would not [1][2].
The tell
When a prediction arrives as one number, ask what the range was. If nobody can tell you, you have been handed a summary, not a forecast — and someone, not nature, deleted the uncertainty.
This is the flip side of reading a confidence score. Often no score is shown at all, because the range got dropped somewhere between the model and you. Sales projections, delivery dates, cost estimates and risk scores all come out of systems that know, inside, how unsure they are — and all arrive as a single confident figure. Asking for the spread costs one question, and it changes what the number means. "40 days" is a plan. "30 to 70 days" is the truth. Only one of those tells you whether to promise a customer anything — the difference between a builder saying "six weeks" and "six to ten weeks, depending on the weather".
The check is a habit, and habits are trained. Statistics, understood is about distributions, spread and confidence — the difference between a number and a number you can act on.
Sources
Every source below was opened and read. Last verified 15 August 2026.
- [1] GenCast predicts weather and the risks of extreme conditions with state-of-the-art accuracy — Google DeepMind (announcing the Nature paper), 4 December 2024
- [2] Google DeepMind's new AI model is the best yet at weather forecasting — Scott J Mulligan, MIT Technology Review, 4 December 2024