Global · 2024
The best weather model doesn't give you one forecast. It gives you fifty.
GenCast beat the world's leading forecast system on 97% of tests — by refusing to answer with a single number.
Published 15 August 2026
What happened
Weather forecasting is the oldest serious prediction problem there is, and the benchmark to beat is ENS, the ensemble system run by the European Centre for Medium-Range Weather Forecasts [1][2].
In December 2024, Google DeepMind published GenCast in Nature. Against ENS it was more accurate on 97.2% of 1,320 tested forecast targets, and on 99.8% of them at lead times beyond 36 hours [1].
The interesting part is how it answers. An earlier model, GraphCast, produced a single best estimate of future weather. GenCast instead produces an ensemble of 50 or more predictions, each a possible trajectory — so the output is not "14°C and raining" but a distribution over what might happen [1].
That distribution is what makes it useful for the things people actually need forecasts for: extreme heat, high winds, the track of a tropical cyclone. A decision about whether to evacuate depends on how likely the bad case is, not on the single most likely case [1].
It is also fast. A 15-day global forecast takes about eight minutes on a single TPU, against hours on a supercomputer for the physics-based systems, and it was trained on four decades of ECMWF's own reanalysis archive up to 2018, then evaluated on 2019 [1].
Its authors and outside experts are clear about the limits: it underpredicts the intensity of cyclones, struggles in the upper troposphere, and is trained on a past climate that is not the one coming. One meteorologist's point is worth keeping — humans still distil this into the forecast anyone acts on [2].
Where it helped
This is the shape of AI worth wanting: it beat a system decades of physics and public investment had produced, and it did so while giving MORE information about its own uncertainty rather than less. The fast, cheap ensemble is a genuine public good — a national met service without a supercomputer budget can now run one. And it stayed honest about where it is weak, which is the part that makes the rest usable [1][2].
Where it can still burn
The danger is downstream, in what happens as the forecast travels. A distribution over 50 outcomes gets collapsed into a single number by the time it reaches an app, a headline or a meeting — the range is the first thing dropped, because ranges are awkward to display and a single number sounds more authoritative. The model did the hard part and said how sure it was; the reporting chain deletes it. Every wrong-sounding forecast you have complained about was probably a 30% chance that happened, presented as a promise that it would not [1][2].
The tell
When a prediction arrives as one number, ask what the range was. If nobody can tell you, you have been handed a summary, not a forecast — and the uncertainty was deleted by someone, not by nature.
This is the flip side of reading a confidence score: often no score is shown at all, because the range got dropped somewhere between the model and you. Sales projections, delivery dates, cost estimates and risk scores all come out of systems that internally know how uncertain they are, and all arrive as a single confident figure. Asking for the spread costs one question, and it changes what the number means: 40 days is a plan, "30 to 70 days" is the truth, and only one of those tells you whether to promise a customer anything.
Share this case
The image has the link printed on it, so it still leads back here.
The check is a habit, and habits are trained. Statistics, understood is about distributions, spread and confidence — the difference between a number and a number you can act on.
Sources
Every source below was opened and read. Last verified 15 August 2026.
- [1] GenCast predicts weather and the risks of extreme conditions with state-of-the-art accuracy — Google DeepMind (announcing the Nature paper), 4 December 2024
- [2] Google DeepMind's new AI model is the best yet at weather forecasting — Scott J Mulligan, MIT Technology Review, 4 December 2024