AI-based weather forecasting models are rapidly being adopted because they are faster, cheaper and require far less computational infrastructure than traditional physics-based models. However, researchers warn these systems handle the growing frequency of previously rare climate extremes poorly. They call such failures "grey swans": physically possible but so rare in historical data that machine-learning models have little or no example to learn from.
Why "grey swans" are a problem
AI systems are trained on past observations and simulated data. When certain extreme events are essentially absent from training sets, AI struggles to generalize to them. This is a growing concern because climate change is producing more "first-of-their-kind" extremes that were not present in historical records.
A historical reference point is the Bhola cyclone: on 12 November 1970 the storm struck the coast of what was then East Pakistan, producing sustained winds of 205 km/h and a storm surge of 10.5 metres, and is estimated to have killed between 300,000 and 500,000 people. The scale of its impact was shaped in part by the limited forecasting and warning systems of the time.
Physics-based models versus AI
Numerical weather-prediction models built on physical laws can simulate very rare events, even if they mark them as statistically unlikely, because they are constrained by the underlying dynamics of the atmosphere and oceans. Machine-learning models, by contrast, rely on examples; if a type of extreme event is missing from the training data, they typically cannot extrapolate correctly.
Pedram Hassanzadeh, an associate professor of geophysical sciences at the University of Chicago, and colleagues published an experiment last April in which they removed category 3–5 hurricanes from an AI model's training set and then tested it on category 5 storms. The AI models failed to forecast those previously unseen high-intensity events accurately — the failure stemmed from a lack of extrapolative ability.
Silent, dangerous failures
Researchers also emphasise the risk of "silent" failures: AI can output confident, routine forecasts even while an unprecedented event is unfolding. A computer science and engineering associate professor at the University of California noted that AI systems can subtly violate conservation laws in ways that common performance metrics do not reveal. When errors occur, they can be harder to diagnose because AI decisions are typically less interpretable.
Institutional risks matter too. If meteorological services switch too fast to AI and let physics-based infrastructure degrade, the redundancy that currently helps catch AI errors could be lost. AI systems also depend on stable observation networks and are vulnerable to pressures on satellite programmes.
Why meteorologists are adopting AI anyway
Despite these risks, many meteorologists deploy AI because it brings clear operational advantages: lower cost, faster runtimes, and increasingly competitive skill for typical weather patterns. Andrew Charlton-Perez, professor of meteorology at the University of Reading, notes that while the best physics-based models improve slowly (roughly a day of additional accuracy per decade), machine-learning models have been improving at a faster rate and now compete with the top physical models.
Practical results have reinforced this shift: during the 2025 Atlantic hurricane season, the Google DeepMind model outperformed many physical models in forecasting storm tracks and intensities. Since about 2023, leading AI systems such as GraphCast, Pangu-Weather and the ECMWF AIFS have reached or exceeded the best physics-based models on mid-range forecast metrics.
AI forecasts are especially valuable where traditional forecasting resources are lacking — often the same regions most exposed to climate impacts. Hassanzadeh led an initiative that provided AI-based monsoon forecasts to 38 million Indian farmers, predicting the onset of the rainy season up to four weeks ahead.
Calls for stricter testing and hybrid approaches
Given the potential consequences, several researchers call for more stringent testing before widespread operational deployment. Shruti Nath, a postdoctoral researcher at the University of Oxford, co-authored an editorial proposing a testing framework that deliberately withholds certain "iconic" extreme events from training data and reserves them for testing. An example would be the 2021 north-west Pacific heatwave, an event that would have been virtually impossible without climate change.
The proposed AIRWIE (AI Retraining Without Iconic Events) protocol would require the meteorological community to agree which high‑impact events serve as strict benchmarks — a challenging but widely supported goal among researchers who see urgent need for such tests.
Other teams, including Hassanzadeh's, are exploring ways to improve AI's extrapolation by combining machine learning with "relevant sampling" techniques that artificially generate rare events for training.
Conclusion: no turning back, but proceed carefully
AI is already reshaping weather prediction, and as the climate becomes more unpredictable forecasters will need every tool available. Understanding and addressing AI's limitations is essential: the aim is to make AI models physically consistent, well calibrated and resilient to changing distributions. Rejecting AI outright because of the grey‑swan problem would forfeit a major generational advance in forecasting; the challenge is to integrate AI with physics-based systems while preserving redundancy, interpretability and public safety.


