Google's AI hurricane model buys forecasters an extra day
Tested live by the US National Hurricane Center in 2025 and open-sourced to run in under a minute, the ensemble matches prior two-day accuracy at three days — but its coarse-resolution strength is still a black box and physics models beat it on record-breaking extremes.
What happenedGoogle DeepMind's WeatherNext Cyclones, published in Nature on Aug 6, 2026, ran live as an experimental forecast at the US National Hurricane Center in 2025 and showed three-day forecasts for path, strength and size as accurate as leading models were at two days.
Why it mattersAn extra day extends time to evacuate, close ports and stage supplies, historically cutting damage by about 30% and yielding up to $9 per $1 spent on early warnings, with the biggest stakes for small island states.
Still openWhether its coarse-resolution skill holds for unprecedented, climate-driven extremes — physics models still beat AI on record-breaking events and the system misses about half of rapid intensifications, with larger ensembles raising false-alarm costs.

An AI ensemble model for hurricanes and typhoons — tropical cyclones, the same warm-ocean storm by different regional names — that predicts path, strength and size together was published in Nature on Aug 6, 2026, and it buys forecasters more than a day. Google DeepMind and Google Research's WeatherNext Cyclones (WN-C) shows its three-day forecasts are now as accurate as leading operational models were at two days, run live by the US National Hurricane Center (NHC), the US agency that issues official hurricane forecasts for the Atlantic and eastern Pacific, during the 2025 season and open-sourced to run a 15-day global ensemble in under a minute on a single TPU (AI chip). That extra day would matter — an extra 24 hours to evacuate, stage supplies and close ports can cut damage by about 30% — but whether the gain holds for the next unprecedented, climate-driven storm is not settled: at coarse resolution the model is empirically strong yet physically unexplained, and physics-based models still beat AI on record-breaking extremes.
How a 3-day AI forecast now matches what physics models could do in 2 days. Mean track error (n mi) and intensity error (kt) versus forecast lead time for WeatherNext Cyclones against ECMWF-ENS, HAFS, and GenCast on global tropical cyclones in 2023–2025 (curves digitized from Nature / DeepMind Fig. 2), with Atlantic 2025 NHC verification for GDMI and the HCCA consensus as open markers. At three days, WeatherNext’s track and intensity errors sit near what the leading physics baselines only achieve about a day earlier—the “extra day” of useful lead time reported in the paper. Global multi-year curves are not identical to one-season Atlantic NHC scores; track values from the Nature figure were converted from km using 1 n mi = 1.852 km. — AI-assisted analytic, built only from real cited or sourced data. Source: Nature, Google DeepMind, NOAA/NWS National Hurricane Center. As of 2026-08-08.
Tropical cyclones are the costliest disaster driver on record — more than 700,000 deaths and $1.4 trillion in losses over 50 years, and 11,778 weather disasters causing just over 2 million deaths and $4.3 trillion between 1970 and 2021 — and the NHC, as the World Meteorological Organization's regional center, provides guidance to nearly 30 countries. Until now, forecasting forced a split: coarse global models steered track, the path set by large-scale winds, while specialized high-resolution regional models tried to capture intensity, how strong the storm gets, driven by small-scale thermodynamics around the eyewall. WeatherNext collapses that split into one model for track, intensity and size — the extent of damaging winds — with an average lead-time advantage of a day or more evaluated on worldwide cyclones from 2023 to 2025, a jump the authors frame as roughly a decade of operational progress.
How one model learned to do two jobs
The gain comes from co-training and a way of making ensembles coherent.
WeatherNext was trained end-to-end on two very different datasets at once: about 20 terabytes of global atmospheric analysis — ERA5, the European reanalysis that reconstructs the past atmosphere, and HRES, the European Centre for Medium-Range Weather Forecasts' high-resolution operational analysis — plus the IBTrACS archive of about 5,000 historical tropical cyclones. That pairing lets the model learn both general weather dynamics and the rare patterns of extreme cyclones.
The architecture is called Functional Generative Networks, or FGN. Instead of adding per-pixel noise, FGN samples a single 32-number noise vector that modulates the whole field at once. Because the same 32 numbers shape the entire globe, each ensemble member is a globally coherent weather scenario — a functional perturbation — rather than speckled noise. Four independently trained versions of the model, each with about 180 million parameters, cover what the model doesn't know, while the 32-number perturbations cover the inherent randomness of weather.
Training rewards the model when the range of outcomes it predicts at each location matches reality, first on single steps then on multi-day rollouts. Yet it learns joint structure anyway. The 32-degree-of-freedom manifold is so constrained that the easiest way to jointly optimize everywhere is to encode physically consistent spatial and cross-variable correlations. Even a stripped-down single-seed, no-rollout, smaller version still beats prior state-of-the-art, and the full model improves derived variables like 10-metre wind that were never directly supervised.
The coarse-resolution surprise
That is why meteorologists were shocked by what happened at low resolution.
Intensity forecasting was assumed to need kilometre-scale resolution to resolve eyewall thermodynamics. WeatherNext achieves state-of-the-art intensity at 0.25 degrees — about 28 by 28 kilometres, roughly 100 times coarser than traditional regional intensity models — and a lightweight mini version at 1 degree, about 111 by 111 kilometres, still performs well.
The authors are explicit that they do not understand how. "When we told the community that our model was only using relatively coarse resolution, they were shocked," lead author Ferran Alet said, adding "It's a black box at the end of the day, but that gives physicists a signal that something is happening that was not previously understood." The paper calls it an open research question and invites investigation. The implication, if the signal is real, is that coarse global analyses already contained more intensity-relevant information than physics models could extract — and that past compute spent chasing higher resolution may have been inefficiently extracting information that was already there. It does not mean the physics is learned, and it does not show the skill generalizes to unprecedented extremes.
From TPU to Colab
Speed is what makes the ensemble size practical, and ensemble size is what makes tail risks visible.
A 15-day global ensemble runs in under a minute on one TPU (AI chip), compared with hours on supercomputers for physics models. WeatherNext 2, the operational update, is described as eight times faster than the prior generation and better on 99.9% of variables and lead times. In 2024 the system ran 50 members, matching conventional physics ensembles; in 2025-2026 it scaled to 1,000 members. Rapid intensification — a jump of at least 35 mph in 24 hours that can turn a weak system into a major hurricane overnight — and landfall are tail events; larger ensembles sample the distribution more fully.
Alongside the Nature paper, Google open-sourced code and weights for WeatherNext Cyclones and WeatherNext 2 and the mini under Apache 2.0 for code and CC BY 4.0 for materials, with a Colab notebook that runs the mini on a free v5e-1 TPU (AI chip) and data feeds via Google Cloud, Earth Engine, BigQuery and Weather Lab. The full models need an H100 GPU or v5p TPU (AI chip) for memory; the mini runs on a P100. Agencies need not run the model at all to benefit — forecasts are viewable on Weather Lab — but ERA5/HRES licensing, data ingest and verification still constrain under-resourced services. Three checkpoints trained through 2022, 2023 and 2024 are provided so results for 2023-2025 can be reproduced.
Where Hurricane Melissa hit Jamaica after the AI flagged a Category 5 landfall five days ahead. Map shows Hurricane Melissa’s NHC best-track path through the Caribbean and its Category 5 landfall near New Hope, Westmoreland Parish, Jamaica on 28 October 2025 (18.1°N, 78.0°W; 160 kt in the ATCF best track). Square and diamond markers are NHC forecast positions for Jamaica about five days (Advisory 10 outlook 28/1800Z at 18.0°N, 78.4°W) and three days (Advisory 17, 28/1200Z inland Jamaica at 17.9°N, 77.2°W) before landfall; Google DeepMind reports WeatherNext assigned 80% confidence of Category 5 Jamaica landfall at five days and near 100% at three days, while some traditional guidance still allowed a weaker Haiti path. Track positions: National Hurricane Center ATCF best track forecast points: NHC Forecast/Advisories 10 and 17; WeatherNext confidence: — AI-assisted analytic, built only from real cited or sourced data. Source: Google DeepMind, NOAA. As of 2026-08-08.
The live test: 2025 and Hurricane Melissa
Retrospective skill often fades in operations. In 2025 it did not, according to independent verification.
2025 was the first year the NHC incorporated AI guidance operationally. WeatherNext ran live as an experimental forecast labelled FNV3, shown by the NHC as GDMI — the NHC's label for the Google DeepMind model —. The NHC Forecast Verification Report of March 30, 2026, found Atlantic mean track errors of 20 nautical miles at 12 hours to 162 at 120 hours, below five-year means at all lead times and up to 14% smaller at 96 hours, with GDMI slightly outperforming the official forecast at 12-72 hours. Intensity was harder — errors exceeded five-year means at all times, 40-50% larger at 60-120 hours, with the climatology-persistence baseline — a simple baseline that assumes the storm keeps doing what it has been doing — up to 90% larger, reflecting an unusually difficult rapid-intensification year — but the official forecast beat all models at 12 and 96 hours and was comparable to GDMI, the best model, at other times. In the eastern North Pacific, the NHC set records for track accuracy at 24-120 hours, with mean errors 19 to 100 nautical miles, 15-30% below five-year means, and GDMI beat the official forecast and all consensus aids (averages of several models) at 48-120 hours. GDMI was not available for the first two Atlantic and first five East Pacific storms, limiting head-to-head samples.
Hurricane Melissa made the abstract concrete. Forming in late October 2025, Melissa rapidly intensified from 70 to 130 mph in 24 hours — a 115-mph and 90-millibar drop in 72 hours ending Oct 28 — and made landfall near New Hope in Westmoreland Parish, Jamaica, on Oct 28 as a Category 5 with 185-190 mph sustained winds, the strongest Jamaica landfall on record and tied for strongest Atlantic landfall, causing about 95 deaths and about $8.8 billion in damage, the costliest in Jamaican history. Traditional models split between a weak Haiti landfall and a major Jamaica hit. WeatherNext predicted a Category 5 Jamaica landfall five days in advance with 80% confidence, rising to near 100% at three days, helping the NHC issue its first-ever Category 5 forecast starting from Category 1 initial intensity. Four days before landfall the NHC track was within about 13 miles of western Jamaica. NOAA confirmed the NHC outperformed every model at nearly every lead time for Melissa and provided almost three days' notice of Category 5 landfall from low initial intensity. The storm still devastated Jamaica — extra lead time aids but does not eliminate impact — and Melissa is one case, not proof of universal skill.
Including WeatherNext in a weighted-average consensus substantially improves that consensus. In a two-way blend, it received 75% of the weight for latitude, 89% for longitude and 42% for intensity; because the existing consensus itself averages five to eight models, its effective weight versus any single traditional model is even larger.
Where it still fails
The same report card that shows a step-change also shows why forecasters call it another tool, not a replacement.
Rapid intensification remains the hardest problem. A standard metric balancing correct detections against false alarms and misses improved from below 0.3 to 0.5 with WeatherNext — real progress that still means the model gets rapid intensification wrong about half the time. Larger 1,000-member ensembles improve tail sampling but raise over-warning costs: evacuations, port and airport closures and staged supplies cost money and trust if the tail scenario does not materialize. For genesis, the trade-off is explicit: the European model detects only about 20% of formations with the lowest false alarms, while other models detect more but cry wolf more often; when two or more models agree, odds rise considerably.
Failure modes are not hypothetical. Hurricane Milton in 2024 was tracked accurately by Google AI but its intensity was forecast below Category 2 versus an actual Category 5 — the most intense Gulf storm since Rita in 2005 — after going from formation to Category 5 in about two days. Nearly 80% of major hurricanes undergo rapid intensification, making intensity the hardest dimension for both AI and physics models.
Out-of-distribution performance is the deeper concern. In peer-reviewed tests of record-breaking heat, cold and wind events exceeding 1979-2017 records, the physics-based ECMWF HRES — the European Centre for Medium-Range Weather Forecasts' high-resolution physics model — consistently outperformed leading AI models GraphCast, Pangu-Weather and FuXi; AI models systematically underestimated both intensity and frequency of records, with underestimation growing the further the event exceeded the training record. The authors attribute this to neural networks' inability to extrapolate beyond their training domain, while physics laws still apply. Training on ERA5 at about 31km, which itself underestimates intensity and cannot resolve mesoscale eyewall structures, bakes in that ceiling. Even a hybrid regional AI system trained on 9km high-resolution typhoon reanalysis and constrained by AI still underestimates super-typhoon intensity by about 10-15 metres per second and shows limited skill above 50 m/s.
Operationally, consensus still wins. In 2024 the NHC official human-in-loop forecast outperformed all individual models at almost all lead times for track, and for intensity except at 4-5 days. As NHC Director Mike Brennan put it, "There's no guarantee that one model, because it did well last year or for this particular storm, is necessarily going to be the best model for the next season or the next storm," and "A hurricane is not just a track or intensity forecast — it requires experts to translate that into what the impacts are going to be — and it's the impacts that kill people." Official warnings remain solely with national meteorological authorities; the model is research code provided as-is without warranty.
Who an extra day helps — if trusted
If the gain is acted upon, the stakes are large and unevenly distributed. An extra 24 hours extends decisions that are time-critical: evacuations, hospital preparation, port and airport closures, and pre-positioning of search and rescue.
The economics cited are structural, not measured for WeatherNext itself. The World Meteorological Organization and the Climate Risk and Early Warning Systems initiative estimate every $1 invested in early warning systems can yield up to $10 in benefits, and that 24 hours' warning of a coming storm can cut damage by 30%. The Global Commission on Adaptation reports a 1:9 cost-benefit ratio and the WMO Bulletin cites more than tenfold return, with an $800 million investment in developing countries potentially avoiding $3-16 billion in annual losses. Those figures sit against a backdrop where storms are the leading cause of economic losses and over 90% of deaths occur in developing countries, with small island states sometimes losing more than 100% of GDP in a single disaster.
Open weights and cloud feeds directly address that capacity gap for the Caribbean, Pacific Islands and South Asia, and collaborations are expanding with the national weather services of the Philippines (PAGASA), Taiwan (CWA), Indonesia (BMKG) and Vietnam (VNMHA), with Japan, Australia and India planned. But the barrier is not zero: full models need H100/V5p hardware, data ingest and licensing remain, and verification and liability frameworks lag. The NHC verification shows what is possible when a well-resourced center runs the model inside a human consensus; whether the same gain transfers to every under-resourced service, and whether false alarms erode trust before the next unprecedented storm, remains untested beyond 2025.
Source recordSources / claims / limits
How this piece is framed: An extra day, bought coarsely and fast — how a single AI ensemble passed its first operational test, and where it still fails
Charts & tables — AI-assisted; provenance on each line
- An extra day bought: AI cuts forecast error by a day's lead time — sourced for this figure · as of 2026-08-08
- Hurricane Melissa's track to Jamaica and the five-day AI forecast — sourced for this figure · as of 2026-08-08
Sources
- (primary) WeatherNext 2: Our most advanced weather forecasting model — Google — https://blog.google/innovation-and-ai/models-and-research/google-deepmind/weathernext-2 · read in full · captured 2026-08-08
- (primary) Skillful joint probabilistic weather forecasting from marginals (arXiv:2506.10772v1) — https://arxiv.org/html/2506.10772v1 · read in full · captured 2026-08-08
- (primary) Google's new AI could give hurricane forecasters an extra day of warning — Straight Arrow News — https://san.com/cc/googles-new-ai-could-give-hurricane-forecasters-an-extra-day-of-warning · read in full · captured 2026-08-08
- (primary) Operational Tropical Cyclone Forecasting with AI — Nature — https://www.nature.com/articles/s41586-026-10953-2 · read in full · captured 2026-08-08
- (primary) How WeatherNext helped the National Hurricane Center better predict Hurricane Melissae historic landfall in Jamaica — Google DeepMind — https://deepmind.google/blog/how-weathernext-helped-the-national-hurricane-center-better-predict-hurricane-melissas-historic-landfall-in-jamaica · read in full · captured 2026-08-08
- (primary) A regional artificial intelligence model for skillful typhoon prediction — Nature — https://www.nature.com/articles/s44304-026-00219-2 · read in full · captured 2026-08-08
- (primary) WeatherNext: AI model achieves breakthrough in forecasting cyclones — Google DeepMind — https://deepmind.google/blog/weathernext-ai-model-achieves-breakthrough-in-forecasting-cyclones/ · read in full · captured 2026-08-08
- (primary) google-deepmind/weathernext GitHub repository — Google DeepMind GitHub — https://github.com/google-deepmind/weathernext · read in full · captured 2026-08-08
- (primary) WMO Bulletin: Early Warning and Anticipatory Action — World Meteorological Organization — https://wmo.int/media/news/wmo-bulletin-early-warning-and-anticipatory-action · read in full · captured 2026-08-08
- (primary) National Hurricane Center Forecast Verification Report 2025 Hurricane Season (30 March 2026) — NOAA/NWS National Hurricane Center — https://www.nhc.noaa.gov/verification/pdfs/Verification_2025.pdf · read in full · captured 2026-08-08
- (primary) The Triple Dividends of Early Warning Systems and Climate Services - WMO Magazine — World Meteorological Organization — https://wmo.int/media/magazine-article/triple-dividends-of-early-warning-systems-and-climate-services · read in full · captured 2026-08-08
- (primary) Skillful joint probabilistic weather forecasting from marginals (arXiv:2506.10772) — arXiv / Google DeepMind — https://arxiv.org/abs/2506.10772 · read in full · captured 2026-08-08
- (primary) Which hurricane models should you trust in 2025? — Yale Climate Connections — https://yaleclimateconnections.org/2025/05/which-hurricane-models-should-you-trust-in-2025 · read in full · captured 2026-08-08
- (primary) DeepMind Says Its AI Can Predict Hurricanes Earlier Than Everyone Else — WIRED — https://www.wired.com/story/deepmind-ai-model-can-predict-hurricanes-earlier · read in full · captured 2026-08-08
- (primary) 2025 Atlantic hurricane season marked by striking contrasts — NOAA — https://www.noaa.gov/news-release/2025-atlantic-hurricane-season-marked-by-striking-contrasts · read in full · captured 2026-08-08
- (primary) Physics-based Weather Models More Reliable Than AI for Extreme Events — KIT — https://www.kit.edu/kit/english/pi_2026_040_physics-based-weather-models-more-reliable-than-ai-for-extreme-events.php · read in full · captured 2026-08-08
- (primary) Economic costs of weather-related disasters soars but early warnings save lives — WMO — https://wmo.int/media/news/economic-costs-of-weather-related-disasters-soars-early-warnings-save-lives · read in full · captured 2026-08-08
- (secondary) Traditional models still 'outperform AI' for extreme weather forecasts — Carbon Brief — https://www.carbonbrief.org/traditional-models-still-outperform-ai-for-extreme-weather-forecasts · read in full · captured 2026-08-08
- (secondary) Sand and snow: Shared financial risk across climate extremes — Global Reinsurance — https://www.globalreinsurance.com/home/sand-and-snow-shared-financial-risk-across-climate-extremes/1459302.article · read in full · captured 2026-08-08
- (secondary) How AI Is Improving Tropical Cyclone Forecasting — Earth.Org — https://earth.org/how-ai-is-improving-tropical-cyclone-forecasting-in-climate-change-era · read in full · captured 2026-08-08
- (secondary) Google DeepMinde WeatherNext 2 Uses Functional Generative Networks For 8x Faster Probabilistic Weather Forecasts — MarkTechPost — https://www.marktechpost.com/2025/11/17/google-deepminds-weathernext-2-uses-functional-generative-networks-for-8x-faster-probabilistic-weather-forecasts · full text not obtained — used its summary
Claims, and how far we tracked each down
- [confirmed] WeatherNext Cyclones paper 'Operational Tropical Cyclone Forecasting with AI' was published in Nature on Aug 6, 2026 (DOI 10.1038/s41586-026-10953-2). · read in full (as of 2026-08-08)
- [confirmed] WeatherNext Cyclones is a single AI model that predicts tropical cyclone track, intensity and size/wind structure together with state-of-the-art accuracy. · read in full (as of 2026-08-08)
- [confirmed] Evaluated on tropical cyclones from 2023025, WeatherNext Cyclones offers an average of a day or more of lead-time advantage over leading operational models; its 3-day forecasts are as accurate as prior models' 2-day forecasts. · read in full (as of 2026-08-08)
- [confirmed] The lead-time gain is described as comparable to roughly a decade of operational meteorological progress. · read in full (as of 2026-08-08)
- [confirmed] Model was co-trained end-to-end on ~20 terabytes of global atmospheric analysis data (ERA5/HRES) and the IBTrACS historical database of ~5,000 tropical cyclones. · read in full (as of 2026-08-08)
- [confirmed] Architecture uses Functional Generative Networks (FGNs): graph neural network encoder/decoder on a 6x-refined icosahedral mesh with graph transformer, ~180M parameters per seed, 768 latent dim, 24 layers, 32-dim Gaussian noise injected via conditional normalization to sample coherent ensemble members. · read in full (as of 2026-08-08)
- [confirmed] WeatherNext Cyclones operates at 0.25 (~28x28 km) resolution, about 100x coarser than traditional high-resolution regional intensity models; WeatherNext 2-mini operates at 1 (~111x111 km) and still shows strong performance. · read in full (as of 2026-08-08)
- [confirmed] A single 15-day global forecast runs in under one minute on a single TPU; WeatherNext 2 is 8x faster than prior WeatherNext Gen and surpasses it on 99.9% of variables and lead times (0-15 days) with ~6.5% average CRPS improvement. · read in full (as of 2026-08-08)
- [confirmed] Ensemble size was 50 members in 2024 (matching conventional physics ensembles) and scaled to 1,000 members in 2025-2026 to better capture rare tail events like rapid intensification. · read in full (as of 2026-08-08)
- [confirmed] Code and model weights for WeatherNext Cyclones, WeatherNext 2 and WeatherNext 2-mini were open-sourced Aug 6, 2026 on GitHub under Apache 2.0 (code) / CC BY 4.0 (materials), with Colab notebook for mini on free v5e-1 TPU and data feeds via Google Cloud, Weather Lab and OpenMeteo. · read in full (as of 2026-08-08)
- [confirmed] 2025 was the first year NOAA's National Hurricane Center incorporated AI model guidance operationally; WeatherNext ran live as FNV3 (NHC post-processed as GDMI) and was the top-performing individual model for track and intensity in NHC's 2025 verification. · read in full (as of 2026-08-08)
- [confirmed] During Hurricane Melissa (Oct 2025), WeatherNext predicted Category 5 landfall in Jamaica 5 days in advance with 80% confidence (near 100% at 3 days), helping NHC issue the first-ever Category 5 forecast from Category 1 initial intensity. · read in full (as of 2026-08-08)
- [confirmed] Hurricane Melissa made landfall near New Hope, Westmoreland Parish, Jamaica on Oct 28, 2025 as a Category 5 with ~185-190 mph sustained winds, the strongest Jamaica landfall on record and tied for strongest Atlantic landfall, with ~95 deaths and ~$8.8B damage (costliest in Jamaican history). · read in full (as of 2026-08-08)
- [confirmed] FGN generates ensembles by passing a single 32-dimensional global noise vector into all conditional layer-norm layers, with learned parameters shared across mesh and grid spatial dimensions, so sampling different vectors per ensemble member produces globally coherent functional perturbations rather than per-pixel noise. · read in full (as of 2026-08-08)
- [confirmed] FGN models epistemic uncertainty with deep ensembles of 4 independently-initialized models (each 180M parameters) and aleatoric uncertainty with learned functional perturbations where forecast samples come from networks whose parameters are sampled from a learned distribution per ensemble member and timestep. · read in full (as of 2026-08-08)
- [confirmed] FGN is trained to minimize the fair (unbiased) Continuous Ranked Probability Score (CRPS) on per-location marginals averaged over all locations, variables and levels, initially with single-step loss then finetuned with autoregressive rollouts up to 8 steps with gradients propagated through the rollout. · read in full (as of 2026-08-08)
- [confirmed] Despite marginal-only CRPS supervision, FGN learns skillful joint spatial and cross-variable structure because the 32-degree-of-freedom globally-shared perturbation heavily constrains the output manifold; the easiest way to jointly optimize CRPS everywhere is to encode physically consistent correlations. Even a stripped-down single-seed, no-autoregressive, smaller (512 dim, 16 layer) FGN still outperforms prior state-of-the-art across the board, and full FGN improves pooled CRPS (8.7% avg-pooled, 7.5% max-pooled) and derived variables like 10m wind speed and z300-z500 thickness. · read in full (as of 2026-08-08)
- [confirmed] WMO primary sources state early warning systems provide more than tenfold return on investment: Global Commission on Adaptation reports cost-benefit ratio of 1:9 (US$1 invested yields US$9 net benefits), and WMO Bulletin states 'more than a tenfold return'; providing just 24 hours' notice of a storm or heatwave can reduce potential damage by 30%; US$800 million investment in such systems in developing countries could prevent US$3-16 billion in annual losses. · read in full (as of 2026-08-08)
- [confirmed] NHC 2025 Verification Report (30 March 2026) independent Atlantic verification: mean official track errors 20 n mi at 12h to 162 n mi at 120h, below 5-year means at all lead times (up to 14% smaller at 96h); official forecasts outperformed consensus aids TVCA, HCCA, FSSE, but Google DeepMind ensemble mean (GDMI) slightly outperformed NHC at 12-72h short-range. Intensity errors exceeded 5-year means at all times, 40-50% larger at 60-120h, with Decay-SHIFOR baseline up to 90% larger, indicating unusually difficult intensity forecasting; official forecasts beat all models at 12h and 96h and were comparable to GDMI (best model) at other times. · read in full (as of 2026-08-08)
- [confirmed] NHC 2025 Verification Report eastern North Pacific: mean NHC track errors 19 n mi at 12h to 100 n mi at 120h, 15-30% lower than 5-year means, breaking records for accuracy at 24-120h; GDMI beat official and all consensus aids at 48-120h. Intensity errors 20-30% lower than 5-year means at 48-120h with record accuracy at 48h and 72h; official comparable to best consensus aids HCCA and NNIC. · read in full (as of 2026-08-08)
- [confirmed] NOAA primary confirms Hurricane Melissa operational value: NHC intensity forecasts for Melissa outperformed every model at nearly every lead time and provided almost three days advance notice of Category 5 Jamaica landfall from low initial intensity; 4 days before landfall NHC projected path over western Jamaica within ~13 miles; Melissa underwent rapid intensification with 115-mph wind increase and 90 mb pressure drop in 72h ending 11 a.m. EDT Oct 28, 2025. · read in full (as of 2026-08-08)
- [confirmed] Nature paper abstract confirms WeatherNext Cyclones evaluated on 2023-2025 cyclones offers average >1 day lead-time advantage over leading operational models for track, intensity and wind radii, comparable to last decade of operational progress, using inputs orders of magnitude coarser than regional models, and that including WN-C in weighted-average consensus substantially improves skill with scalability to 1,000-member ensembles. · read in full (as of 2026-08-08)
- [contested] The coarse-resolution intensity skill (28km, 100x coarser than regional models) is an unexplained black-box result that shocked the meteorological community; DeepMind authors state they do not fully understand how the model extracts intensity signal from coarse inputs, and describe it as "a black box at the end of the day." This qualifies the success narrative that coarse data contains sufficient signal the mechanism is not established and invites skepticism about physical consistency. · read in full (as of 2026-08-08)
- [confirmed] In a peer-reviewed out-of-distribution test of record-breaking extremes (heat/cold/wind records from 2018/2020 exceeding 1979-2017 records), the physics-based ECMWF HRES consistently outperformed leading AI models GraphCast, Pangu-Weather and FuXi; AI models systematically underestimated both intensity and frequency of records, with underestimation growing with the margin of record exceedance. Authors attribute this to neural networks' inability to extrapolate beyond training domain vs physics laws. · read in full (as of 2026-08-08)
- [confirmed] Global AIWP models including ECMWF AIFS systematically underestimate tropical cyclone intensity and show larger intensity errors across all lead times vs regional high-resolution references; a key reason is training on ERA5 reanalysis at ~31km which itself underestimates intensity and cannot resolve mesoscale eyewall structures critical for intensity development. · read in full (as of 2026-08-08)
- [confirmed] WeatherNext Cyclones' rapid-intensification skill remains limited: a standard detection metric balancing correct detections against false alarms and misses improved from below 0.3 to 0.5 under WN-C, meaning the model still gets rapid intensification wrong a meaningful fraction of time; larger 1,000-member ensembles improve tail sampling but do not eliminate misses and raise false-alarm/over-warning costs. · read in full (as of 2026-08-08)
- [confirmed] Operational consensus vs replacement debate: In 2024 NHC official human-in-loop forecasts outperformed all individual models at almost all lead times for track, and outperformed individual models for intensity except 4-5 day LGEM; NHC Director Brennan emphasizes "no guarantee that one model, because it did well last year or for this particular storm, is necessarily going to be best for the next storm" and that "a hurricane is not just a track or intensity forecast it requires experts to translate into impacts." WN-C is framed by NHC as "another tool in the toolbox," not a replacement, and is weighted inside consensus (75% latitude, 89% longitude, 42% intensity in two-way blend) rather than used alone. · read in full (as of 2026-08-08)
- [confirmed] WeatherNext Cyclones operates at 0.256 (~28x28 km) resolution, about 100x coarser than traditional high-resolution regional intensity models; WeatherNext 2-mini operates at 16 (~111x111 km) and still shows strong performance. · read in full (as of 2026-08-08)
- [confirmed] The coarse-resolution intensity skill (28km, 100x coarser than regional models) is an unexplained black-box result that shocked the meteorological community; DeepMind authors state they do not fully understand how the model extracts intensity signal from coarse inputs, and describe it as "a black box at the end of the day." · read in full (as of 2026-08-08)
- [confirmed] Operational consensus vs replacement debate: In 2024 NHC official human-in-loop forecasts outperformed all individual models at almost all lead times for track, and outperformed individual models for intensity except 4-5 day LGEM; NHC Director Brennan emphasizes "no guarantee that one model, because it did well last year or for this particular storm, is necessarily going to be best for the next storm" and that "a hurricane is not just a track or intensity forecast it requires experts to translate into impacts." WN-C is framed by NHC as "another tool in the toolbox," not a replacement, and is weighted inside consensus (75% latitude, 89% longitude, 42% intensity in two-way blend) rather than used alone. · read in full (as of 2026-08-08)
- [confirmed] Training uses Continuous Ranked Probability Score (CRPS) on per-location marginals with fair estimator, plus short autoregressive rollouts up to 8 steps; joint spatial/cross-variable structure emerges despite marginal-only supervision. · read in full (as of 2026-08-08)
- [confirmed] Including WeatherNext Cyclones predictions in a weighted-average consensus ensemble substantially improves consensus skill. · read in full (as of 2026-08-08)
- [confirmed] Melissa underwent rapid intensification (70 to 130 mph in 24h; 115 mph increase and 90 mb pressure drop in 72h), the fourth RI storm of 2025; NHC track forecast 4 days before landfall was within ~13 miles. · read in full (as of 2026-08-08)
- [confirmed] Tropical cyclones have caused >700,000 deaths and $1.4 trillion in economic losses globally over the past 50 years. · read in full (as of 2026-08-08)
- [likely] Every $1 invested in early warning systems can yield up to $10 in benefits, and 24 hours' warning can cut storm damage by 30% (WMO/CREWS estimates). · read in full (as of 2026-08-08)
- [likely] WeatherNext requires orders of magnitude less compute than physics-based models thesis claim of 1,000x less compute enabling national met services with limited compute to run forecasts. · read in full (as of 2026-08-08)
- [confirmed] The coarse-resolution intensity skill remains an open research question; authors state they do not fully understand how the model achieves state-of-the-art intensity at 28km and invite community investigation. · read in full (as of 2026-08-08)
- [confirmed] WeatherNext does not replace official warnings; all official alerts remain issued solely by national meteorological authorities, and the model is research code provided as-is without warranty. · read in full (as of 2026-08-08)
- [confirmed] Each FGN constituent model uses a GNN encoder/decoder mapping lat/lon grid to a 6-times-refined icosahedral mesh with graph-transformer processor, 180M parameters per seed, 768 latent dimension, 24 layers, 6-hour timestep (vs GenCast 57M total, 512 dim, 16 layers, 12-hour timestep). · read in full (as of 2026-08-08)
- [confirmed] WMO Atlas update: extreme weather, climate and water-related events caused 11,778 reported disasters between 1970 and 2021, with just over 2 million deaths and US$4.3 trillion in economic losses; storms are the leading cause of economic losses and the sole hazard with continually increasing attributed portion; over 90% of deaths in developing countries; USA alone US$1.7 trillion (39% of global losses). · read in full (as of 2026-08-08)
- [confirmed] NHC Verification Report confirms 2025 was the first year NHC incorporated AI-based models into real-time operations; GDMI was not consistently available early season, so homogeneous model comparisons exclude first two Atlantic storms and first five eastern North Pacific storms, limiting sample for GDMI vs official head-to-head. · read in full (as of 2026-08-08)
- [confirmed] Even the hybrid regional AI system HITS-LPIPS trained on 9km high-resolution typhoon reanalysis (HiRes) and constrained by AIFS still underestimates super-typhoon intensity by ~10-15 m/s and shows limited skill for intense typhoons >50 m/s, indicating training-data resolution and rarity of extremes remain binding constraints. · read in full (as of 2026-08-08)
- [confirmed] Operational intensity failure mode illustrated by Hurricane Milton (2024): Google AI track was miles-accurate but intensity forecast predicted <Category 2 vs actual Category 5 (most intense Gulf storm since Rita 2005), a rapid-intensification event where storm went from formation to Cat5 in ~2 days; nearly 80% of major hurricanes (Cat3-5) undergo rapid intensification, making intensity the hardest forecast dimension for both AI and physics models. · read in full (as of 2026-08-08)
- [confirmed] False-alarm economics are material: for tropical cyclone genesis, the European model had lowest false alarms but only ~20% correct detection, GFS 20-25% detection with more false alarms, Canadian highest detection but highest false alarms; when two+ models agree odds increase considerably. This tradeoff implies 1,000-member tail scenarios that do not materialize carry evacuation, port/airport closure, and staging costs if acted upon, a cost not quantified in the "extra day" framing. · read in full (as of 2026-08-08)
- [confirmed] AI models remain dependent on physics-based analysis for training and initialization and cannot currently replace classical NWP for high-risk applications; Zhang et al. authors and ECMWF commentary recommend parallel use and hybrid physics-informed approaches, with official warnings remaining solely with national meteorological authorities and liability not transferred to AI providers. · read in full (as of 2026-08-08)
Where we hit a limit / what to double-check
- We did not obtain the full text of Google DeepMinde WeatherNext 2 Uses Functional Generative Networks For 8x Faster Probabilistic Weather Forecasts (https://www.marktechpost.com/2025/11/17/google-deepminds-weathernext-2-uses-functional-generative-networks-for-8x-faster-probabilistic-weather-forecasts); claims resting on it are from its summary — you may be able to reach it directly.
