How each model run actually performed. The headline is the one NOAA/EMC live by:
Northern-Hemisphere 500 hPa anomaly correlation by lead time — area-weighted over 20–80°N,
each forecast scored against its own verifying analysis, anomalies taken vs the ERA5 1991–2020
climatology, in the uncentered NCEP/EMC form so the numbers sit on the same scale as the standard charts.
Higher is better; skill is generally spent by the time ACC falls through 0.6.
How this works ▸
Recent skill trend — Day-5 500 hPa ACC, by verification date
How each model has been performing, plotted by the date each forecast was verifying for
(a Day-5/7 forecast lands 5/7 days after its run). A dip shared by all models marks a
hard-to-forecast day — a bust window. Showing WeatherNext · ECMWF HRES · AIFS by default;
the legend toggles any model on. Click a point to open that run.
Leaderboard — mean NH Day-5 ACC by window
The standing record, not one run's story: each cell is the mean over every forecast that
verified inside the window, against the one common truth. Ranked by the 30-day column; best per
column ringed; the small number is how many runs the mean stands on.
Skill matrix — last 30 days
Mean ACC per model × lead over the last 30 days of verifications — darker = sharper. The scorecard as a record, not a snapshot.
Regional skill — where each model wins, last 30 days
Mean ACC by region at Day 5 / 7 / 10 over the last 30 days — the spatial structure of model superiority. Row winner ringed.
Day
Autopsy
Run post-mortem
One run, end to end — its maps, its fork, its regime, the decisions that would have
perfected it. Defaults to the newest run verified through Day 5; pick any recent init.
InitField
Anomaly correlation by lead
NH 20–80°N, cos-lat weighted · each model against its own analysis at the valid time (the operational-center convention; the common-truth board above instead grades everything against the one ECMWF analysis) · anomalies vs ERA5 1991–2020, uncentered ACC (NCEP/EMC form — the domain-mean anomaly is kept, as on the standard 5-day NH charts) · the 0.60 line marks where useful skill runs out
Regime attribution — what kind of flow, and who handles it
The Day-5 verifying flow, classified into named synoptic regimes (NH vortex state · the Euro-Atlantic
weather regime · the Pacific/PNA pattern). Across the archive we then condition Day-5 skill on the regime —
so "hard flow" stops being a hunch and becomes a per-model, per-regime number. This is the umbrella
why behind the regional table above.
Spatial error — where the height field busted
Z500 forecast minus verifying analysis (m), Northern Hemisphere. Red = forecast
height too high (ridge too strong / trough too weak); blue = too low. The red/blue
dipoles are misplaced troughs & ridges — the actual busts, and where they sit is the why.
Lead
Ensemble scenarios — skill the mean threw away
The 50 ECMWF-EPS members k-means-clustered into scenarios per region (Day 7). When a
minority cluster verifies better than the ensemble mean, the consensus underweighted the scenario
that actually happened — underused skill, and the raw signal for forecasting the forecast.
Scenario lineage — the branching outlook, tracked through time
The Fork's keystone. Instead of re-clustering each run and forgetting, we track coherent
scenario branches for a fixed day across every issuing run — matching this run's clusters to the
last run's branches, carrying their identity forward. Each line is one persistent scenario; its height is
its share (π) of the 50 members, read right (long lead) to left (the day). A branch that climbs is
the outlook converging on it; one that fades is a scenario the models abandoned.
Forecast Fork — bust risk, before the bust
The ensemble's own disagreement, per region × lead — known at forecast time, with no
outcome needed. Redder = the 50 members diverge more (higher spread vs
climatology) = the mean is more likely to bust. Across the archive this predicts skill: —.
Companion site: forecastfork.com.
Divergence — this run (◼ = it did bust, ACC<0.75)
Calibration — spread vs realized ACC (whole archive)
Forecast revision — how settled the outlook is
The Fork's third axis. Spread is how much the ensemble disagrees; bust risk is how
likely the mean is wrong; revision is how much the deterministic outlook keeps changing
run to run — a distinct instability the market can read. Bars show this run's day-by-day z500 revision
vs the seasonal norm: below the line = more settled than usual,
above = still moving.
Fleet independence — how many models do we really have?
Errors vs the common analysis correlate across models — the AI models especially herd
toward ECMWF, whose fields they were trained on. Neff = k²/Σρ² is the effective number of
independent opinions in the fleet: —. A consensus of redundant models is more
confident than it has any right to be. Full instrument: the Herding Index →
Neff by region × lead (of — models)
Most redundant pairs — NH day 5 error correlation
RMSE by lead
Root-mean-square height error (m) — the raw magnitude of the miss