Skill
Verification
Per-Run Model Skill · The NOAA Headline, Scored
How each model run actually performed. The headline is the one NOAA/EMC live by:
Northern-Hemisphere 500 hPa anomaly correlation by lead time — area-weighted over 20–80°N,
each forecast scored against its own verifying analysis, anomalies taken vs an ERA5 climatology.
Higher is better; skill is generally spent by the time ACC falls through 0.6.
Recent skill trend — Day-5 500 hPa ACC, run over run
How each model has been performing across inits. A dip shared by all models marks a
hard-to-forecast regime — a bust window. Click a point to open that run.
Anomaly correlation by lead
NH 20–80°N · verified vs own analysis · the 0.60 line marks where useful skill runs out
Skill matrix
ACC per model × lead — darker = sharper. The at-a-glance scorecard.
Regional skill — where each model wins
ACC by region at a chosen lead — the spatial structure of model superiority, the first cut at why. Row winner ringed.
Regime attribution — what kind of flow, and who handles it
The Day-5 verifying flow, classified into named synoptic regimes (NH vortex state · the Euro-Atlantic
weather regime · the Pacific/PNA pattern). Across the archive we then condition Day-5 skill on the regime —
so "hard flow" stops being a hunch and becomes a per-model, per-regime number. This is the umbrella
why behind the regional table above.
Spatial error — where the height field busted
Z500 forecast minus verifying analysis (m), Northern Hemisphere. Red = forecast
height too high (ridge too strong / trough too weak); blue = too low. The red/blue
dipoles are misplaced troughs & ridges — the actual busts, and where they sit is the why.
Ensemble scenarios — skill the mean threw away
The 50 ECMWF-EPS members k-means-clustered into scenarios per region (Day 7). When a
minority cluster verifies better than the ensemble mean, the consensus underweighted the scenario
that actually happened — underused skill, and the raw signal for forecasting the forecast.
Forecast Fork — bust risk, before the bust
The ensemble's own disagreement, per region × lead — known at forecast time, with no
outcome needed. Redder = the 50 members diverge more (higher spread vs
climatology) = the mean is more likely to bust. Across the archive this predicts skill: —.
Divergence — this run (◼ = it did bust, ACC<0.75)
Calibration — spread vs realized ACC (whole archive)
RMSE by lead
Root-mean-square height error (m) — the raw magnitude of the miss