For people who already know model data — the what, the how, and the why
You know what a 500 hPa anomaly correlation is and what an EPS spread looks like. This page is about what we do with them: score every run the way the operational centers do, then explain why a model won or lost by naming the flow it was in, flag the busts before they verify, and turn all of that into a defensible market line. Every panel below is drawn from live data — the same JSON the trading pages read — so the explanations demonstrate on the real archive, not a screenshot.
| Fork — target today | FF_spread: normalized ensemble spread vs climatology (an uncertainty predictor). FF_bust = P(ACC<c|X) and FF_revision = P(Δforecast>d|X) are Planned — not yet the same quantity. |
| Headline skill | Common-climatology ACC — NH 20–80°N 500 hPa anomaly correlation by lead, area-weighted (cos-lat). Each run vs its own analysis; a common-truth board is Planned. |
| Anomaly climatology | ERA5 fixed climatology (Z500 1991–2020) for both forecast and analysis. Not ECMWF's lead/model-dependent reforecast climatology; a model-climate ACC is Planned. |
| Grid & leads | 0.25° global (721×1440), leads 0–240 h every 12 h; regime + clusters at the Day-5 (120 h) / Day-7 (168 h) valid. |
| Bust thresholds | major degradation ACC<0.75 (the Fork's bust marker); loss of useful synoptic skill ACC<0.60. Lead-normalized Z_ACC + relative/regional bust tiers are Planned. |
| Regions | NH 20–80°N; sectors N. Atlantic, Europe, N. Pacific, N. America, Arctic, Asia. Regime boxes drawn in §2. |
| Scenarios | ECMWF-EPS 50 members, k-means k=4 at Day 7, raw Z500 anomaly. cos-lat weighting / EOF compression / candidate-k / multimodel are Planned. Best-cluster gap is an oracle diagnostic (§5). |
| Shadow House | Shadow weights = global Brier-skill × regime/ENSO tilt, γ=3, k₀=8, dims {Atlantic, Pacific, ENSO}; trained on Skill(m|R̂issue) — the issue-time predicted regime. Soft regime probabilities + proper-score stacking are Planned. Moves no live price. |
| Validation | Live panels show current sample counts. Strict walk-forward + block-bootstrap CIs (by synoptic episode) are Planned — treat "whole-archive" numbers as indicative until then. |
| Provenance | Tamper-evident, content-addressed (sha256) — see §7. Hash-chained manifests + externally-anchored daily roots are Planned. |
| Version | source commit: — |
This box is the honest scope line: what is Operational today versus what the framing points toward. The full external methodology audit driving the roadmap is tracked in the repo.
common-climatology ACC — the operational-center headline, one fixed climatology
The headline is the number the operational centers live by: Northern-Hemisphere 500 hPa anomaly correlation by lead time — area-weighted over 20–80°N, each forecast scored against its own verifying analysis, anomalies taken versus an ERA5 1991–2020 climatology. No station games, no cherry-picked cities — the hemispheric ACC used to declare when a model's useful skill runs out, conventionally where ACC falls through 0.60.
Two honest caveats, so the label is exact. This is common-climatology ACC: a single fixed ERA5 climatology for both forecast and analysis — legible and consistent, but not identical to ECMWF's lead- and model-dependent reforecast climatology (a model-climate ACC is on the roadmap; the gap grows at long leads). And each run is scored against its own analysis here — right for reproducing center-style headline scores, but a common-truth board (every model regridded and verified against one designated analysis) is what should train the cross-model weights, and is Planned.
loading…
the missing why behind the scorecard
ACC tells you how well a run did, never why. So for each run we take the Day-5 verifying analysis, form its Z500 anomaly, and reduce it to a handful of interpretable large-scale indices measured over fixed boxes — the AO/polar-cap, the NAO dipole, the classic Euro-Atlantic blocking centres, and the Wallace–Gutzler PNA points. Those indices classify the flow into named regimes per sector: the NH vortex state, the Euro-Atlantic weather regime (NAO+ / Greenland Block / Atlantic Ridge / Scandi Block), and the Pacific/PNA pattern. Here is exactly where each index is taken:
The payoff is the rollup: bucket every archived run by its regime and report Day-5 ACC per model per regime. "Hard flow" stops being a hunch and becomes a measured, per-model number — a reproducible artifact that exists only over our issue-time archive.
loading…
celsius.earth, folded in
Our box indices are computed off the analysis, so they're honest and work for any date — but they can't see the slow background state. celsius.earth publishes standardized, ERA5-based indices back to 1940: daily for the Z500 family (PNA, EPO, WPO, EA, WP, blocking, the zonal index) and monthly for the SLP family (NAO, AO, and SOI — the ENSO signal). We fold both in. The daily PNA/EPO now drive the Pacific regime label directly (PNA ≷ ±0.5σ for the PNA phases, EPO < −0.75σ for the Alaskan/NE-Pacific blocking ridge); the monthly SOI sets the ENSO phase that conditions the shadow line (§6).
loading…
three distinct quantities, one honest about what's proven
"The Fork" is really three different forecast quantities, and it matters not to conflate them. Today we demonstrate the first — current uncertainty — and show that it associates with realized skill. A calibrated probability of a bust, and a probability the forecast itself revises, are separate models we are building; they are labeled honestly below.
What we can show now: the Fork reads the ensemble's disagreement (normalized spread vs climatology), fully known at issue time, and we test whether it tracks skill — bin runs by their issue-time spread, plot the realized Day-5 ACC in each bin. A clean monotone fall is evidence spread is a bust signal. It is association, not yet a calibrated probability — that is FFbust, above. The public home of the metric is its companion site, forecastfork.com.
loading…
clustering, and an honest upper bound
The ensemble mean is a smooth average of distinct futures. So per region we k-means the 50 EPS members into
a few scenarios at Day 7, score each cluster's mean against the analysis, and compare the best cluster
to the ensemble mean. When a minority cluster verifies better than the mean, the consensus underweighted
the scenario that actually happened — the underused_gap.
Read this honestly: that gap is an Oracle diagnostic. It picks the best
cluster after the outcome is known (G_oracle = max_k ACC(C_k,Y) − ACC(mean,Y)), so with
more clusters it tends to grow by chance. It measures recoverable skill — an upper bound — not skill you
could have banked at issue time. The actionable versions — scenario-mixture skill S(Σ π_k F_k, Y)
and selection skill S(F_k̂, Y) with π_k, k̂ fixed strictly at issue time from cluster
geometry, momentum, cross-model support and analogs — are Planned, and are
where the real edge lives.
regime- and ENSO-conditioned weights, shrunk toward what's proven
The live House posts a skill-weighted line: each model's weight is its Brier skill vs a coin flip over the settled ledger — one global number per model. But §2 shows skill isn't flat across the flow. So the shadow House classifies the regime the current run is predicting (from the multi-model consensus Day-5 forecast — regime-robust) and tilts each model's proven global weight by its measured edge in that regime and ENSO phase:
The shrink term is the whole discipline: a thin regime bucket (small n) pulls the tilt back toward 1, so the shadow barely leaves the proven global weights until the sample earns the lean. It is a shadow only — it moves no live price — precisely so its edge can be watched against outcomes before it's ever trusted with the book.
loading…
content-addressed provenance
A market needs a tamper-evident answer to "what was on the screen at issue time?" Every map, chart and JSON for a run is frozen to an object store under a content hash: the key is the sha256 of the bytes. Re-minting identical content is a no-op; any change lands as a new hash and a new manifest row, with a UTC mint time and its source provenance.
Being precise about what this does and doesn't prove: a content hash proves identical bytes share an
identity and that a changed byte yields a different hash. It does not, on its own, prove when
an object first existed, that no alternate manifest was made, or that the operator never rewrote history. Making
it genuinely hard to rewrite — hash-chained manifest rows (H_t = SHA256(H_{t-1} ‖ M_t)),
daily Merkle roots published to an independent location, and embedded code/calibration-version hashes —
is Planned. That is the difference between "tamper-evident" and "unforgeable."
loading…
See it on a real run: the Forecast page for the forward view,
the Verification board for the scorecards, and any run's full post-mortem at
/verify/<init>.