Method

How Weather Trader reads the models

For people who already know model data — the what, the how, and the why

You know what a 500 hPa anomaly correlation is and what an EPS spread looks like. This page is about what we do with them: score every run the way the operational centers do, then explain why a model won or lost by naming the flow it was in, flag the busts before they verify, and turn all of that into a defensible market line. Every panel below is drawn from live data — the same JSON the trading pages read — so the explanations demonstrate on the real archive, not a screenshot.

1 · The scorecardZ500 ACC, operationally 2 · Regime attributionthe "why" 3 · Teleconnectionscelsius.earth 4 · Forecast Forkbust risk, early 5 · Scenariosunderused skill 6 · Shadow Houseattribution → a line 7 · The recordprovenance
Component status: OperationalShadow ExperimentalOracle diagnostic Planned
Technical specificationmethod v0.3
Fork — target todayFF_spread: normalized ensemble spread vs climatology (an uncertainty predictor). FF_bust = P(ACC<c|X) and FF_revision = P(Δforecast>d|X) are Planned — not yet the same quantity.
Headline skillCommon-climatology ACC — NH 20–80°N 500 hPa anomaly correlation by lead, area-weighted (cos-lat). Each run vs its own analysis; a common-truth board is Planned.
Anomaly climatologyERA5 fixed climatology (Z500 1991–2020) for both forecast and analysis. Not ECMWF's lead/model-dependent reforecast climatology; a model-climate ACC is Planned.
Grid & leads0.25° global (721×1440), leads 0–240 h every 12 h; regime + clusters at the Day-5 (120 h) / Day-7 (168 h) valid.
Bust thresholdsmajor degradation ACC<0.75 (the Fork's bust marker); loss of useful synoptic skill ACC<0.60. Lead-normalized Z_ACC + relative/regional bust tiers are Planned.
RegionsNH 20–80°N; sectors N. Atlantic, Europe, N. Pacific, N. America, Arctic, Asia. Regime boxes drawn in §2.
ScenariosECMWF-EPS 50 members, k-means k=4 at Day 7, raw Z500 anomaly. cos-lat weighting / EOF compression / candidate-k / multimodel are Planned. Best-cluster gap is an oracle diagnostic (§5).
Shadow HouseShadow weights = global Brier-skill × regime/ENSO tilt, γ=3, k₀=8, dims {Atlantic, Pacific, ENSO}; trained on Skill(m|R̂issue) — the issue-time predicted regime. Soft regime probabilities + proper-score stacking are Planned. Moves no live price.
ValidationLive panels show current sample counts. Strict walk-forward + block-bootstrap CIs (by synoptic episode) are Planned — treat "whole-archive" numbers as indicative until then.
ProvenanceTamper-evident, content-addressed (sha256) — see §7. Hash-chained manifests + externally-anchored daily roots are Planned.
Versionsource commit: —

This box is the honest scope line: what is Operational today versus what the framing points toward. The full external methodology audit driving the roadmap is tracked in the repo.

1The scorecard — Z500 anomaly correlation Operational

common-climatology ACC — the operational-center headline, one fixed climatology

The headline is the number the operational centers live by: Northern-Hemisphere 500 hPa anomaly correlation by lead time — area-weighted over 20–80°N, each forecast scored against its own verifying analysis, anomalies taken versus an ERA5 1991–2020 climatology. No station games, no cherry-picked cities — the hemispheric ACC used to declare when a model's useful skill runs out, conventionally where ACC falls through 0.60.

Two honest caveats, so the label is exact. This is common-climatology ACC: a single fixed ERA5 climatology for both forecast and analysis — legible and consistent, but not identical to ECMWF's lead- and model-dependent reforecast climatology (a model-climate ACC is on the roadmap; the gap grows at long leads). And each run is scored against its own analysis here — right for reproducing center-style headline scores, but a common-truth board (every model regridded and verified against one designated analysis) is what should train the cross-model weights, and is Planned.

Demonstration — ACC by lead, latest verified run● live

loading…

How to read itCurves that hold high and to the right are the sharper runs. The gap between two models at a given lead is the skill difference for that run; where a curve crosses the dashed 0.60 line is that model's predictability horizon for the day. A dip shared by all models is not a model problem — it's a hard-to-forecast regime, which is exactly what §2 names.

2Regime attribution — turning "it busted" into "it busted in this flow"

the missing why behind the scorecard

ACC tells you how well a run did, never why. So for each run we take the Day-5 verifying analysis, form its Z500 anomaly, and reduce it to a handful of interpretable large-scale indices measured over fixed boxes — the AO/polar-cap, the NAO dipole, the classic Euro-Atlantic blocking centres, and the Wallace–Gutzler PNA points. Those indices classify the flow into named regimes per sector: the NH vortex state, the Euro-Atlantic weather regime (NAO+ / Greenland Block / Atlantic Ridge / Scandi Block), and the Pacific/PNA pattern. Here is exactly where each index is taken:

The diagnostic boxes (NH, 20–90°N)Z500 anomaly indices
Each labelled rectangle is a cos-lat-weighted area mean of the Z500 anomaly; the NAO is the south-minus-north difference, the PNA the signed 4-point Wallace–Gutzler sum (● +, ○ −). The Pacific label is then overridden by celsius.earth's validated daily PNA/EPO when available (§3).

The payoff is the rollup: bucket every archived run by its regime and report Day-5 ACC per model per regime. "Hard flow" stops being a hunch and becomes a measured, per-model number — a reproducible artifact that exists only over our issue-time archive.

Demonstration — Pacific regime × model, Day-5 ACC (whole archive)● live

loading…

Worked example — the edge, namedThe Pacific PNA− (zonal / east ridge) pattern is typically the hardest: the wavetrain is flat and models phase the downstream troughs differently. That is where a deterministic model with weaker Pacific data assimilation bleeds the most skill — and where the market should trust the ensemble mean less and the sharper model more.

3Teleconnections — the low-frequency drivers a single Z500 field can't see

celsius.earth, folded in

Our box indices are computed off the analysis, so they're honest and work for any date — but they can't see the slow background state. celsius.earth publishes standardized, ERA5-based indices back to 1940: daily for the Z500 family (PNA, EPO, WPO, EA, WP, blocking, the zonal index) and monthly for the SLP family (NAO, AO, and SOI — the ENSO signal). We fold both in. The daily PNA/EPO now drive the Pacific regime label directly (PNA ≷ ±0.5σ for the PNA phases, EPO < −0.75σ for the Alaskan/NE-Pacific blocking ridge); the monthly SOI sets the ENSO phase that conditions the shadow line (§6).

Demonstration — the background state right now● live

loading…

The daily σ values are the validated authored indices at the current run's Day-5 valid date; ENSO is the monthly SOI-derived phase. This is the low-frequency context a single hemispheric forecast map cannot carry.

4The Forecast Fork — bust risk before the bust

three distinct quantities, one honest about what's proven

"The Fork" is really three different forecast quantities, and it matters not to conflate them. Today we demonstrate the first — current uncertainty — and show that it associates with realized skill. A calibrated probability of a bust, and a probability the forecast itself revises, are separate models we are building; they are labeled honestly below.

FFspread Operational
Normalized current uncertainty — the ensemble's own disagreement per region × lead, known at issue time.
what the chart below shows
FFbust Planned
A calibrated P(ACC < c | X) from a fitted model — not a spread-vs-ACC bin plot. With reliability, Brier/BSS, ROC.
§ roadmap 2.1
FFrevision Planned
P(the forecast materially changes next cycles) — the actual "forecasting the forecast" (Revision Tensor).
§ roadmap 2.2

What we can show now: the Fork reads the ensemble's disagreement (normalized spread vs climatology), fully known at issue time, and we test whether it tracks skill — bin runs by their issue-time spread, plot the realized Day-5 ACC in each bin. A clean monotone fall is evidence spread is a bust signal. It is association, not yet a calibrated probability — that is FFbust, above. The public home of the metric is its companion site, forecastfork.com.

Demonstration — issue-time spread vs realized Day-5 ACC (archive)● live

loading…

How to read itDown-and-to-the-right is what you want to see: as issue-time divergence rises, realized ACC falls. The steeper and cleaner that fall, the more a trader can lean on spread alone to size a position — cutting exposure into a high-divergence window before any analysis exists to confirm the bust.

5Ensemble scenarios — recoverable scenario skill Oracle diagnostic

clustering, and an honest upper bound

The ensemble mean is a smooth average of distinct futures. So per region we k-means the 50 EPS members into a few scenarios at Day 7, score each cluster's mean against the analysis, and compare the best cluster to the ensemble mean. When a minority cluster verifies better than the mean, the consensus underweighted the scenario that actually happened — the underused_gap.

Read this honestly: that gap is an Oracle diagnostic. It picks the best cluster after the outcome is known (G_oracle = max_k ACC(C_k,Y) − ACC(mean,Y)), so with more clusters it tends to grow by chance. It measures recoverable skill — an upper bound — not skill you could have banked at issue time. The actionable versions — scenario-mixture skill S(Σ π_k F_k, Y) and selection skill S(F_k̂, Y) with π_k, k̂ fixed strictly at issue time from cluster geometry, momentum, cross-model support and analogs — are Planned, and are where the real edge lives.

step 1
50 members
the raw EPS at Day 7, one field each
step 2
k-means → 4 scenarios
group members by pattern, per region
step 3
score each vs analysis
which scenario verified?
step 4
best − mean = gap
a minority cluster beating the mean = underused skill
Worked example — oracle boundIf a 16%-of-members cluster over the N. Atlantic verifies at ACC 0.89 while the full-ensemble mean sits at 0.83, that +0.06 recoverable gap says the blocking scenario the mean smoothed away was present and could have been ridden — in hindsight. Whether it was identifiable beforehand is the open, Planned question (mixture/selection skill), not something we claim the market should already have priced.

6The shadow House — turning attribution into a market line Shadow

regime- and ENSO-conditioned weights, shrunk toward what's proven

The live House posts a skill-weighted line: each model's weight is its Brier skill vs a coin flip over the settled ledger — one global number per model. But §2 shows skill isn't flat across the flow. So the shadow House classifies the regime the current run is predicting (from the multi-model consensus Day-5 forecast — regime-robust) and tilts each model's proven global weight by its measured edge in that regime and ENSO phase:

wshadow(m) = wglobal(m) · tilt(regime, m) · tilt(enso, m)    where  tilt = 1 + shrink·γ·( accregime/accbaseline − 1 ),  shrink = n / (n + k₀)

The shrink term is the whole discipline: a thin regime bucket (small n) pulls the tilt back toward 1, so the shadow barely leaves the proven global weights until the sample earns the lean. It is a shadow only — it moves no live price — precisely so its edge can be watched against outcomes before it's ever trusted with the book.

Now conditioned correctlyThe tilt is trained on Skill(m | R̂issue) — skill bucketed by the regime each run predicted at issue time, the same variable used to apply it — so training and operation condition on the same thing, and the table automatically absorbs the chance the regime classification itself was wrong. (The retrospective verifying-regime table is kept separately for the /verify science display.) Two refinements remain on the roadmap: soft regime probabilities — a weighted mix over regimes instead of one hard label, to kill the 0.5σ boundary jumps — and replacing the heuristic tilt with proper-score stacking. It is still a directional shadow, not yet a fitted weight.
Demonstration — the shadow tilt for the current run● live

loading…

How to read itGreen is the regime-tilted shadow weight; the navy edge is the proven global weight. The gap between them is the lean this flow earns — and if the buckets are still thin, that gap is deliberately small. Δ is the normalized weight shift per model.

7The minted record — a tamper-evident account of what was available, when Operational

content-addressed provenance

A market needs a tamper-evident answer to "what was on the screen at issue time?" Every map, chart and JSON for a run is frozen to an object store under a content hash: the key is the sha256 of the bytes. Re-minting identical content is a no-op; any change lands as a new hash and a new manifest row, with a UTC mint time and its source provenance.

Being precise about what this does and doesn't prove: a content hash proves identical bytes share an identity and that a changed byte yields a different hash. It does not, on its own, prove when an object first existed, that no alternate manifest was made, or that the operator never rewrote history. Making it genuinely hard to rewrite — hash-chained manifest rows (H_t = SHA256(H_{t-1} ‖ M_t)), daily Merkle roots published to an independent location, and embedded code/calibration-version hashes — is Planned. That is the difference between "tamper-evident" and "unforgeable."

Demonstration — a run's minted manifest● live

loading…

Served straight from the object store. Each row is immutable; the sha is the identity.

Putting it together — one run, end to end

score
ACC by lead
how well, operationally (§1)
explain
regime
what flow, who owns it (§2–3)
warn
the Fork
bust risk, before the outcome (§4)
mine
scenarios
skill the mean discarded (§5)
price
shadow line
tilt toward who wins this flow (§6)
freeze
mint
the permanent record (§7)

See it on a real run: the Forecast page for the forward view, the Verification board for the scorecards, and any run's full post-mortem at /verify/<init>.