Method

How Weather Trader reads the models

For people who already know model data — the what, the how, and the why

You know what a 500 hPa anomaly correlation is and what an EPS spread looks like. This page is about what we do with them: score every run the way the operational centers do, then explain why a model won or lost by naming the flow it was in, flag the busts before they verify, and turn all of that into a defensible market line. Every panel below is drawn from live data — the same JSON the trading pages read — so the explanations demonstrate on the real archive, not a screenshot.

1 · The scorecardZ500 ACC, operationally 2 · Regime attributionthe "why" 3 · Teleconnectionscelsius.earth 4 · Forecast Forkbust risk, early 5 · Scenariosunderused skill 5b · Fork familylineage + Reverse Fork 6 · Shadow Houseattribution → a line 7 · The recordprovenance
Component status: OperationalShadow ExperimentalOracle diagnostic Planned
Technical specificationmethod v0.3
Fork — target todayFF_spread: normalized ensemble spread vs climatology (an uncertainty predictor). FF_bust = P(ACC<c|X) is Live (walk-forward-validated), and FF_revision — the run-over-run z500 revision tensor + settled-vs-moving signal — is Live.
Headline skillCommon-climatology ACC — NH 20–80°N 500 hPa anomaly correlation by lead, area-weighted (cos-lat). Display board = each run vs its own analysis; a common-truth board (all models vs the ECMWF analysis) Operational trains the shadow House.
Anomaly climatologyERA5 fixed climatology (Z500 1991–2020) for both forecast and analysis. Not ECMWF's lead/model-dependent reforecast climatology; a model-climate ACC is Planned.
Grid & leads0.25° global (721×1440), leads 0–240 h every 12 h; regime + clusters at the Day-5 (120 h) / Day-7 (168 h) valid.
Bust thresholdsmajor degradation ACC<0.75 (the Fork's bust marker); loss of useful synoptic skill ACC<0.60. Lead-normalized Z_ACC + relative/regional bust tiers are Planned.
RegionsNH 20–80°N; sectors N. Atlantic, Europe, N. Pacific, N. America, Arctic, Asia. Regime boxes drawn in §2.
ScenariosECMWF-EPS 50 members, k-means k=4 at Day 7, raw Z500 anomaly. cos-lat weighting / EOF compression / candidate-k / multimodel are Planned. Best-cluster gap is an oracle diagnostic (§5).
Shadow HouseShadow weights = a proper-score stacking prior × regime/ENSO tilt. The prior is the minimum-variance blend w* = Σ⁻¹1 on the model error covariance — it rewards skill and penalises redundancy (Neff ≈ 2.3 of 4, so the ECMWF~AIFS pair isn't double-counted). On top, a shrunk regime/ENSO tilt (γ=3, k₀=8, soft-blended over the predicted-regime distribution), trained on Skill(m|R̂issue). Leave-init-out backtest: the prior beats equal weight by +5.5% blend RMSE; the regime tilt adds a further +0.2% — it helps and doesn't fight the prior. Moves no live price.
ValidationLive panels show current sample counts. Strict walk-forward (rolling-origin, past-only) and block-bootstrap CIs (by synoptic episode) are now Live — the P(bust) model reports both leave-init-out and walk-forward scores; regime attribution carries block-bootstrap CIs.
ProvenanceTamper-evident, content-addressed (sha256) — see §7. Manifests are now hash-chained (Ht=SHA256(Ht−1‖Mt)) with a cross-init anchor root; publishing that root off-host is the remaining step.
Versionsource commit: —

This box is the honest scope line: what is Operational today versus what the framing points toward. The full external methodology audit driving the roadmap is tracked in the repo.

1The scorecard — Z500 anomaly correlation Operational

common-climatology ACC — the operational-center headline, one fixed climatology

The headline is the number the operational centers live by: Northern-Hemisphere 500 hPa anomaly correlation by lead time — area-weighted over 20–80°N, each forecast scored against its own verifying analysis, anomalies taken versus an ERA5 1991–2020 climatology. No station games, no cherry-picked cities — the hemispheric ACC used to declare when a model's useful skill runs out, conventionally where ACC falls through 0.60.

The correlation is taken in the uncentered (NCEP/EMC) form: the domain-mean anomaly is kept rather than subtracted, which is what puts these numbers on the same scale as the published 5-day NH charts. The choice is not cosmetic. Scoring 2026 against a 1991–2020 baseline leaves a domain-mean NH height anomaly of about +29 m — the warming signal — and that offset is, in effect, perfectly forecast. Measured over 30 verification dates at day 5, the ECMWF/WMO centered form (which removes it) runs 0.017 lower for ECMWF and 0.015 lower for GFS. Same forecasts, same analyses, same climatology; only the normalization differs.

Two honest caveats, so the label is exact. This is common-climatology ACC: a single fixed ERA5 climatology for both forecast and analysis — legible and consistent, but not identical to ECMWF's lead- and model-dependent reforecast climatology (a model-climate ACC is on the roadmap; the gap grows at long leads). And each run is scored against its own analysis on this display board — right for reproducing center-style headline scores. Cross-model weights must not train on that (a forecast sits slightly closer to its own analysis than to an independent one), so a second common-truth board Operational verifies every model against one target — the ECMWF operational analysis, on the identical 0.25° grid — and that trains the shadow House. As expected it leaves ECMWF/AIFS (which share the IFS analysis) unchanged and removes GFS's small self-verification edge.

Demonstration — ACC by lead, latest verified run● live

loading…

How to read itCurves that hold high and to the right are the sharper runs. The gap between two models at a given lead is the skill difference for that run; where a curve crosses the dashed 0.60 line is that model's predictability horizon for the day. A dip shared by all models is not a model problem — it's a hard-to-forecast regime, which is exactly what §2 names.

2Regime attribution — turning "it busted" into "it busted in this flow"

the missing why behind the scorecard

ACC tells you how well a run did, never why. So for each run we take the Day-5 verifying analysis, form its Z500 anomaly, and reduce it to a handful of interpretable large-scale indices measured over fixed boxes — the AO/polar-cap, the NAO dipole, the classic Euro-Atlantic blocking centres, and the Wallace–Gutzler PNA points. Those indices classify the flow into named regimes per sector: the NH vortex state, the Euro-Atlantic weather regime (NAO+ / Greenland Block / Atlantic Ridge / Scandi Block), and the Pacific/PNA pattern. Here is exactly where each index is taken:

The diagnostic boxes (NH, 20–90°N)Z500 anomaly indices
Each labelled rectangle is a cos-lat-weighted area mean of the Z500 anomaly; the NAO is the south-minus-north difference, the PNA the signed 4-point Wallace–Gutzler sum (● +, ○ −). The Pacific label is then overridden by celsius.earth's validated daily PNA/EPO when available (§3).

The payoff is the rollup: bucket every archived run by its regime and report Day-5 ACC per model per regime. "Hard flow" stops being a hunch and becomes a measured, per-model number — a reproducible artifact that exists only over our issue-time archive.

Demonstration — Pacific regime × model, Day-5 ACC (whole archive)● live

loading…

Worked example — the edge, namedThe Pacific PNA− (zonal / east ridge) pattern is typically the hardest: the wavetrain is flat and models phase the downstream troughs differently. That is where a deterministic model with weaker Pacific data assimilation bleeds the most skill — and where the market should trust the ensemble mean less and the sharper model more.

3Teleconnections — the low-frequency drivers a single Z500 field can't see

celsius.earth, folded in

Our box indices are computed off the analysis, so they're honest and work for any date — but they can't see the slow background state. celsius.earth publishes standardized, ERA5-based indices back to 1940: daily for the Z500 family (PNA, EPO, WPO, EA, WP, blocking, the zonal index) and monthly for the SLP family (NAO, AO, and SOI — the ENSO signal). We fold both in. The daily PNA/EPO now drive the Pacific regime label directly (PNA ≷ ±0.5σ for the PNA phases, EPO < −0.75σ for the Alaskan/NE-Pacific blocking ridge); the monthly SOI sets the ENSO phase that conditions the shadow line (§6).

Demonstration — the background state right now● live

loading…

The daily σ values are the validated authored indices at the current run's Day-5 valid date; ENSO is the monthly SOI-derived phase. This is the low-frequency context a single hemispheric forecast map cannot carry.

4The Forecast Fork — bust risk before the bust

three distinct quantities, one honest about what's proven

"The Fork" is really three different forecast quantities, and it matters not to conflate them. Today we demonstrate the first — current uncertainty — and show that it associates with realized skill. A calibrated probability of a bust, and a probability the forecast itself revises, are separate models we are building; they are labeled honestly below.

FFspread Operational
Normalized current uncertainty — the ensemble's own disagreement per region × lead, known at issue time.
what the chart below shows
FFbust Experimental
A fitted, calibrated P(ACC < c | X) — logistic on spread + lead, scored out-of-sample by leave-init-out (block by run).
fitting…
FFrevision Live
How much the outlook keeps changing run over run — the Revision Tensor (z500 revision by lead) + a settled-vs-moving read, on /verify.
§ roadmap 2.2

What we can show now: the Fork reads the ensemble's disagreement (normalized spread vs climatology), fully known at issue time, and we test whether it tracks skill — bin runs by their issue-time spread, plot the realized Day-5 ACC in each bin. A clean monotone fall is evidence spread is a bust signal. It is association, not yet a calibrated probability — that is FFbust, above. The public home of the metric is its companion site, forecastfork.com.

Demonstration — issue-time spread vs realized Day-5 ACC (archive)● live

loading…

How to read itDown-and-to-the-right is what you want to see: as issue-time divergence rises, realized ACC falls. The steeper and cleaner that fall, the more a trader can lean on spread alone to size a position — cutting exposure into a high-divergence window before any analysis exists to confirm the bust.

5Ensemble scenarios — recoverable scenario skill Oracle diagnostic

clustering, and an honest upper bound

The ensemble mean is a smooth average of distinct futures. So per region we k-means the 50 EPS members into a few scenarios at Day 7, score each cluster's mean against the analysis, and compare the best cluster to the ensemble mean. When a minority cluster verifies better than the mean, the consensus underweighted the scenario that actually happened — the underused_gap.

Read this honestly: that gap is an Oracle diagnostic. It picks the best cluster after the outcome is known (G_oracle = max_k ACC(C_k,Y) − ACC(mean,Y)), so with more clusters it tends to grow by chance. It measures recoverable skill — an upper bound — not skill you could have banked at issue time. The actionable versions — scenario-mixture skill S(Σ π_k F_k, Y) and selection skill S(F_k̂, Y) with π_k, k̂ fixed strictly at issue time from cluster geometry, momentum, cross-model support and analogs — are Planned, and are where the real edge lives.

step 1
50 members
the raw EPS at Day 7, one field each
step 2
k-means → 4 scenarios
group members by pattern, per region
step 3
score each vs analysis
which scenario verified?
step 4
best − mean = gap
a minority cluster beating the mean = underused skill
Worked example — oracle boundIf a 16%-of-members cluster over the N. Atlantic verifies at ACC 0.89 while the full-ensemble mean sits at 0.83, that +0.06 recoverable gap says the blocking scenario the mean smoothed away was present and could have been ridden — in hindsight. Whether it was identifiable beforehand is the open, Planned question (mixture/selection skill), not something we claim the market should already have priced.

5bPersistent scenario lineage & Reverse Fork Experimental

the branching probability graph — the distinctive core

The scenarios above are re-clustered fresh each run. The distinctive step is to stop doing that: for a fixed valid time, track the same coherent branches across successive runs, so the ensemble becomes a mixture F_t(x|v) = Σ_k π_{k,t} G_{k,t}(x) whose branches persist and whose probability mass flows as the forecast matures. Matching each run's clusters to the previous run's branches (optimal assignment on pattern correlation) yields a branching graph — births, deaths, splits, merges — and per-branch Δπ (Fork Momentum). Verified on a real case: over 14 runs toward one valid time, one branch collapsed 46%→8% while another grew 28%→48% — a scenario resolving, invisible to the mean or spread alone.

Forward Fork Experimental
From the current state → which outcomes are possible and how their probabilities are evolving.
Reverse Fork Experimental
Given an outcome E → P(E)=Σ q_k(E)π_k, which branches carry it (Bayes P(B_k|E)), and the trigger / kill conditions.
Cascade / Revision / Autopsy Planned
One branch → all downstream targets (HDD/CDD, load, wind, contracts); P(forecast revises); post-verification lineage replay.
Reverse Fork, validatedFor "KPHX daily Tmax ≥ 115 °F" at Day 3 the branch-weighted Σ q_k π_k equals the raw ensemble frequency exactly (0.62) — the partition check — and decomposes it into which atmospheric branches carry the heat, plus the trigger (the branch that must survive) and the kill (what makes it collapse). Two guards held throughout: no hindsight leakage (prospective vs retrospective kept separate) and association ≠ causation (we say pathway/antecedent, not cause). Branch labels are coarse today (proper regime labels + EMOS-calibrated q_k are the refinements).

This whole family runs near-data on our own multi-terabyte ensemble archive — we own the issue-time ensemble rather than re-fetching it. Fork Bust (§4) and Fork Independence (§6) are the already-built confidence members of the same family.

6The shadow House — turning attribution into a market line Shadow

regime- and ENSO-conditioned weights, shrunk toward what's proven

The live House posts a skill-weighted line: each model's weight is its Brier skill vs a coin flip over the settled ledger — one global number per model. But §2 shows skill isn't flat across the flow. So the shadow House classifies the regime the current run is predicting (from the multi-model consensus Day-5 forecast — regime-robust) and tilts each model's proven global weight by its measured edge in that regime and ENSO phase:

wshadow(m) = wglobal(m) · tilt(regime, m) · tilt(enso, m)    where  tilt = 1 + shrink·γ·( accregime/accbaseline − 1 ),  shrink = n / (n + k₀)

The shrink term is the whole discipline: a thin regime bucket (small n) pulls the tilt back toward 1, so the shadow barely leaves the proven global weights until the sample earns the lean. It is a shadow only — it moves no live price — precisely so its edge can be watched against outcomes before it's ever trusted with the book.

Now conditioned correctly, and smoothedThe tilt is trained on Skill(m | R̂issue) — skill bucketed by the regime each run predicted at issue time, the same variable used to apply it (so training and operation condition on the same thing, and the table absorbs the chance the regime classification was wrong) — and it is applied by blending over a soft regime distribution P(R=r | Ft), not one hard label, so the weight no longer jumps when an index crosses a 0.5σ threshold. The retrospective verifying-regime table is kept separately for the /verify science display. One refinement remains: replacing the heuristic tilt with proper-score stacking (roadmap 2.7). Still a directional shadow, not yet a fitted weight.
Demonstration — the shadow tilt for the current run● live

loading…

How to read itGreen is the regime-tilted shadow weight; the navy edge is the proven global weight. The gap between them is the lean this flow earns — and if the buckets are still thin, that gap is deliberately small. Δ is the normalized weight shift per model.

One more thing a skill-weighted blend must not ignore: the models are not independent sources. AI models (AIFS, GraphCast-family) train toward ERA5/IFS, so their errors correlate with ECMWF's — and as agencies nudge full-physics systems with AI, the "independent" models converge onto one shared error structure. A blend weighted by skill or count then overweights what is really a single vote. We measure it directly: the correlation of each pair's error fields (vs the common truth), reduced to an effective number of independent models Experimental.

Demonstration — model error correlation & effective independence (Day-5, NH)● live

loading…

Why it mattersNeff below the model count is redundancy you're double-counting. The fix is a diversity penalty in the blend — Ω_dep = Σ wₘwₙ·ρ(eₘ,eₙ) — so genuinely independent information is weighted up and near-duplicates down. That penalty is roadmap 2.7 (proper-score stacking); this diagnostic is what parameterizes it.

7The minted record — a tamper-evident account of what was available, when Operational

content-addressed provenance

A market needs a tamper-evident answer to "what was on the screen at issue time?" Every map, chart and JSON for a run is frozen to an object store under a content hash: the key is the sha256 of the bytes. Re-minting identical content is a no-op; any change lands as a new hash and a new manifest row, with a UTC mint time and its source provenance.

Being precise about what this does and doesn't prove: a content hash proves identical bytes share an identity and that a changed byte yields a different hash. It does not, on its own, prove when an object first existed, that no alternate manifest was made, or that the operator never rewrote history. Making it genuinely hard to rewrite — hash-chained manifest rows (H_t = SHA256(H_{t-1} ‖ M_t)), daily Merkle roots published to an independent location, and embedded code/calibration-version hashes — is Planned. That is the difference between "tamper-evident" and "unforgeable."

Demonstration — a run's minted manifest● live

loading…

Served straight from the object store. Each row is immutable; the sha is the identity.

Putting it together — one run, end to end

score
ACC by lead
how well, operationally (§1)
explain
regime
what flow, who owns it (§2–3)
warn
the Fork
bust risk, before the outcome (§4)
mine
scenarios
skill the mean discarded (§5)
price
shadow line
tilt toward who wins this flow (§6)
freeze
mint
the permanent record (§7)

See it on a real run: the Forecast page for the forward view, the Verification board for the scorecards, and any run's full post-mortem at /verify/<init>.