Casebook

The Casebook

One verified run a week, taken apart in public

Every week the most instructive fully-verified model run becomes a written case study — what the models said at issue, where the ensemble forked, what the shadow House did, what actually verified, who won, and what a reader should take away. Every sentence is composed from numbers in the frozen verification record; nothing is written after the fact that couldn't have been read at the time.

How a study is chosen. Each Monday the pipeline scores every fully-verified init of the trailing week for instructiveness: how far its Day-5 skill sat from the season mean (busts and nails both teach), how widely the models disagreed (a winner and a loser to name), and whether the issue-time Fork call — the calibrated bust probability — turned out right. The standout becomes that week's study. Everything is traceable: each figure comes from the same frozen artifacts served on the Verification board and each study links to the full per-run post-mortem, so any claim can be checked against the record it cites.