Atlas Ahead

Stormfront

A second family of models, named for cyclone structure rather than cloud formations, and built to answer a different question: not how wrong were the polls, but which races were about to be called for the wrong side. Every one is scored across 1,401 elections — President, Senate, Governor and House — and every race it got right and wrong is on this page.

125
Upsets caught
38
Upsets missed
198
False alarms
77%
Recall
39%
Precision

What these numbers mean before you read anything else

Hypercane flags 323 races out of 1,401 — the top 22.5% of each cycle by risk, minus any race where one candidate still leads by 6 points or more after the pull toward the state's own voting history. Of the 163 races where the polls named the losing candidate, it flags 125 and misses 38. It also flags 198 races that the polls got right after all.

Three flags in five are false alarms. That is what 39% precision means, and it is the honest cost of catching 77% of the upsets. A model page showing only the hits would be advertising; the false alarms are drawn on the wall below in the same size as everything else.

The panel holds 1,869 races. 1,649 have a state voting history, which every model here needs, and 24 of those have a polling average of exactly zero and are left out, for the reason recorded further down. 1,401 are scored, from 2005 to 2024, because the first three cycles can only ever be training data for a model that is never allowed to see its own future. Every percentage on this page is over those 1,401 unless it says otherwise.

Does it actually beat what we had?

Wall Cloud is the best model in The Cloud Room, and three earlier attempts to beat it failed. Wall Cloud's scores stop at 2022, so this comparison covers the 1,300 races from 2005 to 2022, 153 of them upsets. Wall Cloud is given exactly as many flags as Hypercane in every cycle, so the two rows are directly comparable.

modelflaggedcaughtmissed recallprecisionAUC
Wall Cloud300104 4968%35%0.853
Hypercane300117 3676%39%0.876

Bootstrap, 4,000 resamples, Hypercane minus Wall Cloud, on exactly the flags drawn above:
Δrecall    +8.5 pts, 95% CI [+1.2, +16.0]   P(better) 0.983
Δprecision +4.3 pts, 95% CI [+1.0, +7.8]   P(better) 0.995
ΔAUC      +0.022, 95% CI [+0.006, +0.039]   P(better) 0.995

On recall, precision and discrimination (AUC) the 95% interval excludes zero, so Hypercane is demonstrably better there. It has the higher AUC in 6 of 9 cycles.

An earlier version of this page claimed intervals that excluded zero, and at the time that was wrong: 22 races with a polling average of exactly zero were being scored as automatic upsets, which inflated both models and the headline signal. Once they were removed, the six-input model's lead no longer cleared zero. The intervals above are for the two-input model introduced in September 2026, on the corrected panel.

Flags are assigned within each cycle — the top 22.5% of that year's races — rather than by one threshold across twenty years. A single global cut lets one volatile cycle consume the whole flag budget and leave later cycles unflagged, which would make the per-cycle comparison meaningless. The confidence intervals above are computed on this same within-cycle rule, not a different one.

Tightened in September 2026: fewer false alarms, the same catches

Until 15 September 2026 this page flagged the top quarter of each cycle from a six-input model. Two changes cut the false alarms, chosen with the weight on the elections since 2016.

A 6-point guard. A race is not flagged if one candidate still leads by 6 points or more after the pull toward the state's own voting history, and that pull does not change who leads. Across the scored panel, 893 races looked like that, and the polls named the loser in 1.2% of them.

Two inputs instead of six, and a smaller quota. Poll count, the gap to the state's lean and the undecided share never earned their place, and once the adjusted margin is in the model, whether the pull flips the leader adds nothing either. The two-input model ranks races better since 2016 (AUC 0.858 against 0.851), which lets the quota drop from 25% to 22.5% of each cycle.

rule2016–2024 caughtfalse alarms 2005–2024 caughtfalse alarms2024 caughtfalse alarms
Before: six inputs, top 25%57 of 77103 128 of 1632308 of 1018
Now: two inputs, top 22.5%, guard57 of 7788 125 of 1631988 of 1015

2016–2024, bootstrap, 4,000 resamples, new rule minus old:
Δrecall       -0.0 pts, 95% CI [-3.8, +3.8]
Δfalse alarms -15, 95% CI [-23, -8]
Δprecision   +3.7 pts, 95% CI [+1.6, +6.1]

A caveat on the evidence. The models are always fitted out of cycle, but the new rule was chosen by looking at how each option did from 2016 to 2024, so those years are not a clean test of the choice itself. The check that is clean is walk-forward selection, which only ever looks at earlier cycles: it picked this same rule for 2022 and for 2024.

Every race, and what happened to it

One square per election. Click any square for the race. Filter by cycle or by office — House races are 643 of the 1,401 and 111 of the 163 upsets, so they dominate the wall by weight.

Caught — flagged, polls named the loser Missed — an upset it did not flag False alarm — flagged, polls were right Quiet — not flagged, polls right

The four models

What drives it — one number, and one guard

Hypercane now reads two things: how big the polling lead is once Coriolis has pulled it toward the state's own last presidential margin (by a fitted weight, currently 0.112), and whether the race is for the House. The smaller that adjusted lead, the likelier the polls have named the wrong winner.

43.4%
upset rate when the pull changes who leads  (n=53)
10.4%
upset rate when it does not  (n=1,348)

A 4.2× difference. That sign flip was once the model's headline term. With the adjusted margin in the model it adds nothing more, because a lead small enough to flip is already a small lead. So it now works the other way round, as the guard: a race is cleared only when the adjusted lead is large and the pull does not flip it.

The idea did not come from the polls. It came from watching a prediction market: across the 2026 races where both exist, corr(state lean − poll, market probability − ours) = +0.602. The market shrinks polls toward fundamentals, so we tested whether doing the same helps. On 1,401 historical races it does.

The bug that inflated an earlier version of this page

22 of the races in this panel had a final polling average of exactly zero. A race is scored as an upset when the sign of the poll disagrees with the sign of the result — and the sign of zero is zero, which never equals plus or minus one. So every dead-level race was silently counted as an upset, giving that group an upset rate of 1.000.

A poll that is dead level cannot name the wrong side, because it names no side. Those races are now excluded. The effect was not cosmetic: the flip signal fell from a published 5.8× to a true 4.8×, the base upset rate fell from 12.8% to 11.8%, and the confidence intervals on the six-input Hypercane's advantage — which had excluded zero — no longer did. Every number on this page is computed on the corrected panel.

Removing the contamination widened Hypercane's lead over Wall Cloud on AUC, because the degenerate races were automatic upsets that both models flagged. It narrowed the confidence in that lead, because there are fewer real upsets to be confident about. Both moves are recorded rather than the flattering one alone.

Applied to 2026

Coriolis on the ten 2026 Senate races with polling, alongside what the market is charging for the same outcome. This is not the pre-registered call, which freezes 2026-11-02.

One race fires the signal, and the market independently agrees with it. Two different methods, one state.

Built 2026-09-15 from src/stormfront.py. All models scored out of cycle: trained only on elections that had already happened. Market prices from Kalshi, captured because resolved markets do not appear to be retained.