US election data
A working record of how well US election polling has actually done, and where it has gone wrong.
This sits apart from the rest of the site. It is not an investigation and it names nobody. It is a dataset and a set of models, published because the underlying question — how much should you trust a poll? — is one the numbers can answer better than intuition can.
What is in it
1,766 races, 1998 to 2022 — President, Senate, Governor and House — each with the average of every poll taken in its final 21 days, the actual result, and the signed difference between them. Alongside that, a county panel covering seven presidential cycles, 2000 to 2024.
Seven models are scored against those races. Every one is tested out-of-cycle: trained only on elections that had already happened when it made its call, never by shuffling the years together. Three of the seven lose to predicting no error at all, and the page says so. The best of them, Wall Cloud, catches just over half the races the polls call for the wrong side.
What it does not do
It does not forecast winners, and it does not claim to know which way a poll will be wrong. Measured across those 1,766 races, the direction of polling error is not predictable from anything tested here — not the economy, not the previous cycle’s miss, not the make-up of the polling field, not the national environment.
What is partly predictable is how big an error will be, and — separately — which races are likely to be called for the wrong side. Races polled by few firms, or whose polling sits far from the state’s own voting history, miss by more; races that are simply close are the ones that flip. The best model here, Wall Cloud, catches 122 of the 226 races the polls got backwards, at 38% precision. That is the whole of the claim.
Those are different questions and the data separates them cleanly: the model that best predicts the size of an error is nearly useless at predicting upsets, and the page says so on every view.
Where the data comes from
Every number traces to one of these, and each is public:
- FiveThirtyEight poll archive —
raw_polls.csvfrom the pollster-ratings data, cycles 1998 to 2022. One row per poll question, carrying the poll’s own margin and the actual result. FiveThirtyEight was closed in 2025; this is the archived file, and there is no 2024 or 2026 in it. - MIT Election Data and Science Lab — county presidential returns 2000-2024, US Senate statewide returns 1976-2024, US House district returns 1976-2018. Harvard Dataverse.
- US Census Bureau — Citizen Voting Age Population special tabulation, 2006-2010 through 2020-2024, used as the turnout denominator. Total population is not used: it counts non-citizens and would understate turnout precisely in the places where that matters most.
- Bureau of Labor Statistics — Local Area Unemployment Statistics, county by month, 1995-2024.
- Bureau of Economic Analysis — regional tables CAINC1 and CAINC35, county personal income and transfer receipts.
- USDA Economic Research Service — county educational attainment, derived from the American Community Survey.
- Silver Bulletin — published pollster ratings.
- Wikipedia — 2026 general-election polling, for the forward-looking panel only. This is a wiki rather than a curated archive and is treated as lower-trust than everything above; every row keeps its source article and the date it was read.
- us-atlas — county and state boundaries for the map.
How current it is
Updated occasionally. No cadence is promised, because none is kept. The historical record changes only when a source publishes a revision; the 2026 panel changes as polling arrives, and the field thickens through October.