The 2017 tax fight ended with a pile of contradictory numbers and a vote. The Joint Committee on Taxation said the Tax Cuts and Jobs Act would reduce revenue by $1.46 trillion over ten years; Penn Wharton's dynamic analysis pointed to roughly $2 trillion of added debt ;[1] congressional supply-siders waved both off and predicted the cuts would pay for themselves through growth. Somebody had to be wrong. The unusual thing — the reason this chapter exists — is that this time, reality eventually filed a report.
By 2024 the receipts were in. The growth windfall never arrived: revenue did not come in above what the scorekeepers had projected under the law, and the self-financing prediction failed on its own terms. The models fared differently. The Committee for a Responsible Federal Budget compared real, inflation-adjusted revenue for 2018 through 2024 against CBO's 2018 projections and found it within 0.5 percent — setting aside 2022, a year CRFB attributes to a surge of capital-gains realizations and inflation rather than to the tax law .[2] That comparison deserves its caveats stated plainly: receipts landing near a projection grades the forecast, not the law's true cost, because the world without the TCJA was never observed, and inflation, a pandemic, tariffs, and later legislation all moved revenue too. With those confounds acknowledged, the microsimulations were approximately right about direction and magnitude, and the free-lunch prediction was not. The One Big Beautiful Bill Act extended the law's individual provisions in July 2025 ,[3] so the 2018–2024 window stands as a completed natural experiment rather than the end of the story.
But "approximately right" is the best anyone can honestly say.
Three ways to be right
Microsimulation validation happens at three levels, in rising order of difficulty.
The first is component validation: does each individual calculation match the rules? A married couple earning $100,000 with two children owes a specific federal income tax, derivable from the statute, checkable against IRS worksheets or commercial tax software. Either the model computes it or it does not. This is the tractable case — tedious at scale, but tractable.
The second is aggregate validation: summed across the population, does the model reproduce reality's totals? Simulated SNAP enrollment against USDA administrative records; simulated income-tax revenue against IRS collections. A model can pass every component check and still fail here, because the aggregate depends on the population you ran the rules over, not just the rules.
The third is predictive validation: when the model forecast the effect of a change, did the forecast hold? This is the hard one. A policy change never arrives alone; it lands bundled with economic shifts, other reforms, and behavioral responses nobody anticipated, and the model's error is tangled with everything else that happened.
Good models pass the first two levels reliably. The third is where humility enters, and the right posture toward it comes from weather forecasting: tomorrow's forecast is reliable, next week's roughly right, next month's a best guess, and the responsible move is to calibrate confidence to the horizon rather than abandon the forecast.
The ACA test
The Affordable Care Act supplies the best natural experiment we have on the third level, because the CBO made a specific, dated, checkable prediction and then a decade happened to it.
In March 2010, the CBO projected that the ACA would cut the share of non-elderly Americans without insurance from more than 18 percent to about 7.6 percent by 2016, on the assumption — universal in every model at the time — that every state would expand Medicaid . Two years later the Supreme Court made expansion optional, and nineteen states declined: a legal shock no microsimulation had a module for. FactCheck.org later reworked the arithmetic to ask what the 2010 projection implies once you account for the states that opted out; the adjusted figure is 9.4 percent. The measured 2016 rate, from CDC data, was 10.4 percent [4] — within a point of the adjusted projection, across a six-year horizon interrupted by a constitutional ruling.
That is the headline, and it flatters the model. The components tell a different story:
| Coverage source | CBO 2010 prediction | Actual 2016 |
|---|---|---|
| Exchange enrollment | 21–23 million | 10.4 million |
| Medicaid expansion | 10 million | 14.4 million |
| Total uninsured | 30 million | 27.9 million |
Source: [4]
The CBO overestimated exchange enrollment by more than half and underestimated Medicaid enrollment by nearly half. The total came out close in part because the two errors ran in opposite directions and partly cancelled. As the FactCheck.org retrospective put it, the CBO's error "was in estimating where the uninsured would get covered, not how many of them would gain coverage" .[4]
Both readings of this episode are true, and the book needs both. Read generously: predicting a national coverage rate within a point, six years out, through a Supreme Court shock, is a real achievement no rule of thumb could match. Read strictly: an aggregate that lands while its components miss has passed the second validation level by partial cancellation while failing the first — and a decomposition performed years later, with the answers in hand, is not the same thing as the original forecast being right. A number that comes out right for the wrong reasons has not been validated; it has been lucky, and luck does not generalize. The only way to know which you have is to check the components — which is exactly why the levels are ordered the way they are.
The forecaster grades itself
The CBO does something most forecasting institutions never do: it grades its own work in public, publishing systematic retrospectives that compare past projections against what happened .[5] The retrospectives show real improvement. For deficit projections six years out, the average absolute error fell from 3.2 percent of GDP over 1989–2001 to 1.0 percent over 2002–2019 — a threefold gain in two decades, built on richer IRS administrative data, more capable computing, and the retrospectives themselves, which exposed systematic biases that could then be corrected .[5] Self-grading is the feedback loop that makes improvement possible; the agencies that skip it stay exactly as wrong as they started.
The same retrospectives also set the ceiling. CBO's projections are about as accurate as the Blue Chip consensus of some fifty private forecasters, and about as accurate as the Administration's [6] — everyone converges on a similar wall. And the wall moves. The 2021 projection was CBO's largest overestimate on record; the 2023 was its largest underestimate, off by 3.9 percent of GDP, more than triple the historical average .[5] Models tuned on ordinary decades met a pandemic economy and missed in both directions, two years apart.
Then there is the finding I like least and trust most. Kevin Kliesen and Daniel Thornton, economists at the Federal Reserve Bank of St. Louis, analyzed CBO forecasts published from 1976 to 2007 and found that a random walk — assume next year's deficit simply equals this year's — would have beaten the CBO on average, over both short and medium horizons . I cite this against my own interest, since this book spends several chapters on forecasting infrastructure. The right reading says less about CBO than about the problem: deficit forecasting is saturated with irreducible uncertainty — recessions, crises, and pandemics are precisely the events no model foresees — and beating a naive baseline is genuinely hard. Computing a household's taxes under known law and predicting next year's deficit are different problems wearing the same word "model," and an honest accuracy question keeps them apart. The first has an exact answer. The second has a distribution, and institutions that publish it as a point are hiding the interesting part.
Markets, and the people who beat polls
If forecasting is the weak flank, who does it well? One answer that actually gets graded: markets where people bet. An NBER study of prediction markets found their forecasts "weakly more accurate" than survey forecasts — the study's own careful phrase — across GDP, inflation, and employment .[7] The Monday before the 2024 presidential vote, with polls showing a coin flip, Polymarket had Donald Trump at 58 percent .[8] Academic comparisons of market prices against FiveThirtyEight's model found the markets competitive or better .[9]
The other answer is a particular kind of person. For Federal Reserve decisions, Good Judgment's superforecasters reportedly beat financial futures markets by about 30 percent in 2024–2025 .[10] Behind that result sits Philip Tetlock's two decades of evidence: most experts forecast little better than chance, but a small minority — roughly 2 percent of Good Judgment Project participants — consistently outperform, by updating often, thinking in probabilities, and refusing to marry a prediction to an ideology .[11][12]
Markets and superforecasters share one discipline the scorekeepers lack: they are graded, publicly and unavoidably, every time. But they have a structural blind spot that keeps them from replacing anything in this book. A market needs a question that resolves — an event that either happens or does not. It can price will the bill pass; it cannot price what would the bill have done had it passed, because a never-enacted counterfactual resolves nothing. Counterfactuals are exactly microsimulation's terrain. The interesting move, which Part IV builds an institution around, is the synthesis: structural models to compute the counterfactuals, market-style grading to keep the forecasts honest — every estimate published with an interval and scored when the official number lands. As this book went to press, that scoreboard's first forecasts were still waiting on reality; nothing there is validated yet, only staked.
The data are broken too
Everything so far assumed the model's inputs describe the country. They do not, and the person who proved it is Bruce Meyer, an economist at the University of Chicago who has spent two decades documenting a quiet crisis in household survey data .[13]
Meyer and colleagues linked survey responses to administrative records — what people told the Census Bureau against what the program files show — and the gaps are not rounding errors. Roughly 40 to 50 percent of SNAP recipients did not report receiving benefits in the Current Population Survey. More than 60 percent of Temporary Assistance for Needy Families and General Assistance went unreported. About a third of housing-assistance recipients did not mention it; even Social Security was missing from about a tenth of recipients' answers. Feed those surveys to a flawless model and it will confidently understate the safety net's reach, because the people it simulates deny receiving what they receive.
The rot compounds. Survey response rates have fallen since the 1990s, and the people who stop answering skew low-income, young, and mobile, so the Census Bureau increasingly fills the gaps by imputation — statistically inferring the answers people never gave — which means growing stretches of the "data" reflect the imputation model rather than any respondent. Retirees misreport in their own direction: a Census Bureau study linking survey responses to IRS and Social Security records found median income for Americans 65 and older was 30 percent higher in the administrative data than in their survey answers, enough to put measured senior poverty at 9.1 percent when the validated records showed 6.9 .[14] And at the top, the public files are top-coded by design ,[15] compressing exactly the incomes that high-end tax reforms target.
Administrative data fixes much of this and introduces its own holes: tax returns miss non-filers, omit non-taxable transfers, and describe legal tax units rather than the households people actually live in; IRS research files run about two years behind; and no common identifier links records cleanly across agencies. Nor do the rules themselves tell you who shows up: EITC take-up runs around 78 to 80 percent ,[16] SNAP around 82 percent with wide state variation ,[17] and SSI take-up among eligible seniors may run as low as 50 to 60 percent citation pending. A model that assumes everyone claims what the law owes them overstates every program it touches — and take-up moves when policy does, as the 2021 Child Tax Credit expansion showed by reaching non-filers in patterns no prior model would have forecast.
The lesson generalizes: accuracy is as much a property of inputs as of equations. Getting the rules exactly right — chapter 2's project — buys nothing if the population you run them over is a funhouse mirror of the country. That is why the synthesize-and-calibrate strategy the last chapter ended with is a load-bearing wall, and why so much of Part II is about data rather than rules.
When models disagree
In 2017, four institutions put numbers on the TCJA and the spread was enormous: JCT's conventional $1.46 trillion revenue loss; its dynamic $1.07 trillion; the Tax Foundation's far sunnier $448 billion; Penn Wharton's $1.9-to-$2.2 trillion — and even reading that list takes care, because the figures measure different objects. The first three are revenue losses under different feedback conventions; Penn Wharton's is added debt. Part of the apparent chaos is institutions answering subtly different questions.
The rest decomposes into three sources, and naming them is more useful than despairing over the spread. First, data: JCT computes on confidential IRS returns, TPC on less granular public files, and different shops assume different baselines — the assumed path of current law a reform is measured against, a choice that can move an estimate by hundreds of billions before the reform itself is even modeled. Second, behavioral assumptions: how much people rearrange work, saving, and business form when rates change. The Tax Foundation's large assumed growth response versus Penn Wharton's modest one drove more of the gap between them than any other factor — and reality supplied a live example of the stakes, when the TCJA's new 20 percent deduction set off a surge of restructuring into pass-through form, exactly the kind of response every scorekeeper had to guess at in advance. Third, modeling choices: how income shifts between corporate and individual returns, how growth is projected, how expiring provisions are treated.
Model disagreement, decomposed this way, stops looking like failure. It is evidence that policy analysis rests on judgment about data, behavior, and method — and running several models turns a shouting match between conclusions into a comparison of stated premises, which is the only kind of policy argument that can actually be adjudicated.
There is also a quieter mechanism that keeps shared tools honest. Daniel Feenberg maintained TAXSIM for more than four decades, and the thousand-plus published papers that used it formed an informal validation network: every author had an incentive to catch and report the bug that would embarrass their own results, and the model improved under collective scrutiny .[18] For computing federal tax from given inputs, TAXSIM's accuracy is high — the rules are deterministic and documented — with discrepancies clustering in state edge cases and credit interactions; it computes taxes only, record by record, and produces no aggregates. In software testing's vocabulary, TAXSIM became an oracle: an independent implementation of the same rules against which another model's answers can be checked, one record at a time — an idea that chapter 10 promotes from convenience to method.
What can never be checked
In 2017 the CBO estimated that repealing the ACA under the American Health Care Act would cost 23 million people their coverage by 2026 . The bill failed, partly on the strength of that number. So the prediction was never tested — and never can be. The genre is built that way. Enact a reform and the world without it goes unobserved; reject it and the world with it does. A counterfactual cannot be checked against reality, because being unrealized is what makes it a counterfactual.
Which forces the question of what a counterfactual's claim to trust can rest on. The answer is the parts. The whole — "this reform would have cost X and covered Y" — is untestable. But each rule in the model can be proved against the statute; the population can be checked against the census and the administrative totals; the behavioral assumptions can be labeled, sourced, and varied to show what swings on them. The pieces can be verified even when the whole cannot, and a counterfactual assembled from verified pieces is a different epistemic object from one that is turtles all the way down.
Occasionally reality even grades a piece of the whole. The 2021 Child Tax Credit expansion arrived with model-based predictions that it would cut child poverty by roughly 40 percent citation pending. Census data then showed child poverty falling from 9.7 to 5.2 percent in 2021 ;[19] Columbia's poverty center tracked the decline month by month as payments went out ;[20] and when the expansion lapsed, child poverty climbed back. No controlled experiment — the economy was reopening, other pandemic programs were moving — but direction and rough magnitude landed, on one of the few occasions when a modeled counterfactual became enacted law fast enough to watch.
George Box gave the field its permanent epigraph: "all models are wrong, but some are useful" .[21] The record of this chapter sharpens it. The TCJA models were approximately right; the ACA model was right in total and wrong in composition; the deficit forecasts lost to a random walk; the survey data under everything was quietly rotten in measured, correctable ways. Wrong, useful, improvable — and the difference between responsible modeling and false precision turned out, in every case, to be the same thing. Not accuracy. Checkability.
So here is the discipline this book runs on, stated once and enforced everywhere: a simulation is admissible only where its verification chain terminates in ground truth — a component checked against the statute, an aggregate checked against administrative totals, a forecast graded when the official number finally lands. Where the chain holds, a number can be trusted no matter who produced it, agency or amateur. Where it breaks, the number is assertion wearing the costume of arithmetic, and no amount of computational sophistication redeems it.
The rest of this book is an attempt to build things that pass that rule, and an honest accounting of where they do not. It begins, though, somewhere much smaller than a national model: with one household, one benefit cliff, and the frustration that turned a family's spreadsheet into a reason to build any of this at all.
References
- Penn Wharton Budget Model (2017). The Tax Cuts and Jobs Act, as Reported by Conference Committee (12/15/17): Static and Dynamic Effects on the Budget and the Economy.
- Committee for a Responsible Federal Budget (2024). Has TCJA Paid For Itself?.
- 119th United States Congress (2025). H.R. 1, One Big Beautiful Bill Act.
- Kiely (2017). CBO's Obamacare Predictions: How Accurate?.
- Congressional Budget Office (2024). An Evaluation of CBO's Projections of Deficits and Debt From 1984 to 2023.
- Congressional Budget Office (2025). CBO's Economic Forecasting Record: 2025 Update.
- Wolfers (2012). Prediction Markets for Economic Forecasting.
- Polymarket (2024). 2024 Presidential Election Winner.
- Crane (2020). Prediction Market Accuracy in the Long Run.
- Good Judgment Inc. (2024). Superforecasters Outperform Markets on Fed Predictions.
- Tetlock (2005). Expert Political Judgment: How Good Is It? How Can We Know?.
- Tetlock (2015). Superforecasting: The Art and Science of Prediction.
- Meyer (2015). Household Surveys in Crisis.
- Bee (2017). Do Older Americans Have More Income Than We Think?.
- Larrimore (2008). Consistent Cell Means for Topcoded Incomes in the Public Use March {CPS.
- Internal Revenue Service (2024). EITC Participation Rate by States, Tax Years 2012 through 2019.
- USDA Food (2024). Trends in Supplemental Nutrition Assistance Program Participation Rates: Fiscal Year 2019 to Fiscal Year 2021.
- Feenberg (1993). An Introduction to the TAXSIM Model.
- Fox (2022). The Supplemental Poverty Measure: 2021.
- Parolin (2021). Monthly Poverty Rates among Children after the Expansion of the Child Tax Credit.
- Box (1976). Science and Statistics.