Chapter 13

The uncertainty gap, and the scoreboard

3,698 words · 19 min · draft, July 2026


In May 2017 the Congressional Budget Office published its analysis of the American Health Care Act, the Republican bill to repeal and replace much of the Affordable Care Act. The headline finding: by 2026, 23 million fewer Americans would have health insurance than under current law . Not "approximately 23 million." Not "between 18 and 28 million."

Just: 23 million.

The number did what headline numbers do. The bill's supporters attacked it, its opponents carved it into granite, and for weeks the debate treated 23 million as a measurement — a fact to be defended or debunked — rather than what it was: an estimate of the gap between two futures, made nearly a decade before either would arrive. CBO had not erred. Twenty-three million was a defensible central estimate, produced by careful people using the best available methods. But a nine-year forecast of insurance coverage rides on economic conditions, behavioral responses, and a labor market nobody controls, and a single integer expresses none of that. It is 2026 now, the year the forecast pointed at, and the number can never be checked: the bill died in the Senate that July citation pending, so the world it described never ran. The estimate was precise, uncertain, and ungradable all at once, and nothing in how it was published let anyone see the second and third properties.

The 2002 film Minority Report built its dystopia on this exact suppression. Three "precogs" foresee murders; the police arrest people before they kill; the predictions display as facts, exact to the minute and the street corner, with no probability attached. The twist is that the system produces three forecasts that usually agree and sometimes diverge, and the dissenting vision — the minority report — is filed away to preserve the illusion of certainty. Every budget score that ships as a single number files away a minority report of its own. The interval exists, down in the sampling error and the contested elasticities. It just never makes the page.

Every result in this book so far has been a single number. A taper cut costs £2.8 billion. A carbon dividend cuts poverty 10 percent. A single mother faces an 80 percent marginal rate. That is not a defect of one model; the microsimulation paradigm, from Orcutt forward, was built to answer "what would happen if?" — it produces scenarios, not probability distributions. Nor is it ignorance: statisticians have known how to put standard errors on microsimulation output for decades. The headline scores that decide votes simply ship without them. This chapter is about what those clean figures hide, what it costs to expose it, and the institution I am helping build to put the hidden part on a public scoreboard.

Four kinds of not-knowing

Start with where the model's people come from. The Current Population Survey's March supplement asks fewer than ninety thousand households to stand in for a country of more than 130 million. For a national aggregate that is plenty; for the cell you actually care about — self-employed workers in Montana with investment income — it may be a handful of records, and the sampling error in each cell propagates when the model sums weighted impacts across the country. Calibration helps: the machinery chapters 6 and 7 described cut the Enhanced CPS's deviations from its administrative calibration targets by roughly 97 percent. to verify But calibration pins down the questions the government already publishes answers to. It does nothing for the variance of a question nobody has posed. Call this input-data uncertainty: the survey is a sample, and every number computed from it inherits the sample's noise.

Second, the parameters. Tax brackets and benefit rates are usually known exactly — the statute prints them. Behavior is not. A meta-regression of 1,720 estimates from 61 studies put the mean elasticity of taxable income near 0.30, with individual estimates running from near zero to above 1.0 depending on method, sample, and country .[1] The stakes are not academic: a reform raising the top marginal rate by ten points might add $100 billion in revenue at an elasticity of 0.2, or $60 billion at 0.5. Same reform, same model, opposite headlines. Parameter uncertainty is the honest name for a number the literature brackets and the model hard-codes.

Third, the model itself. Every microsimulation encodes assumptions about how programs interact, how households respond, how the economy adjusts — and different assumptions, written into code by different reasonable people, produce different results. Structural uncertainty is the possibility that the model's shape, not just its inputs, is wrong. It is the hardest kind, and it gets its own section below.

Fourth, the future. Any projection past the current year rests on forecasts of wage growth, inflation, and demographic change, which enter the model as fixed numbers and are, of course, forecasts. Future uncertainty is the oldest kind, and models swallow it whole.

Each layer compounds the ones before it. The tidy figure on the page is the sum of four kinds of not-knowing, reported as one kind of knowing.

Why the single number survives

False precision breeds false confidence. A number published without a range invites its audience to defend it to the decimal, which is how legislators end up litigating 23 million as if it were a census count. When a model reports that a reform costs £12 billion, the question that comes back is "is it worth £12 billion?", almost never "how sure are you?" Score two competing reforms at $50 billion and $48 billion and the $2 billion gap reads as a real difference, even when each estimate carries ±$10 billion and the gap is noise wearing the costume of a finding. The blindness runs deeper than optics. A point estimate cannot say "there is a one-in-ten chance this costs twice the projection," and the downside case is often what a policymaker actually needs to weigh. And when two models disagree — Tax-Calculator says one thing, PolicyEngine another — a reader without intervals cannot tell whether one is wrong or both sit comfortably inside the same uncertainty.

So why does the single number persist? Because everyone downstream needs it. Reconciliation rules require one cost figure, not a distribution. A journalist needs a headline. An advocate needs a talking point — "this reform lifts two million children out of poverty" mobilizes in a way "between 1.4 and 2.6 million" never will. Producers of estimates know they are uncertain; consumers treat them as certain; producers, learning what gets quoted, stop reporting the spread. The equilibrium is stable and everyone in it is behaving rationally, which is exactly why no one inside it will fix it.

The producers are not oblivious, and the record here deserves stating fairly. The Penn Wharton Budget Model, one of the few groups to compare its projections to outcomes systematically, found its estimates generally accurate but with meaningful variance .[2] CBO grades its own aggregate work in public: its June 2024 projection underestimated fiscal-year 2025 revenues by about 6 percent — $334 billion — in significant part because it had not anticipated a round of tariff increases .[3] Across 1984 to 2023, its retrospectives show revenue is harder to forecast than spending and that errors compound with the horizon ;[4] its economic forecasts tend to be more accurate than the Administration's and the Blue Chip consensus, and run broadly similar to the Survey of Professional Forecasters .[5] That is real self-scrutiny, rarer than it should be.

But notice what it is: an evaluation of the aggregate forecast, years later, in a study. No standing public loop grades each policy score, one by one, against what happened. Part of the reason nobody built one is deep rather than lazy: most policy scores are comparisons between the world that ran and a world that never will. The AHCA's 23 million was a difference between two branches, and the world ran one. You cannot look up a counterfactual. Hold that thought; the institution at the end of this chapter is designed around it rather than in denial of it.

Putting bands on the number

The methods for exposing uncertainty exist, and each buys a different slice of it. Monte Carlo simulation runs the model many times with inputs drawn from probability distributions and reports the spread. I use it in EggNest, a retirement-planning tool I built: instead of "you will have $1.2 million at 65," it runs ten thousand scenarios over distributions of market returns, inflation, and wage growth and reports "a 90 percent chance of between $800,000 and $1.8 million." The first formulation is what people want to hear; the second is what the arithmetic supports. For policy microsimulation the obstacle is cost. The engine already applies every bracket and phase-out to an entire simulated country; running that ten thousand times multiplies the work by four orders of magnitude, and seconds become hours.

Bootstrap resampling is cheaper and narrower: draw many samples from the microdata, reweight each to population totals, run the policy on each, and read the distribution of results. The computation is embarrassingly parallel: 500 replicates add minutes on a cluster, not hours. What it captures is input-data uncertainty, the fact that the survey is a sample rather than a census, and nothing else; a bootstrap says nothing about a wrong elasticity or a wrong model. But for the everyday question — how much should I trust this number? — input-data uncertainty is usually the leading term, which makes the bootstrap the right place for an honest program to start.

Beyond those two sit the deeper renovations. Bayesian methods treat a contested elasticity as a distribution — a prior centered near 0.30 that updates as evidence accumulates — instead of a constant, at the price of rebuilding models that hard-code the number. Squiggle, a probabilistic-programming language, propagates explicit distributions through a calculation ;[6] it suits a Fermi estimate and fights you across ten thousand lines of microsimulation. Scenario analysis reports a baseline, an optimistic case, and a pessimistic one, conveying sensitivity without pretending to know the probability of each branch.

To make the leading term concrete, I ran a small paired-subsample experiment in PolicyEngine US on March 31, 2026. I drew ten subsamples of 1,000 households each from the then-current Enhanced CPS, ran the same stylized reform to the Earned Income Tax Credit on each baseline-and-reform pair, and recorded the change in aggregate EITC and in household net income. On the full sample, the reform cut aggregate EITC by $12.6 billion. Across the ten subsamples, the same reform cut it by anywhere from $11.1 billion to $15.0 billion, with a mean of $12.9 billion; net income moved on the same pattern, a $13.9 billion full-sample drop against a range of $12.1 to $16.2 billion. Ordinary survey sampling alone — before any behavioral elasticity, any forecast, any modeling assumption — moved the national estimate from about 12 percent below the point estimate to about 19 percent above it. A swing that size carries a reform across budget thresholds the single figure made look safe.

The caveats, flatly: this is not a production-grade confidence interval. The subsamples are deliberately small and few, the exercise was a one-off run by hand, and it captures input-data uncertainty only — none of the other three kinds. What it demonstrates is narrower and still worth having: the number the model reports is the center of a distribution, not the result.

Making that demonstration routine instead of artisanal is a data-infrastructure problem, and it is one of the jobs populace — the calibrated-microdata commons Part II introduced — exists to do. The machinery that produced the imputations and the weights, from the quantile regression forests of chapter 7 to the synthesis across integrated sources and the calibration against administrative totals, ships inside certified release bundles rather than vanishing behind a finished file. That means imputation and synthesis error can be quantified and versioned instead of buried: a later analyst can re-derive the weights and measure how much they could have differed. The March 31 experiment took an afternoon of hand work. The point of the infrastructure is that it should take a flag.

The weather standard

One field solved this problem so thoroughly that we forgot it was ever a problem. Modern weather models run an ensemble — dozens of simulations from slightly perturbed starting conditions — and report the spread as a probability. "A 60 percent chance of rain" is a calibrated probability: across many such forecasts, it rains on close to 60 percent of the days so labeled.

Forecasters call that property calibration, and it is worth pausing on the word, because this book has been using it for something else. Data calibration — chapter 2's kind — adjusts survey weights until a dataset reproduces administrative totals that already exist. Forecast calibration is a property of stated probabilities: events assigned 60 percent happen about 60 percent of the time. The first is checked against the present. The second can only be earned against the future, by tracking whether the 60-percent days deliver rain at the promised rate and recalibrating when they drift. The loop is a procedure, not a metaphor: forecast, resolve, score, recalibrate. Weather forecasting has run it for decades, on millions of forecasts, and its skill has improved steadily and measurably because every forecast is eventually checked against a sky that either rained or did not.

Other fields run partial versions. Clinical trials report effects with confidence intervals, and "15 percent lower mortality (95% CI: 8–22)" is a different finding from "15 percent (95% CI: −2–32)," even though both center on 15. Financial risk management lives on Value at Risk and implied volatility; the 2008 crisis exposed how badly those models underweighted tail risk, and the response was better uncertainty modeling, not abandonment. Climate science runs large model ensembles and writes in calibrated language — the IPCC's "likely" means at least 66 percent, "very likely" at least 90. Election forecasting has absorbed the grammar too: a model that gave one 2016 candidate a 29 percent chance was not refuted when he won, because 29 percent is not zero. The structural parallel to policy modeling is exact — complex nonlinear systems, many interacting variables, forecasts people stake decisions on. The difference is that meteorology spent decades building the verification loop, and policy analysis has retrospectives but has never run the loop: no committed forecast, no fixed resolution date, no score, no recalibration.

When the model itself is wrong

Bootstraps and ensembles put bands around a model's answer. They cannot say whether the model asks the right question. You can propagate uncertainty through a wrong model and get a precise answer that is precisely wrong. Structural error is the uncertainty we do not know we have, and nothing inside the model catches it. Only reality does — which is a hard sentence for a modeler to type, and the evidence for it comes from the two natural experiments this field has actually been handed.

In 2017 and 2018 the Finnish government ran a randomized controlled trial of basic income: 2,000 unemployed people received €560 a month, unconditionally, measured against a control group of about 175,000 to verify. Microsimulation had a clear prediction. The design cut participation tax rates by 23 percentage points, and in the models, cutting participation tax rates raises employment. The first-year result: no statistically significant effect on days worked .[7] The eventual published analysis found modest gains in the second year, concentrated in particular subgroups .[8] The models had the incentive exactly right — participation tax rates really did fall 23 points — and the response wrong, because a permanent unconditional floor is not a marginal tax tweak, and the historical data the response parameters were fitted on contained nothing like one.

The Affordable Care Act's individual mandate tells the same story from the other direction. When Congress zeroed the mandate penalty in 2017, CBO projected 13 million fewer insured Americans by 2027, from models in which the mandate drove enrollment . The effect came in far smaller. People had enrolled for the subsidies, for the security, for the momentum of already being enrolled — and the historical record could not separate "enrolled because of the mandate" from "enrolled while the mandate happened to exist," so the models credited the mandate with behavior that belonged to everything around it. Finland's models overestimated a response to incentives; the mandate models overestimated a response to compulsion. Both were reasonable readings of history. Both were wrong in a way that no bootstrap, no prior, and no ensemble could have flagged in advance, because the flaw was in the model's shape, and only a committed forecast that missed by more than its interval allowed could have caught it.

The scoreboard

The last chapter split the stack into a pole that states and a pole that predicts. On the stating side the check is immediate: an encoded rule is right or wrong against the statute and the reference calculators the day it is written. The predicting side has no such settlement. The four uncertainties live in the gap between a deterministic calculation and a claim about the future, and no statutory exactness closes that gap. The only thing that closes it is the weather forecaster's loop — make the claim in advance, in public, and let the world grade it.

The Thesis Institute is that loop built as an institution :[9] the last chapter's public docket of forecasts of government metrics, each locked with an interval before the number is known and graded when the official number lands. Remember the problem with grading policy scores: the counterfactual never prints. The docket is built to respect that rather than promise around it. Its unconditional entries — next month's unemployment print, next quarter's caseload under the law as it stands — resolve directly. Its conditional entries, the Medicaid wait-time page among them, resolve only on the branch the world takes; on the branch not taken, the entry expires unscored, and the ledger says so. The causal gap between branches — the thing the AHCA's 23 million actually claimed — never resolves at all, and the ledger labels which kind of object each entry is instead of letting a scenario impersonate a testable claim.

That is weaker than grading every policy score. It is also the honest version, and it still transforms the field it touches: each estimate the earlier chapters produced that names something reality will print — a caseload, a revenue line, the poverty rate the Census Bureau will publish — can now become a forecast with a published interval and a resolution date, held to the weather standard: verified probabilities, recalibrated on misses. The discipline this book keeps applying to other people's tools lands hardest here, on the institution built to enforce it.

There is no track record — only a mechanism for acquiring one, in public.

Reading a number until then

Grades take years to accumulate, and policy will not wait. So the practical posture, for however long the scoreboard stays young: treat every point estimate as the center of a distribution. Read "£12 billion" as "our best guess is near £12 billion, and the true value could reasonably sit twenty percent to either side, or further." Read two models' disagreement as information about the uncertainty rather than a fight to referee. Ask which assumptions move the answer most, because the answer usually has one or two hinges. Reserve the deepest skepticism for the most novel policies — Finland's lesson — where structural uncertainty dominates and no historical data has the answer's shape. And do not mistake imprecision for uselessness: "roughly $50 billion" rules out $5 billion and $500 billion, which is often all a decision needed to know.

Here is what that looks like in practice, on a reform this book has touched before. PolicyEngine scores making the Child Tax Credit fully refundable — removing the refundable cap and the earnings phase-in that hold the poorest families below the full credit — against 2026 law .[10] The static result: a cost of roughly $23 billion for the year, with child poverty falling about 2.6 percentage points, from 16.6 to 14.0 percent. That is my own calculation, run at a specific model version on a specific data vintage, to be rerun before this book prints — not an official score. The uncertainty-aware version attaches a band to the cost, driven mainly by one behavioral unknown: how many currently non-filing families would actually claim. And it attaches a second band to the poverty effect, driven by the systematic undercount of income among low-income households in survey data. The central estimates do not move. What changes is what you can honestly say: the poverty reduction survives even the pessimistic branch, while the cost has real range. A legislator holding the banded version knows more than one holding a clean $23 billion — including knowing which part of the claim is sturdy and which part is a bet.

The scoreboard can eventually grade all of this, because all of it names things reality prints: an unemployment rate, a caseload, a call-center wait time, a child-poverty figure with a release date. But the docket has an edge, and standing at it you can see the next problem. What people believe about a policy, what they prefer, what they would trade away to get it — none of that has a first print. No administrative office certifies it; no future date resolves it. The next chapter steps past the edge of what ground truth can grade, into the simulation of opinion itself, carrying the same discipline into territory where the ground is soft.

Society in silico · draft in public · source