Chapter 2

The tax model wars

3,870 words · 20 min · draft, July 2026


In 1983, a small London think tank did something no outside body had done before to verify: it built a working model of the British tax and benefit system and turned it on the government's own budget. The model was called TAXBEN, and it belonged to the Institute for Fiscal Studies. When the Chancellor announced new rates and thresholds, IFS economists could feed the rules in and simulate their effects on families up and down the income distribution within hours — publishing who gained and who lost before the official gloss had time to set. The institute's annual Green Budget became required reading for journalists, civil servants, and the politicians it embarrassed.

The institution behind the model began, like a surprising number of things in this book, with frustration at a tax law nobody could decode. The 1965 Finance Act had introduced capital gains tax to Britain in a form so opaque that four financial professionals — the banker Will Hopper, the investment-trust manager Bob Buist, the stockbroker Nils Taube, and the tax consultant John Chown — decided the country needed an institution whose job was understanding fiscal policy from outside the government. They settled it over dinner on July 30, 1968, at the Stella Alpina restaurant in London, and incorporated the Institute for Fiscal Studies the following year .[1] Fifteen years from a dinner-table grievance to a model that could check the Treasury's arithmetic: that is roughly the speed at which this kind of independence gets built, and it is worth remembering later, when the building speeds up.

TAXBEN was a tax-benefit model — a microsimulation that computes both the taxes a household owes and the benefits it receives, capturing how the two interact. The distinction from a mere tax calculator matters more than it sounds, because the sharpest effects in fiscal systems live in the interaction: a family that pays little tax can still lose most of a raise to benefit withdrawal, and no model that sees only one side of the ledger will ever show it.

But before the outsiders could fight the government's numbers, the government had to build them. That machine is where the war starts.

Washington builds the machine

In 1974, after Richard Nixon impounded funds Congress had appropriated, Congress created the Congressional Budget Office to give itself analytical capacity independent of the executive branch. Its first director was Alice Rivlin — the same Alice Rivlin who, as Guy Orcutt's doctoral student, had co-authored the 1961 book that first made microsimulation run. She wanted a nonpartisan office willing to publish numbers that discomfited both parties, and she staffed it with economists rather than political operatives .[2] The culture held: CBO built microsimulation models for health insurance, Social Security, and tax analysis, and for long-run questions it built CBOLT, a long-term model of the entire federal budget .[3]

Congress's tax arithmetic ran through an older body. The Joint Committee on Taxation had existed since 1926 ,[4] and by the 1970s its staff had developed a suite of microsimulation models — an individual income model, a corporate model, an estate and gift model, an international model for cross-border flows .[5] The budget process built on the Congressional Budget Act of 1974 made JCT's numbers the official revenue estimates for tax legislation, and "official" is doing heavy work in that sentence. When a senator drafts a tax amendment, JCT staff feed it into the models, which apply the proposed rules to a large, confidential sample of actual tax returns; within days, sometimes hours, out comes a score — the projected change in federal revenue over ten years. That single number decides whether the amendment fits the budget resolution, survives reconciliation, or is fiscally viable at all. A score too high kills a proposal before any vote. This is budget scoring: a gate legislation must pass through, operated by a model most members of Congress have never seen.

The executive branch had its own. Treasury's Office of Tax Analysis had run computerized tax models since George Sadowsky's work in the early 1960s [6] — Sadowsky later called that first, three-month effort "probably the most useful program I ever wrote" ,[7] a fair self-assessment for a program that taught the federal government to test tax law before enacting it. By the 1980s his hack had matured into the Individual Income Tax Model, refreshed regularly from IRS returns, producing the revenue estimates behind each Administration's proposals as JCT's served Congress.

And because taxes were only half the ledger, the Urban Institute — Orcutt's old home — ran the TRIM series, updated every year from the Current Population Survey to model the programs the tax models left out: SNAP, Medicaid, TANF, SSI, housing assistance. It captured the interaction between taxes and transfers, the thing TAXBEN had shown mattered. It served government agencies under contract; the public had no access.

By the mid-1980s, then, the United States had a complete analytic apparatus for fiscal policy: models at CBO, JCT, Treasury, and Urban, collectively capable of telling you what nearly any tax or benefit change would do to revenue and to households. It was a genuine intellectual achievement. It was also, from the outside, a set of black boxes — and the walls around them were not primarily made of code.

The moat

Section 6103 of the Internal Revenue Code forbids the disclosure of individual tax-return information, with a narrow set of exceptions. Under section 6103(f), the tax-writing committees and the Joint Committee on Taxation have statutory access to the returns themselves; Treasury's analysts hold theirs as part of administering the tax system, under 6103(h). External researchers have neither.

The consequence is a permanent asymmetry in raw material. The government's models run on large, carefully edited samples of actual returns, drawn by the IRS's Statistics of Income division, carrying detail no outsider ever sees — the joint pattern of incomes, deductions, and credits as filers actually report them. Everyone outside works from public-use files: sampled down, anonymized, and top-coded, meaning the highest incomes are collapsed into a single category or swapped for averages so that no billionaire can be re-identified from the data. Top-coding is a reasonable privacy safeguard with an awkward side effect: it blurs the data exactly where the biggest tax fights are decided. Debates about taxing the top one percent are conducted, outside government, on files designed to conceal the top one percent.

The rule exists so nobody can look up a neighbor's tax return. Its byproduct is an analytical moat. You can publish your code, your methods, and your assumptions, and you still will not have the data.

The moat had costs that compounded quietly for decades. An estimate people disliked could be dismissed as biased, and there was no external check to settle the question. An estimate that was simply wrong could not be reconstructed from outside, so errors survived. Debate narrowed to whatever the scorekeepers were willing and able to model — a proposal the models could not score was, politically, a proposal that could not exist. Expertise pooled inside a few institutions, because nowhere else could a tax economist practice on real machinery. And citizens were asked to take the numbers on faith, which worked about as well as faith usually does in a polarized country: each side kept its own arithmetic.

The rest of this chapter is the fifty-year campaign against that arrangement. The law forbids breaching the moat, so the outsiders rebuilt, in public, nearly everything inside it.

The outsiders arm themselves

The first serious tool arrived from academia. At the National Bureau of Economic Research, Daniel Feenberg began building TAXSIM in the early 1980s: a program that computed federal and state income-tax liability for any household you described to it .[8] By the 1990s TAXSIM was reachable over the internet — one of the first tax calculators anyone could run remotely — and Feenberg maintained it with a craftsman's persistence for the rest of his career. Economists used it in over a thousand published papers. What TAXSIM did not do was politics: it calculated taxes for individual records but produced none of the aggregate revenue estimates and distributional tables that drive legislative fights.

The institution built for the fights came in 2002, when Gene Steuerle, Len Burman, and William Gale founded the Tax Policy Center as a joint venture of the Urban Institute and the Brookings Institution .[9] The founders knew the machine from the inside: Steuerle had helped build the official models at Treasury, Burman at CBO and Treasury. TPC built its microsimulation on the same public-use files available to everyone outside government — less granular than the confidential returns, but enough to estimate the revenue and distributional effects of a proposal within days.

Its arrival changed what a campaign promise could get away with. In 2012, Mitt Romney ran on a tax plan that promised to cut rates sharply while remaining revenue-neutral through unspecified base-broadening. TPC ran the arithmetic and showed the promise was mathematically impossible: no combination of the available base-broadeners could pay for the rate cuts without raising taxes on the middle class or losing revenue .[10] The campaign could dispute the finding; it could not rebut it with numbers, because it had no comparable model and the one that existed said no. Five years later, during the debate over the 2017 Tax Cuts and Jobs Act, TPC's tables showed how heavily the bill's benefits tilted toward high-income households — analysis Congress's official distributional tables were never going to lead with.

Others followed. Kent Smetters launched the Penn Wharton Budget Model at the University of Pennsylvania in 2016, distinguishing it by producing dynamic estimates — scores that let a tax change move the size of the economy rather than holding it fixed .[11] The Budget Lab at Yale arrived later with another lens. In little more than a decade, the number of independent shops able to score a major tax proposal went from essentially zero to half a dozen, each with a discernible temperament: the Tax Policy Center and the Institute on Taxation and Economic Policy leaned left on distributional questions, the Tax Foundation leaned right on growth effects, Penn Wharton positioned itself in the middle. Pluralism had replaced monopoly — at least for anyone senior enough to know which shop's assumptions to discount.

The dynamic-scoring war

With multiple models came a fight about what a score should even measure, and it is the clearest case study in how methodological choices carry political weight.

Traditionally, JCT scored a tax cut against an economy of fixed size: cut revenue by $100 billion and the score says $100 billion. The convention is usually called static scoring, and the label misleads in an important way — conventional estimates already let people respond to the law, shifting the timing of income, reorganizing businesses, rebalancing portfolios (behavioral responses, in the trade's vocabulary: the ways filers rearrange their affairs when the rules change). What the convention holds fixed is the macroeconomy itself: total output, employment, investment. Supply-siders argued this systematically overstated the cost of tax cuts, since lower rates should spur growth that claws some revenue back. Score the same cut dynamically — letting it change the size of the economy — and it might book at $70 billion instead.

The academic verdict, worked through in forums and papers across the 2000s and stated formally by Alan Auerbach in 2005, was that dynamic scoring is conceptually correct and practically treacherous: the macroeconomic models needed to estimate the feedback are themselves deeply uncertain, so the choice of model smuggles in the answer .[12] Congress made the choice anyway. In January 2015, the new Republican House majority required JCT to produce dynamic scores for any legislation with budgetary effects above a quarter of a percent of GDP .[13]

The Tax Cuts and Jobs Act of 2017 became the first major test, and the numbers are worth keeping straight because chapter 3 grades them. JCT's conventional score put the ten-year revenue loss at $1.46 trillion . Its dynamic score came in near $1.07 trillion: faster growth adding roughly $451 billion of revenue, higher interest costs eating $66 billion of that, for a net deficit reduction of about $385 billion — trimming the bill's cost by roughly 27 percent, a real effect and nothing like self-financing .[14] Penn Wharton, running the feedback through its own model, judged the growth effects more modest still: by its estimate the law would add $1.9 to $2.2 trillion to federal debt over the decade even after growth .[11] No credible model found a free lunch. The law passed anyway, its individual provisions scheduled to expire in 2025 — a schedule Congress then overrode, extending them that July in the One Big Beautiful Bill Act .[15]

A decade into the experiment, Douglas Elmendorf, Glenn Hubbard, and Heidi Williams — a former CBO director, a former chair of the Council of Economic Advisers, and one of the country's leading empirical economists — assessed dynamic scoring and concluded it had improved budget analysis at the margins without resolving the fundamental uncertainty about macroeconomic feedback .[16] That is roughly where the fight rests: the models multiplied, the assumptions became somewhat more explicit, and the deepest disagreements moved from arithmetic into premises, where at least they can be argued about honestly. Whether any of the numbers were right is a different question, and the next chapter takes it up directly.

The idea crosses borders

While Washington fought over scoring conventions, the method itself was going global, and each country's version tells you something about what independence requires.

Australia got an institution. From 1993, Ann Harding built a microsimulation program at the National Centre for Social and Economic Modelling in Canberra; her edited volume Microsimulation and Public Policy served for years as the field's standard reference ,[17] and NATSEM's models turned up everywhere from government agencies to parliamentary committees to the evening news, on everything from tax reform to childcare.

Europe got a federation. From 1996, a team led by Holly Sutherland set out to build EUROMOD: one tax-benefit model spanning, eventually, every EU member state, with national systems harmonized so a reform in Spain could be compared with one in Finland on the same terms .[18] It was developed at the University of Essex on European Commission research grants; researchers applied for access, the ethos was sharing even where the code was not yet fully open, and the Commission eventually took over its maintenance through its Joint Research Centre. EUROMOD is the quiet giant of this field — twenty-seven countries' rules, maintained for decades, validated country by country — and later chapters lean on it as a reference point.

Britain got a spin-off. With funding from the Nuffield Foundation, a project led by the economist Mike Brewer turned EUROMOD's UK component into a standalone model, UKMOD — built and housed at Essex's Centre for Microsimulation and Policy Analysis, which Matteo Richiardi directs, and introduced in a 2021 paper by Richiardi, Diego Collado, and Daria Popova .[19] "We wanted to democratize access to tax-benefit analysis," Brewer said of the project citation pending. Its first free public release, in October 2019, made it the UK's first freely available tax-benefit model, and public bodies picked it up — the Scottish Parliament's research service, NHS Health Scotland, the Welsh Government citation pending.

And individuals could do it alone. Howard Reed ran TAXBEN inside the IFS from 2000 to 2004 ,[20] spent a spell as chief economist at IPPR, then founded Landman Economics in 2008 and built his own tax-transfer model, which he pointed at basic income, welfare reform, and the cumulative toll of austerity — work in service of an ambition he described as "a new settlement of the same scale and sustainability as the Beveridge-inspired reforms of 1945" citation pending. Malcolm Torry, directing the Citizen's Basic Income Trust, used EUROMOD across nearly a decade to cost basic-income schemes in granular detail .[21] Independence of analysis, it turned out, could follow from a freely downloadable model as surely as from a new institution.

The French break

The sharpest conceptual jump came from France, a country that gave its parliament no analytical shop at all — no CBO, no independent scorer; evaluating a proposed amendment meant asking the Finance Ministry to evaluate the proposal it might well oppose.

In January 2011, Camille Landais, Thomas Piketty, and Emmanuel Saez published Pour une révolution fiscale — "For a Tax Revolution" — arguing for restructuring the French income tax into a single progressive levy .[22] The book mattered less than its companion. At revolution-fiscale.fr, any French citizen could design a tax reform and watch its budgetary and distributional consequences resolve on screen; hundreds of thousands did .[22] For the first time anywhere to verify, the general public could run the kind of analysis that finance ministries ran, on a question the finance ministry would rather not have publicly analyzed.

Four months after the book appeared, a small team inside the French government's own strategy unit — the Centre d'analyse stratégique, later renamed France Stratégie — began building something quieter and more consequential: OpenFisca, written by Mahdi Ben Jelloul and Clément Schaff, released to the public later that year. to verify It was the source code beneath a simulator, built on the premise that a country's tax and benefit rules should exist as executable code alongside the legal prose .[23] Where the révolution fiscale site opened one model's results to the public, OpenFisca opened the machinery — anyone could read the encoded rules, correct them, or build a different tool on top.

The premise acquired a name — rules as code, with variants like "legislation as code" and "machine-consumable rules" — and a small international movement. In 2018, New Zealand's Service Innovation Lab spent three weeks translating a piece of legislation into Python and human-readable rules simultaneously, demonstrating that a law could be drafted in both forms at once .[24] The OECD published Cracking the Code in 2020, a primer for governments on the approach .[25] OpenFisca itself spread across four continents — powering France's Mes Aides benefits calculator, a New Zealand rates rebate, adaptations in Tunisia and Senegal — and collected an OECD innovation award and recognition from the European Commission as its most innovative open-source software .[23] At the 2016 Open Government Partnership summit in Paris, a team of French and Tunisian volunteers modeled Senegal's income tax in under thirty-six hours and won the hackathon .[23] By 2019, France's National Assembly had built LexImpact on top of OpenFisca, letting legislators simulate the effects of proposed amendments; within two years the tool had informed well over a hundred simulations in parliamentary debate.

Keep the Senegal number in mind — thirty-six hours, volunteers, one tax — because a version of it, scaled to entire countries and verified against reference models, is where this book eventually goes.

The American open-source turn

The United States' contribution to the open movement came from an unexpected quarter: the American Enterprise Institute, a market-oriented think tank not usually accused of undermining institutional authority. In 2013, Matt Jensen founded the Open Source Policy Center there [26] and recruited Martin Holmer, an MIT-trained economist, to write Tax-Calculator: an open-source model of US federal income and payroll taxes, in Python, with its code on GitHub where anyone could read the formulas, run the model, and check the results .[27] The provenance did the project a strange favor. Open modeling pitched from the left would have been read as advocacy in disguise; coming from AEI, it was harder to dismiss as partisan, and contributors arrived from across the spectrum.

Tax-Calculator grew into the Policy Simulation Library, a loose confederation of independent open-source models sharing transparency and interoperability standards while each specialized — overlapping-generations macro models for dynamic scoring, a capital-cost-recovery model from the Tax Foundation, computational-economics tools from QuantEcon — with contributors ranging from CBO staff to the City of New York. Even the Federal Reserve joined the drift toward glass boxes, releasing a Python implementation of FRB/US, its several-hundred-equation macroeconomic model, in 2022 .[28]

By the early 2020s, the open movement had produced something remarkable and something incomplete. Remarkable: working, credible, inspectable models of major tax systems, free to anyone. Incomplete: every one of them demanded a programmer. The engine was open. The last mile — to something an ordinary person could use, and rely on — was not yet built. A tool called PolicyEngine would eventually put the open approach in front of anyone with a web browser; it enters this story in Part II.

What the moat still held

Step back from the fifty-year campaign and audit what it won. TAXSIM, TPC, Penn Wharton, OpenFisca, Tax-Calculator: every one of these was a gain in method — open code, published assumptions, reproducible results. None was a gain in data. Section 6103 stood exactly where it stood in 1976, and it should: the confidentiality of tax returns is not a bug to be routed around, because you cannot open-source a dataset you are legally forbidden to hold.

And rules as code, for all its promise, has a second limit hiding inside the first. Encode a country's rules and you get an engine that can compute any individual scenario — this household, that income, those benefits. But the questions legislation turns on are aggregate: what does the reform cost, who gains, how many fall below the poverty line. To answer those, the engine needs something it does not contain — a representative population to run the rules over — and building one runs straight back into the moat. The returns that would make the population faithful are the returns nobody outside may touch.

There is exactly one answer the constraint permits, and it is the strategy the rest of this book builds on. Take the public files you are allowed to hold. Synthesize from them an artificial population — households that exist only in silicon. Then calibrate it: adjust each record's weight, the number of real households it stands for, until the weighted totals reproduce the administrative aggregates the government does publish — total returns filed, total wages by bracket, total SNAP enrollment by state. The government will not show you the microdata, but it publishes thousands of true totals about it, and each one is a constraint an artificial population can be forced to satisfy. Statistics Canada had proven a version of the idea decades earlier, shipping its SPSD/M model on exactly such a synthetic database; what remained was to build it at the scale and openness the moment now demands.

You do not breach the moat. You reconstruct, statistically, what stands behind it.

Whether the reconstruction can be made faithful enough is a question that occupies much of Part II. But it sits inside a harder and older question, one the closed apparatus spent fifty years half-avoiding and the open movement now had to face about its own work: how would you know whether any of these models is right?

References

  1. Institute for Fiscal Studies (2024). History of the IFS.
  2. Rivlin (1984). Economic Choices 1984.
  3. Congressional Budget Office (2018). An Overview of CBOLT: The Congressional Budget Office Long-Term Model.
  4. Joint Committee on Taxation (2024). History of the Joint Committee on Taxation.
  5. Joint Committee on Taxation (2024). Revenue Estimating.
  6. Sadowsky (1991). Computing Technology and Microsimulation.
  7. Sadowsky (2005). An Interview with George Sadowsky.
  8. Feenberg (1993). An Introduction to the TAXSIM Model.
  9. Tax Policy Center (2024). The Urban-Brookings Tax Policy Center Microsimulation Model.
  10. Brown (2012). On the Distributional Effects of Base-Broadening Income Tax Reform.
  11. Penn Wharton Budget Model (2017). The Tax Cuts and Jobs Act, as Reported by Conference Committee (12/15/17): Static and Dynamic Effects on the Budget and the Economy.
  12. Auerbach (2005). Dynamic Scoring: An Introduction to the Issues.
  13. Congressional Research Service (2015). Dynamic Scoring for the Congressional Budget Process: Pros and Cons.
  14. Joint Committee on Taxation (2017). Macroeconomic Analysis of the Conference Agreement for {H.R..
  15. 119th United States Congress (2025). H.R. 1, One Big Beautiful Bill Act.
  16. Elmendorf (2024). Dynamic Scoring: A Progress Report on Why, When, and How.
  17. (1996). Microsimulation and Public Policy.
  18. Sutherland (2013). EUROMOD: the European Union tax-benefit microsimulation model.
  19. Richiardi (2021). UKMOD -- A New Tax-Benefit Model for the Four Nations of the UK.
  20. Northumbria University (2024). Professor Howard Reed.
  21. Torry (2019). Static Microsimulation Research on Citizen's Basic Income for the UK: A Personal Summary and Further Reflections.
  22. Landais (2011). Pour une r\'{e.
  23. OpenFisca (2024). About OpenFisca.
  24. New Zealand Service Innovation Lab (2018). Better Rules for Government Discovery Report.
  25. Mohun (2020). Cracking the Code: Rulemaking for Humans and Machines.
  26. American Enterprise Institute (2015). Open Source Policy Center Launches Tax Brain.
  27. Holmer (2024). Martin Holmer Profile.
  28. Board of Governors of the Federal Reserve System (2024). FRB/US Model.

Society in silico · draft in public · source