"For the first time, history has an author."
The book opened on that line, and it has been waiting here ever since. Engerraund Serac says it with pride, standing in a control room while Rehoboam's predictions cascade across the displays — every life a trajectory, every trajectory arranged, a society optimized to one man's definition of optimal .[1] He does not claim to see the future so much as to write it; in his machine, prediction and control have fused into a single act. The fiction earns its horror honestly. What it gets right is that the modeling of society was never the dangerous part — the danger was an author: one model, one owner, one objective, no one outside the room who could say it was wrong. Seventeen chapters later, the answer this book proposes fits in a sentence.
History should have auditors, not an author.
Both visions begin from the same recognition — that complex social systems can be modeled computationally, policies simulated before enactment, populations represented household by household, consequences estimated before anyone has to live them. They part on everything after. Serac's model is his alone; its predictions are hidden from the people living inside them; its objective is singular and imposed; it speaks in certainties; those without access are subjects. The open path inverts each property in turn: code anyone can read, predictions anyone can query against their own household, objectives contested among many stakeholders because no single optimizer gets to define "optimal," uncertainty printed as intervals rather than suppressed, and the capacity to analyze held as public infrastructure rather than private advantage. Both paths run on models. They differ in who can see the models, question them, and correct them — which is to say, they differ in everything.
What we built
A year ago, this chapter's job was to describe what somebody would need to build. It reads differently now for a reason the preface confessed: a chapter drafted in 2025 as a description of missing infrastructure had to be rewritten as a field report, the future tense overtaken by the past tense between one draft and the next. A book about the pace of change that was itself outrun by it is an awkward artifact, but an honest one — and the being-outrun is the most direct evidence for the claim that matters most here. The window in which this infrastructure gets built open or closed is not a generation away; it is now. So here is the ledger, scoped and dated as of July 2026: what exists, and then what does not.
PolicyEngine, the open engine of Part II, encodes thousands of federal and state tax and benefit rules and shows them from two angles — the household view, what does this policy do to my family, and the society view, who gains, who loses, and what does it cost. Its population data is calibrated to more than 9,000 administrative totals to verify, its calculations are cross-checked against independent implementations through validation partnerships with the National Bureau of Economic Research and the Federal Reserve Bank of Atlanta to verify, and partner benefit screeners — MyFriendBen among them — run families across the same engine to verify. HM Treasury has piloted it beside the closed model that produces the government's own estimates .[2] It is still running and still free to anyone with a browser.
The rules commons is the agent turn made institutional. The Axiom Foundation launches on July 28, 2026 ,[3] carrying the Part III work: roughly 3,000 US rule files spanning federal law and twenty-eight states, with companion repositories for the United Kingdom, Canada, New Zealand, and Belgium to verify. Its stated mission is to encode all public policy — statute, regulation, agency manual, down to a single London borough's council-tax reduction scheme — with every monetary amount traced to a quoted excerpt of its governing source. In one week this July, its pipeline encoded five African tax-benefit systems and checked them against the SOUTHMOD reference models that UNU-WIDER and its national partners have maintained for a decade, reporting the findings back to the models' authors .[4] On every surface compared against a reference model so far, every mismatch carries an evidence-backed explanation — zero unexplained on the tested surfaces, which is a narrower and more defensible sentence than "correct," and the right one. That week is the strongest evidence this book has for its thesis, and it happened while the manuscript sat open.
Beneath both views sits the data commons: populace, which assembles the calibrated population — synthetic households, tuned to administrative totals, that let a model represent anyone in the country rather than only the respondents a survey happened to reach. It is the successor to the enhanced survey files of chapter 6, and the answer to the asymmetry chapter 2 diagnosed: the government's confidential records were never coming out from behind Section 6103, so the open stack synthesizes its own people and proves them against the totals everyone can see.
The scoreboard is the youngest piece. The Thesis Institute publishes forecasts of government metrics — the first print of an unemployment number; a Medicaid call-center wait time under a policy that may or may not take effect — each carrying an explicit interval, each to be graded when the official figure lands .[5] It launches publicly in mid-September 2026. Its status belongs in the record exactly as it stands: the docket is live, the grades arrive with reality, and not one forecast has resolved.
And one project exists to measure the boundary the others defend. PolicyBench puts today's best language models on a public board and scores them on complete household tax-and-benefit calculations .[6] As of mid-2026, the best model still gets roughly one in nine of those calculations wrong by more than a dollar, and the typical model misses about one in four — family-level errors, a wrong Medicaid eligibility or a wrong SNAP amount, not rounding slips. That measured gap is why the deterministic layer has to exist at all, and the board exists so the claim stays checkable as the models improve.
The five entries are one structure pulled apart along its natural seams — three of the introduction's four questions, given institutional homes with different verification loops; the fourth stays where the introduction left it, with the public. PolicyEngine keeps its name and its users; underneath, it is being recomposed into Axiom, which holds what the law says, and populace, which holds who the people are, with Thesis as the third pole for what will happen: the decomposition chapter 12 argued for, with repositories and launch dates where the argument used to be.
What the split buys is a kind of accountability the closed stack never carried. Rules are checked against statute and reference calculators the day they are written; populations are checked against administrative totals; forecasts are checked against reality when the number prints. An estimate that names something reality will print — a caseload, a revenue line, a poverty rate with a release date — need no longer be a single number issued and forgotten; it can carry an interval and, on the branch the world takes, a grade. The CBO and its peers publish real retrospectives of their aggregate projections; what none of them runs is the per-estimate loop, commitments logged in advance and graded one by one, misses kept on the page. Building that loop is the reason to split the stack, and the loop has not yet closed on a single forecast.
What we haven't
The debit side of the ledger is long, and it belongs in the same list as the credits. Start with uncertainty, which is still mostly unquantified where ordinary users meet the tools. A microsimulation answers "this reform costs fifty billion dollars" when it should answer "fifty billion, give or take this much, and here is the assumption doing the work." Chapter 13 showed the banded version running on one reform; wiring it through production, so that every score arrives as a distribution instead of a point dressed as precision, is unfinished work.
Adoption is early. When Congress passed the One Big Beautiful Bill Act in July 2025 — Public Law 119-21, extending the 2017 tax cuts — the public argument ran on CBO scores and think-tank estimates .[7] Scores mattered enormously: they shaped a live, enacted law worth trillions of dollars. But they were institutional scores, produced by the same closed machinery Part I described, and citizens running their own analyses played no visible part. The tools existed; the debate proceeded as if they did not. The shift from available to expected is a cultural transition, and weather forecasting — the book's favorite precedent — needed decades to turn "chance of rain" into something ordinary people plan around. Policy analysis has barely started that clock.
The prediction pole has no track record. The Thesis Institute has not resolved a single forecast. Its scoreboard is a set of promises with intervals attached; until reality grades them, they are predictions, not evidence. Saying this out loud, repeatedly, is the entire reason to build a scoreboard instead of a megaphone.
And the AI layer is half-arrived. The rules and the data now exist at production scale, and an AI assistant can call them today; that part is real and shipped. What has not happened is the habit. Ask most assistants a tax-policy question and the answer still comes from training data rather than from a calculation engine — the infrastructure is built, and the reflex of reaching for it is not.
Add it up and this book describes as much aspiration as achievement: the rules layer is further up the mountain than the last draft dared claim, and the prediction pole is barely off the trailhead.
We are partway up, and honest about the altitude.
Why the fork matters
If a society cannot reason about itself, it cannot govern itself. And the alternative to open simulation was never human judgment uncorrupted by models — it is the set of failure modes already in place. Agencies make black-box decisions with proprietary tools; the Joint Committee on Taxation produces the revenue estimates Congress votes on with models Congress cannot inspect. Debate runs on vibes: politicians assert, pundits assert louder, and numbers float past without sources or audit trails. AI systems answer policy questions from training data, confidently and often wrongly — the family-level errors PolicyBench measures, shipped at the scale of every chatbot. And analytic power pools where the compute and the data already are: a hedge fund can model the distributional impact of a tax reform in minutes while a community organization cannot, so the fund's version of the story arrives first and arrives armed. None of it is hypothetical — this is how policy is analyzed today.
Open simulation answers each failure in kind — auditable code where there were black boxes, reproducible results where there were vibes, tools an AI can call where it would otherwise invent, public infrastructure where there was private advantage. None of the substitutions is automatic, and each can be done badly. An open model can still encode a bad assumption; transparency does not make a mistake correct — it makes the mistake findable, which is the property everything else in this book compounds on.
The claim has edges worth drawing. Some government models earn their opacity: publish exactly how the fraud detector weighs its signals and you have published the evasion manual, and parts of policing and tax enforcement sit under the same constraint. For those, accountability runs through the output rather than the source — error rates measured and published, appeals that work, auditors with access — a real discipline, and a different one. And the argument here concerns government; insurers and credit scorers raise a cousin of the question under different law, a fight for other books. What survives the carve-outs is the narrow claim the book actually makes: where the model is the policy — what the law says, who qualifies, what a reform does to whom — the rules are law, and law the governed cannot read is not functioning as law. Openness is owed exactly where the computation decides.
The deepest version of the case is democratic. Democracy asks people to vote on policies without knowing their effects, to debate taxes without calculating impacts, to judge officials without any way to audit their claims — and for most of democratic history that ignorance was not chosen, it was structural: the models sat inside institutions, and expert analysis was expensive and late. What the open stack changes is concrete enough to picture. A candidate releases a tax plan; within hours, independent analyses appear — not from one think tank but from many people running the same open model under different named assumptions. A voter enters her own household and sees the specific dollar effect. A journalist checks the distributional claim against the calculation instead of against a rival quote. The disagreement that survives that process is the real one, about values, argued from a shared factual floor. Chapter 15 sized this claim honestly — information is one lever among identity, institutions, and strategy — but it is the lever infrastructure can move, and now the option to know exists where it did not.
The machines will do the asking
As AI systems become the interface to everything else, more people will meet policy through them than through any government website or newspaper table. Someone asks what a child-credit expansion would do for them, or how a carbon tax lands on a family like theirs, or which candidate's health plan leaves them better off. The system answering can do one of two things. It can synthesize plausible text from training data — inventing eligibility rules, hallucinating statistics — which is what happens today, and why the best frontier models still complete fewer than a third of full federal returns correctly [8] and still miss household calculations at the rate PolicyBench records. Or it can call a tool: look the parameters up in an open encoding, run the calculation deterministically, and return an answer with an audit trail.
The value of the second path compounds. A microsimulation model that takes ten minutes to learn serves one researcher at a time; the same model, reachable by an AI agent, serves anyone who can ask a question in plain language. Every hour spent encoding a rule exactly, tracing it to its source, and racing it against an oracle pays off in proportion to how many agents eventually call it — and that number is climbing. Reliability here is a mechanism, not a promise. When PolicyBench's own explanatory notes were caught inventing the derivations behind scores they described, the fix was not better prose discipline; it was mechanical grounding — regenerate every explanation from the engine's actual internals and reject any that cannot be traced. Even the audit layer needs ground truth.
Nor is the architecture a private bet. The same shape — an AI front end over a deterministic, checkable back end — keeps appearing wherever institutions try to let people ask real questions: the IRS has put AI agents to work summarizing cases and searching guidance ,[9] and California's Poppy assistant helps state employees navigate dense policy catalogs .[10] None of those systems is PolicyEngine, and that is the point. The pattern converges because it is the shape that works. The alternative, at the scale of every assistant on Earth, is more access to less reliable information — which no one should mistake for progress.
Brittle and improvable
In the fiction, Serac's machine fails. Rehoboam cannot handle the hosts — beings who do not fit its model, whose choices it cannot predict — and the closed system shatters on contact with genuine novelty. Television resolves these things too tidily, but the lesson underneath is sound: closed systems optimize for their own assumptions, and when the world moves, they break in ways their designers never anticipated. The government modeling monopoly of the 1980s was a mild, real version of the same disease. When the Joint Committee on Taxation produced an estimate, you could disagree with it, but you could not demonstrate the error — models that cannot be challenged cannot be corrected, and so they weren't. The open-source movement changed this not by producing better models, but by making models improvable. And in the 2026 stack, openness is machinery rather than gesture: the same encodings anyone can read sit behind gates where tested coverage may only rise and unexplained mismatches may only fall, so the worry that open code quietly rots has an operational answer instead of a hopeful one.
Run the through-line once, end to end. Guy Orcutt proposed simulating a nation household by household in 1957 .[11] Government agencies realized the idea behind closed doors, and for decades the machinery answered only to its owners. Think tanks broke the monopoly; open-source projects put the tools in browsers; and agents have now made them buildable in days and checkable line by line against the models governments already trust. Every step in that sequence widened the circle of people who could take part in the argument about what policy does, and to whom. The last step widened something else: who can build the argument's instruments, and how fast an error anywhere in them can be found.
An invitation
The fork, restated one last time, is not science fiction. Both paths are live institutional choices being made right now — in code, in procurement contracts, in the defaults of AI assistants, in decisions that look technical and are not. One path rebuilds the analytic machinery of government behind glass: closed, singular, speaking in certainties. The other rebuilds it in daylight: inspectable, plural, graded when reality arrives. The difference was never sophistication — Serac's machine had plenty — but whether anyone outside the room can check it, correct it, and build on it.
So the invitation is quiet, because the work does not need believers; it needs checkers. The code is public. The encodings can be read, the models run against your own questions, the forecast docket watched while it earns or fails to earn a record — the admission rule chapter 3 stated and every chapter since has enforced, now in your hands. Find something wrong and file it; the findings ledgers already run in both directions — divergences dispositioned against our encodings, bugs filed upstream against the reference models themselves. Or reject the premise entirely: argue that these tools entrench technocracy rather than open analysis, that what is measurable will crowd out what matters. I think that argument is wrong, and I would rather lose it in the open, against specifics anyone can examine, than win it by default behind glass.
What the tools actually do is narrow, and it is enough. They cannot tell a society what to value, cannot guarantee that insight gets used wisely, cannot stop a bad-faith actor from trying to game them. What they can do is make the invisible visible: the benefit cliff buried in a formula, the uncertainty a confident headline hides, the interaction between programs that no single-program analysis reveals, the distributional table folded into a bill's fine print. A voter who can see that a plan raises her family's income by twelve hundred dollars may still vote against it, for reasons of value or identity or loyalty — that is her right, and nothing in this book has an opinion about it. The difference is that she chooses with open eyes.
Guy Orcutt died in 2006, forty-nine years after the paper almost nobody read. He lived to see microsimulation become institutional infrastructure, but not public infrastructure; the browsers, the open engines, and the agents that build and check the rules one household at a time all came after him. His insight was per-unit — model society one household at a time, and a picture emerges that no aggregate equation can hold. The agent era supplied the half he never got to see: per-unit verifiability, each household's answer checked against a statute, an oracle, or a real determination, the discipline the whole stack now rests on. The arc from his idea to here bends toward openness — not inevitably, not smoothly, but detectably: monopolies gave way to competition, proprietary models to open code, expert-only interfaces to anyone's browser, and now to machines that must answer for their arithmetic. The next turn of the arc is the one this book was written inside, as AI becomes the way millions of people meet what policy does to them. Whether that turn expands access or expands misinformation depends on a choice that is still open, and on who shows up to check the checkers.
Serac's machine had an author and no auditors, and the show was right about where that ends. The machinery can simulate the futures on offer; only the public it serves can decide which one is worth building. Society in silico is not a destination. It's a method.
References
- HBO (2020). Westworld, Season 3, Episode 2: ``The Winter Line''.
- HM Treasury (2024). HMT: Policy Engine UK - Algorithmic Transparency Recording Standard.
- Axiom Foundation (2026). Axiom Foundation -- The world's rules, encoded.
- UNU-WIDER (2026). SOUTHMOD: Simulating tax and benefit policies for development.
- Thesis Institute (2026). Thesis Institute forecast docket.
- PolicyEngine (2026). PolicyBench: Benchmarking language models on household tax and benefit calculations.
- 119th United States Congress (2025). H.R. 1, One Big Beautiful Bill Act.
- Bock (2025). TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task.
- Muoio (2025). IRS deploys AI agents through Salesforce.
- California Department of Technology (2026). Poppy: California's Digital Assistant.
- Orcutt (1957). A New Type of Socio-Economic System.