Chapter 16

Simulating values

3,680 words · 19 min · draft, July 2026


Every chapter before this one obeyed the same rule. A simulation earned attention only where its verification chain ended in ground truth: an encoding checked against the statute and a reference calculator, a synthetic population calibrated to administrative totals, a forecast that an official number will eventually grade. This chapter asks the deepest question in the book's set — not what people want today, which surveys can measure, but what people would want after reflection, and whether that is something a model could forecast. The admission rule cannot follow the question that far. No office certifies reflective values; no future print resolves them. So read what follows as horizon, not foundation.

Here is what reaching past the rule looks like in practice. In 2024 I ran the first systematic test I know of on whether a language model can forecast value change, and the headline result was a miss. The miss is the center of this chapter — not the method, not the modest success that preceded it, and not the uses a mature version might someday have. The failure carries the lesson, and the two years since have mostly deepened it.

Why forecast a value

In 1996, Gallup began asking Americans whether same-sex marriages should be legally valid; 27 percent said yes. By the early 2020s the figure was about seven in ten .[1] The change was not random. It followed exposure, generational replacement, and the slow crystallization of arguments, and in retrospect the trajectory looks almost legible — the kind of curve you feel you could have drawn in advance. Could you have, in 1996? And can you now tell which of today's contested values will look settled in a generation, and which will reverse?

This is not a search for moral truth. It is a forecasting question, the same shape as forecasting inflation or rainfall: a survey, a time series, a model, a calibrated interval, and a grade when the next reading lands. The apparatus of chapter 13, with an attitude on the vertical axis instead of a dollar. It matters because decisions get made against contested values constantly, and increasingly by automated systems. "Do what humans want" fractures the moment you ask which humans, and whether you mean what they say now or what they would endorse on reflection. Stuart Russell built his account of controllable AI on exactly that uncertainty: machines should stay unsure what humans want and learn it from behavior .[2] But behavior reveals present preferences — precisely the ones most likely to move. If value change has structure, some of it might be forecastable, and the only way to find out is to forecast it and take the grade.

The backtest

The General Social Survey has asked Americans the same attitude questions since 1972, with wording stable enough to form real time series. I took seventeen of them — spanning sexual morality, drug policy, gender roles, race, religion, confidence in science, guns, capital punishment, and free speech — and gave GPT-4o each variable's history through 2021, asking not for a point guess but for a spread: the 10th, 25th, 50th, 75th, and 90th percentiles of where the 2024 reading would land. The spread was there to force the model to express uncertainty rather than default to false confidence.

Two deliberately simple baselines set the bar. A linear trend fits a straight line to the history and extends it; a no-change baseline predicts the historical average. Any method worth having should beat both, and beating the straight line is the harder, more meaningful test. The model's 2024 forecasts came in at a mean absolute error of about 4.8 percentage points, against about 7.2 for the linear trend and about 10.6 for no-change — roughly 1.5 times better than the straight line, roughly 2.2 times better than assuming nothing ever changes. Naming the baseline is most of the honesty here: "2.2 times better" is true only against the weakest opponent. The claim worth keeping is that the model beat a straight line by about half again, and that its advantage concentrated where trends bent: acceptance accelerating on some questions, change decelerating on others, exactly where a straight line goes wrong. Its raw intervals came out too narrow, the familiar overconfidence of language models, and had to be widened with a standard calibration step borrowed from weather forecasting before the 80 percent intervals covered 80 percent of the outcomes.

Read carefully, that is a modest, real result: on three-year forecasts of American social attitudes, a language model beat naive baselines, by more against the weaker one and less against the sterner one. Then came the variable I would have bet the house on.

The miss

Among the seventeen series, same-sex acceptance had the cleanest story. The GSS asks whether sexual relations between two adults of the same sex are wrong; the share answering "not wrong at all" had climbed for half a century, driven by generational replacement and exposure, with every hallmark of a one-way trend. It read roughly 57 percent in 2018, 62 in 2021, and 61 in 2022 .[3][4] to verify It was exactly the series a forecaster feels safest about. In 2024 it fell to roughly 55, about six points below its 2022 reading.

Every method predicted up; the number went down.

The language model missed the reversal. The linear trend missed it. The no-change baseline missed the level and the direction both. Six points is not large by the standards of survey movement, and the reversal may yet prove to be noise, an artifact, or the leading edge of a genuine backlash; that question is still open. The lesson is the direction. On the one series where the trend looked most like destiny, everything pointed one way and reality went the other. Value change carries genuine surprises that no model, statistical or neural, reliably anticipates.

The break in the ruler

The GSS changed how it collects answers, and the change sits directly under this chapter's centerpiece. Facing the pandemic, the 2021 round moved from in-person interviewing to an address-based, push-to-web design, and its response rate fell to roughly 17 percent, against roughly 50 percent in 2022 .[4] Mode matters for exactly this kind of question: answering a screen instead of a person removes social-desirability pressure, and self-administration tends to raise the share giving the franker answer — including "not wrong at all." NORC, which runs the survey, cautions users to "carefully examine how a change they are analyzing relates to this methodological shift."

So part of the climb the model trusted — the 2021–22 peak — may be a mode artifact, and part of the 2024 reversal may be that artifact unwinding. Both the rise and the fall sit partly on a methodological break, not on pure attitude change. The caveat does not rescue the forecasts: none of the methods knew about the mode change either, and reality graded them on the numbers as published. What it does is make the miss doubly instructive. The model failed to anticipate the reversal, and the series it extrapolated was itself less solid than it looked. A forecaster's ground truth is a measurement, measurements are made by instruments, and instruments have seams. Anyone proposing to forecast humanity's values should first notice how much trouble we have measuring last year's.

The replication

The miss sent me back to rebuild the experiment properly, and the rebuild cut deeper than the miss. In July 2026 I pre-registered a replication before making a single model call .[5] Twenty GSS items this time, survey-weighted; each model shown the series only through 2022 and asked for 2024; every training cutoff checked against the data's release date, so that no arm could be "forecasting" a number it had already read. And the baselines got sterner. The one that matters for a slow-moving series is persistence — the forecast that 2024 simply equals 2022.

Persistence beat everything.

Across the twenty items, the language models' errors ran 3.9 to 4.1 points against persistence's 3.15. A frontier reasoning model did no better than the cheap ones, and raising a model's reasoning effort made its accuracy slightly worse. The models' 90 percent intervals covered the truth about half the time; properly computed classical intervals covered at their stated rate. And the 2024 wave turned out to be stranger than one famous reversal: eight of the twenty items broke their pre-2022 trend, and four of them — approval of the school-prayer ban, same-sex relations, rejection of traditional gender roles, and marijuana legalization — posted the largest single-wave declines in their items' recorded histories. Every arm, model and baseline alike, missed the same-sex reversal again, with point forecasts between 60 and 66 against an actual of roughly 55; the only interval that contained the truth managed it by being twenty-one points wide.

Squaring the two runs is deflating and clarifying at once: the baselines are most of the result. The first experiment set the bar at a trend line and a decades-long average, and against an average taken over decades of much lower readings, any method that merely notices the recent level looks brilliant. Persistence is the bar that matters, and nothing cleared it.

Recall in a forecast's clothes

The rebuilt experiment included a control the earlier ones lacked, and it explains where much of the apparent forecasting skill in this literature comes from. Run a model on a series it can name, inside its training window, and it is not only extrapolating — it is remembering. Shown the same-sex series through 2010 with its identity visible, one model "forecast" the present at 78, a level the series has never reached in fifty years of asking. Shown the same numbers with the name stripped, the same model said 47.5. Thirty points, from the label alone.

An identified backtest inside a model's training window is a recall test wearing a forecast's clothes. The diagnosis reaches well past this chapter, because a growing literature evaluates language models by "predicting" surveys, elections, and experiments the models have already read about. Two protocols fix it: strip the identities, or forecast strictly forward, where no training corpus can help. The rebuilt experiment adopts both — its forecasts for the 2026 and 2028 GSS waves were registered before the data exist.

A stranger fix is coming into view. Language models trained exclusively on text from before 1931 now exist citation pending; such a model has read no poll it could parrot, because Gallup's first national surveys came in 1935, the GSS began in 1972, and the entire quantitative record of public opinion lies beyond its horizon. Early evaluations put their weakness at the elicitation layer rather than the knowledge layer — they can complete text they cannot converse about — which looks like an engineering problem rather than a wall, and stronger checkpoints at the same cutoff are expected within months. The eventual prize is ninety years of graded actuals: every reversal, plateau, and backlash in the polling record, out-of-sample by construction.

The same apparatus will also happily overreach, which is its own warning. Point it at 2050, or 2100, and it dutifully returns a tidy median and interval for each year. I once tabulated exactly those long-horizon numbers. After 2024, I cut the table. The only honest way to read such figures is as an illustration of what the method emits when asked, never as a finding about the future — and the further past the data you push, the more the interval, not the median, is the result worth reporting.

Heterogeneity is the target, not the noise

A tempting mistake runs through this whole subject: imagining that value forecasting converges on a single answer — the value system humanity is "heading toward." It does not, and it should not. Grant the method perfect foresight and people would still hold different things dear; the object worth forecasting is the whole distribution across a population. The reframing has teeth for anything that would consume the forecast. A point estimate invites a system to optimize toward it. A distribution refuses: it says that at the end of any plausible reflection there remain people who weigh liberty over welfare and people who weigh welfare over liberty, and that both are part of the answer rather than error bars around it. A system that took the distribution seriously would favor actions that score acceptably across several value systems over actions that score maximally within one, grow cautious where the value systems disagree most, and treat a move that catastrophically violates any large minority as disqualified rather than merely outvoted. That is value pluralism recast as an engineering constraint — closer to a robustness requirement than a moral theory.

The temptation to print the distribution anyway has to be resisted. An earlier draft of this chapter contained a table of invented population shares — tidy percentages for how much of a reflective humanity would land on which cluster of values. It was exactly the confident fiction this book stands against, dressed as a finding, and it is gone. The distribution is the thing to estimate, not to assume.

Two uncertainties

Any forecast of values carries two kinds of uncertainty that do not reduce to each other. Aleatoric uncertainty is the genuine spread of values across people: a perfect model of a population that will never agree still reports disagreement. Epistemic uncertainty is our uncertainty about what that spread even is: a distribution over distributions, which better data, longer histories, and better methods narrow. They demand different responses. Epistemic uncertainty is the part you can pay to reduce — more surveys, more countries, sterner calibration. Aleatoric uncertainty is the part you must design around, because it remains when the study is done. Both belong in the output, quantified separately: not "humanity will value X," but a probability that it does, an interval around that probability, and an account of how much of the interval is ignorance and how much is real variety. This is the same discipline chapter 13 applied to policy costs, turned on values — and it is harder here, because the ground truth itself is a survey with seams.

Reflection is not time

The deepest problem is not statistical. The appeal of forecasting values is the hope that where attitudes are heading is better than where they are — that the forecast points forward in judgment, not just in time. Nothing in the data licenses that hope. That support for same-sex marriage rose from 27 percent to about seven in ten within a generation shows that time passed and attitudes moved; it does not show that anyone reasoned more carefully. The engines of value change are sociological — generational replacement, media exposure, information cascades, tribal signaling. A model that predicts them predicts a social process, not a deliberative one. Some shifts plainly track moral progress: abolition, the widening of the franchise. Others may be drift with no progress in them. Others might be backlash to perceived overreach — one live reading of the 2024 reversal, and if that reading is right, a model that "correctly" projected continued acceptance would have been wrong in a morally loaded direction.

The philosophers who pressed hardest on this offer warnings more than solutions. Rawls distinguished reasonable from unreasonable comprehensive doctrines — not every value system deserves equal standing in political life — but "reasonable" resists being turned into a forecasting rule; you cannot regress on it. Habermas specified the conditions of undistorted deliberation, equal participation and the absence of coercion among them, and no historical stretch of attitude change is guaranteed to have met those conditions; most plainly did not. Sen and Nussbaum warned about adaptive preferences: values formed under oppression, stable and sincerely held, that it would be perverse to extrapolate forward as though they were considered choices. A person raised to expect little may sincerely want little. A society that has normalized an injustice may show stable, well-measured preferences for it. A forecasting model does not know the difference between a preference reasoned into and a preference ground into place — both are points in a time series — and predicting the second kind accurately would be an achievement in the service of nothing good.

We have no clean account of when temporal change tracks reflection and when it is mere drift. Predicting well is evidence of underlying structure; it is not evidence that the structure deserves honoring. That is a permanent reason for humility, not a bug a larger model removes.

Nor is there one trajectory to defer to. Everything above was measured on one country's attitudes. The World Values Survey, run across roughly a hundred countries since 1981 citation pending, shows convergence and its opposite: on some dimensions, a broad drift toward self-expression values as societies grow wealthier and more secure — Ronald Inglehart's post-materialist pattern — and on others, religion and family structure and the norms around sexuality among them, persistent cultural zones in which different starting points produce genuinely different paths rather than the same path at different speeds. Recent years have added pointed reversals in several countries. A global forecast would have to model heterogeneity at the level of whole societies, and there is no neutral vantage from which to call one society "ahead" and another "behind."

The governance questions

Suppose the method someday worked: calibrated, replicated, graded forward for years. It would immediately become a technology of power, and who held it would matter as much as how accurate it was. A government could dress paternalism as foresight. A firm could steer preferences rather than serve them. A lab could align a system to projected values that happened to coincide with its own interests. And a forecast would not need a steering hand to move its own target: publishing that a shift is coming could accelerate the shift or suppress it. That feedback runs through all social science and is no reason to stop; it is a reason the adversarial questions would need answers before any of it earned authority.

Who generates the forecasts, and under whose incentives? A university lab answers to peer review, a statistical agency to its legislature, a company to its shareholders — different accountabilities, and none should hold a monopoly. Distributed production under open methods can at least be checked.

Which conditioning assumptions are on the table? A forecast conditioned on broad deliberation differs from one conditioned on polarized media, and the choice between them is itself a value judgment that must be stated where it can be argued with.

What is democratic input for? Forecasts should inform deliberation and never stand in for it; citizens must be able to examine the method, challenge the assumptions, and reject a forecast-based justification outright. How do disputes get resolved? Contested forecasts need appeals — a rival team rerunning the method, a public comment window, an independent review — not an oracle's say-so.

And the hardest one: what happens when the forecast and today's expressed values disagree? The honest answer stays uncertain, and the 2024 reversal is the standing cautionary tale, because a 2020 projection of continued acceptance, used to justify anything at all, would have been overconfident within four years. Whatever governs this work has to be built to absorb surprise and revise — built like the scoreboard, not like a prophet.

Interest without belief

If value forecasting ever earned admission on this book's own terms — calibrated across variables, countries, and horizons; leakage-controlled; adversarially tested; graded forward rather than backtested — it would bear on the largest open question in AI: how systems that act on our behalf should weigh contested human values. Even then its role would be bounded: one input, evidence about where considered preferences might sit, never an authority over them. Today it has earned no such standing. What exists is a research question with one suggestive backtest, one humbling miss, one pre-registered replication in which nothing beat persistence, and one methodological caveat sitting under the whole series .[5] citation pending If the work matures, its institutional shape is already described in this book: value forecasting is at bottom a forecasting problem — a probability distribution over what people will come to value, scored against what actually happens — exactly the kind of claim chapter 13's docket exists to post with an interval and grade. The forward-registered forecasts for the 2026 and 2028 GSS waves would be its first entries; none has resolved. There is no track record here, only a method and a discipline for eventually being wrong in public.

One can imagine chaining this book's tools into a single machine — a forecast value distribution feeding a democratic simulation feeding a policy model feeding measured welfare. I have drawn that diagram. It is an unbuilt composition of mostly toy parts, and drawing a diagram is not building a machine.

What can honestly be said is smaller and sturdier than the grand version. Ask people what they value. Model the patterns. Project them forward with calibrated uncertainty, and stay loud about how uncertain the projection is. We cannot know what humanity would choose after reflection, and we should distrust anyone who claims they can. We might, someday, forecast it — badly at first, and a little better with each miss admitted in public. Everything else in this book earned its place by terminating in a check against the world; this chapter did not, and I have tried to keep the difference sharp, because it is the one place where I am reasoning past the edge of what can yet be verified. The correct posture toward it is interest without belief.

And nothing in the choice ahead waits on it. The fork the final chapter returns to — open or closed, graded or oracular, plural or singular — is being decided now, with the tools already built.

Society in silico · draft in public · source