Study 1
Simulating a population
Draft availableWhen personas are built from survey data, do their answers reproduce the population distribution, and can changes to that population be audited?
On 74 ANES demographic cells, calibrated personas had about one-third the vote-distribution error of naive repeated prompting and preserved within-cell variation. Direct readout, verbalized sampling, a calibrated label, and a training-data lookup matched or beat the persona arm on these coarse static cells. Steering was tested for ideology and religious attendance; it worked on DeepSeek and only partly on GPT-4o-mini.
Why it matters
Simpler methods are stronger baselines for static shares in well-surveyed groups. Persona value must be measured on individual variation and downstream interaction, with calibration and steering audits published alongside the result.
Read the study →Can simulated discussion reproduce which people change their answers in real groups?
In a 50-group Wason evaluation, the original engine scored worse than no conversation. A private step that reasoned from each participant's seeded misconception reduced composite error by 45% on DeepSeek. Three-seed transfer tests found the opposite effect on Qwen 9B and an inconclusive result on Llama 3.1 8B.
Why it matters
A conversation mechanism can improve one model and harm another. Revalidate it whenever the model or engine changes.
Read the study →Does a larger model produce more human-like deliberation?
On the Wason task, solo drift toward the textbook answer rose across nine Qwen sizes, although the series was not monotonic. The 0.8B model's wrong-to-correct rate was 0.044, compared with 0.573 for DeepSeek. Under generated deliberation, error was statistically tied from 4B upward. Both member fidelity and change dynamics followed the thinking model; the speaking model had a smaller effect.
Why it matters
Model choice trades member-level reading against task-specific drift. Select the thinking model from measured requirements. The single-run split suggests a cheaper speaking model may preserve most measured behavior, but this needs interval-bearing replication.
Read the study →Study 4
Evidence in the room
Draft availableWhen sources enter a discussion, do simulated voters change at the human rate and for supported reasons?
Without evidence arrival, generated agents changed at 0.09 times the real rate. On 48 committed voters in content-bearing debates, an external belief state plus a provenance gate matched the gross change magnitude, 0.3125; the instability-adjusted result was marginal. Across six seeds per evidence-bearing cell, evidence helped at the level of passing the preregistered band, while rankings by evidence volume and model size remained unresolved. In a 39-turn audit, 19 of 34 source-claiming turns, 56%, were judged unsupported by their excerpt.
Why it matters
Evidence must reach the room, and source claims need a content check. This study does not rank evidence volume or model size.
Read the study →Study 5
Who sees the votes
Draft availableDoes seeing the room's votes change what a simulated group decides?
On one model and task, an anonymous running count increased within-group agreement by 17.8 percentage points. Adding names produced no detected additional effect. One specified false tally produced no detected change in the share ending on its target relative to the true-tally arm. On within-group agreement, it produced no detected difference from the private control and was 16.6 points below the true tally. Simulated groups reached about half the convergence observed in the real groups.
Why it matters
Vote visibility is a causal configuration choice. The experiment detected an informational channel and left identity-based social pressure unresolved for this task.
Read the study →Study 6
Does it read like a person?
Draft availableCan visible AI-writing tells be removed, and does removing them improve behavioral fidelity?
A speech checklist removed 69% of the measured writing tells, with no detected behavioral change. An instruction aimed at hedging in private thought barely changed the visible tells and improved composite fidelity by 0.282. In a 34,000-message census, tell rates varied sharply by model.
Why it matters
Transcript style and behavioral fidelity require separate measurements. Cleaner prose does not establish a more faithful simulated decision.
Read the study →Study 7
Documented individuals
Draft availableDo measured facts about a person improve prediction of that person's answer?
Across three models, one predictive political attribute raised balanced accuracy from 0.60-0.64 to 0.86-0.88. Traits sampled from the person's demographics scored 0.49-0.52. Role-playing the person and asking about them directly were statistically equivalent, agreeing on 92-96% of individuals.
Why it matters
Measured provenance carries useful signal. A structured profile provides an audit trail; the persona framing added no measured accuracy in this experiment.
Read the study →Study 8
Audiences no survey measured
Draft availableCan deeper demographic descriptions reach a population that existing surveys missed?
Across specifications containing one to five coarse demographic attributes, direct prompting remained the most accurate tested arm, and the training-data lookup beat the fusion generator at every depth. Even the deepest survey cells still had training support. The support gap appeared when a variable was missing. That result comes from masking an observed variable; a genuinely never-collected variable remains untested.
Why it matters
A defensible unsupported-audience claim must identify the missing variable, show that it carries signal, and report the survey's own uncertainty floor.
Read the study →