Research

What AnthroSim has been tested against.

The program tests selected claims against ANES 2020, 500 recorded DeliData discussions, and 369,215 Wikipedia Articles-for-Deletion debates.

The current program reports 8 studies in 5 accessible drafts. Review scores come from the automated Stanford Agentic Reviewer at paperreview.ai. Human peer review is pending for every draft.

Eight measured questions

Findings so far

Study 1

Simulating a population

Draft available

When personas are built from survey data, do their answers reproduce the population distribution, and can changes to that population be audited?

On 74 ANES demographic cells, calibrated personas had about one-third the vote-distribution error of naive repeated prompting and preserved within-cell variation. Direct readout, verbalized sampling, a calibrated label, and a training-data lookup matched or beat the persona arm on these coarse static cells. Steering was tested for ideology and religious attendance; it worked on DeepSeek and only partly on GPT-4o-mini.

Why it matters

Simpler methods are stronger baselines for static shares in well-surveyed groups. Persona value must be measured on individual variation and downstream interaction, with calibration and steering audits published alongside the result.

Read the study →
Study 2

Group discussion

Draft available

Can simulated discussion reproduce which people change their answers in real groups?

In a 50-group Wason evaluation, the original engine scored worse than no conversation. A private step that reasoned from each participant's seeded misconception reduced composite error by 45% on DeepSeek. Three-seed transfer tests found the opposite effect on Qwen 9B and an inconclusive result on Llama 3.1 8B.

Why it matters

A conversation mechanism can improve one model and harm another. Revalidate it whenever the model or engine changes.

Read the study →
Study 3

Model size

Draft available

Does a larger model produce more human-like deliberation?

On the Wason task, solo drift toward the textbook answer rose across nine Qwen sizes, although the series was not monotonic. The 0.8B model's wrong-to-correct rate was 0.044, compared with 0.573 for DeepSeek. Under generated deliberation, error was statistically tied from 4B upward. Both member fidelity and change dynamics followed the thinking model; the speaking model had a smaller effect.

Why it matters

Model choice trades member-level reading against task-specific drift. Select the thinking model from measured requirements. The single-run split suggests a cheaper speaking model may preserve most measured behavior, but this needs interval-bearing replication.

Read the study →
Study 4

Evidence in the room

Draft available

When sources enter a discussion, do simulated voters change at the human rate and for supported reasons?

Without evidence arrival, generated agents changed at 0.09 times the real rate. On 48 committed voters in content-bearing debates, an external belief state plus a provenance gate matched the gross change magnitude, 0.3125; the instability-adjusted result was marginal. Across six seeds per evidence-bearing cell, evidence helped at the level of passing the preregistered band, while rankings by evidence volume and model size remained unresolved. In a 39-turn audit, 19 of 34 source-claiming turns, 56%, were judged unsupported by their excerpt.

Why it matters

Evidence must reach the room, and source claims need a content check. This study does not rank evidence volume or model size.

Read the study →
Study 5

Who sees the votes

Draft available

Does seeing the room's votes change what a simulated group decides?

On one model and task, an anonymous running count increased within-group agreement by 17.8 percentage points. Adding names produced no detected additional effect. One specified false tally produced no detected change in the share ending on its target relative to the true-tally arm. On within-group agreement, it produced no detected difference from the private control and was 16.6 points below the true tally. Simulated groups reached about half the convergence observed in the real groups.

Why it matters

Vote visibility is a causal configuration choice. The experiment detected an informational channel and left identity-based social pressure unresolved for this task.

Read the study →
Study 6

Does it read like a person?

Draft available

Can visible AI-writing tells be removed, and does removing them improve behavioral fidelity?

A speech checklist removed 69% of the measured writing tells, with no detected behavioral change. An instruction aimed at hedging in private thought barely changed the visible tells and improved composite fidelity by 0.282. In a 34,000-message census, tell rates varied sharply by model.

Why it matters

Transcript style and behavioral fidelity require separate measurements. Cleaner prose does not establish a more faithful simulated decision.

Read the study →
Study 7

Documented individuals

Draft available

Do measured facts about a person improve prediction of that person's answer?

Across three models, one predictive political attribute raised balanced accuracy from 0.60-0.64 to 0.86-0.88. Traits sampled from the person's demographics scored 0.49-0.52. Role-playing the person and asking about them directly were statistically equivalent, agreeing on 92-96% of individuals.

Why it matters

Measured provenance carries useful signal. A structured profile provides an audit trail; the persona framing added no measured accuracy in this experiment.

Read the study →
Study 8

Audiences no survey measured

Draft available

Can deeper demographic descriptions reach a population that existing surveys missed?

Across specifications containing one to five coarse demographic attributes, direct prompting remained the most accurate tested arm, and the training-data lookup beat the fusion generator at every depth. Even the deepest survey cells still had training support. The support gap appeared when a variable was missing. That result comes from masking an observed variable; a genuinely never-collected variable remains untested.

Why it matters

A defensible unsupported-audience claim must identify the missing variable, show that it carries signal, and report the survey's own uncertainty floor.

Read the study →
Where the workflow fits

A different early question for each team

Each route uses simulation to decide what deserves a real test. The result remains exploratory.

Research

Problem

Recruitment makes early changes to a discussion guide or experimental condition slow and expensive.

Evidence-bounded use

Pilot the instrument with synthetic cohorts, inspect failure cases, and take the revised design into participant recruitment.

Enterprise and product

Problem

Product documents and recorded support failures do not automatically become varied test cases for a live system.

Evidence-bounded use

Let configured synthetic users discuss the supplied material, then turn recurring questions into tests against the real product.

Market research

Problem

The first fieldwork round often exposes confusing concepts or questions after recruitment is already complete.

Evidence-bounded use

Compare early concepts across configured synthetic segments and revise the discussion guide before fieldwork.

Political and policy

Problem

Policy language can reach a public meeting before the team has heard a broad set of objections.

Evidence-bounded use

Collect simulated objections and unclear terms, then test the revised message with constituents.

Legal

Problem

A real mock jury is costly to spend on discovering that an instruction or case fact is hard to follow.

Evidence-bounded use

Find confusing evidence and instructions with a synthetic panel before a real mock jury.

Sensitive and regulated

Problem

Early research design can stall when recruitment or participant records carry privacy and review requirements.

Evidence-bounded use

Explore the questions without recruiting people or uploading identifiable participant records, then plan the approved human study.

Source material

Five draft papers

Each page carries the method, measurements, limitations, and current automated review record.

IE-ASIM-2026-01 · Draft available

From Silicon Sampling to Population Steering

Automated review: 6/7, 7/7, 5/7 · latest assessment: justify publication

IE-ASIM-2026-02/03 · Draft available

The Inner Loop and the Ladder

Automated review: 4/7, 4/7, 5/7, 7/7 · latest assessment: recommend acceptance

IE-ASIM-2026-04 · Draft available

Evidence in the Room

Automated review: 5/7, 4/7, 4/7, 5/7, 5/7 · latest assessment: recommend acceptance, or a strong borderline in favor

IE-ASIM-2026-05/06 · Draft available

Two Kinds of Visibility in Simulated Deliberation

Automated review: 6/7, 5/7, 6/7 · latest assessment: recommend acceptance

IE-ASIM-2026-07/08 · Draft available

Two Boundaries for Simulated People

Automated review: 6/7, 7/7, 7/7 · latest assessment: recommend acceptance