SurdaticsSURDATICS
All posts
Industry commentary

Pew’s synthetic respondents did not stand in for survey participants

Surdatics Research

AI-assisted, reviewed by the Surdatics team. Drafted with AI from the sources listed at the end, then checked and approved by a person.

The appeal of a synthetic respondent is easy to see: give a model information about a participant, and perhaps it can answer later survey questions on that person’s behalf. Pew Research Center tested that proposition by supplying panelists’ demographic information and earlier answers, then asking an AI model to answer questions from 3 American Trends Panel surveys as those panelists. Pew says its main synthetic results used Claude Opus 4.6. This was a test of substitution, not a test of whether bots could enter a live survey.

The substitution did not hold up well. Pew found an average difference of 12 percentage points between synthetic and human estimates across nearly 300 individual questions; the average difference exceeded 15 points on around 28% of questions. Those gaps matter because a survey team does not usually need a plausible-sounding set of answers. It needs estimates that can support decisions about what people actually report. A model can produce fluent responses while still changing the finding the research was meant to establish.

The errors are not just an average

Aggregate error can obscure where a synthetic survey is least dependable. Pew reported average absolute errors across the 3 waves of 16.1 percentage points for Republicans and Republican leaners and 15.1 for Black adults. Pew also found that 25% of human respondents said they had heard a lot about data centers, compared with 3% of synthetic respondents. A team using synthetic answers to plan outreach or interpret differences between groups would need to know whether the particular estimate it relies on is sound, not merely whether the overall output looks coherent.

Prior answers may help a model form a convincing portrait of a participant. They do not guarantee that the portrait captures what the participant knows, has encountered or would say to a new question. The practical lesson is to validate the specific measures and groups relevant to a study before considering any use of simulated answers. Similarity on familiar questions should not be treated as permission to assume accuracy on unfamiliar ones.

Missing answers can disappear from view

The problem is not limited to estimates that drift. Pew Research Center found that 47% of questions in its synthetic poll had at least one answer choice selected by no synthetic respondent; none of the questions in the 3 human survey waves did. If a response option disappears from simulated results, a researcher might mistake a model’s reluctance to choose it for an absence of that view among people. That is especially consequential when the purpose of a survey is to find experiences a team did not anticipate.

The output also depends on the model chosen. Pew found that GPT-5.1 and Claude Opus 4.6 produced different portrayals of public opinion on a subset of the same questions. If changing the model changes the apparent public, a synthetic estimate needs to be treated as an output of a modelling setup, not as an independent observation of public opinion.

Keep the useful distinction

Rejecting synthetic respondents as replacements does not mean rejecting every experimental use. NORC at the University of Chicago identifies questionnaire testing and methodological experimentation as potential uses, while calling for rigorous evaluation. A research team could explore whether simulated answers expose confusing wording or suggest questions to test. It should then check those ideas with people rather than treating the simulation as the completed study.

Survey substitution is also distinct from AI-assisted participation in surveys that seek human respondents. In separate research, Wang, Mamaev and Leckie found that detectors readily identified naïve AI-generated responses but performed near chance against agents designed to mimic whole respondents. That finding concerns the difficulty of detecting certain assisted entries; it does not make synthetic stand-ins more accurate. Teams need separate plans for validating any simulated research tool and for protecting the integrity of human-response data.

For monitoring-and-evaluation and AI data teams alike, the boundary should be explicit in the research record: which outputs came from people, which came from a model, and what evidence supports using either for the stated purpose. Synthetic answers can prompt better questions. Pew’s comparison is a warning against letting them silently become the answers.

Sources

  1. 1Silicon Samples and Synthetic Surveys: Can AI Stand In for Human Respondents? | Pew Research Centerpewresearch.org
  2. 2AI surveys fail to capture the diversity of public opinion | Pew Research Centerpewresearch.org
  3. 3Human Data Remain Essential in the Age of Synthetic Respondents | NORC at the University of Chicagonorc.org
  4. 4Towards Detecting AI-Assisted Responses in Online Surveysarxiv.org