A synthetic respondent is a research subject played by a language model — a written identity, situation and manner of speaking that a model is asked to hold to while a moderator asks it questions. The output is transcript-shaped: answers, objections, hesitations, the occasional refusal to engage with a premise.
The interesting question is not whether that output reads as plausible. It does, immediately and effortlessly, and that is the problem rather than the proof. The question is which decisions it can carry. The honest answer is narrower than the product demos suggest and wider than the flat dismissals allow.
Two different things wear the same name
The first family is statistical. A model is conditioned on real survey respondents’ demographic backstories and asked to answer the next question the way that person would. Argyle and colleagues called the result a silicon sample and reported that a properly conditioned model reproduced response distributions from human subgroups closely enough to be worth studying, a property they named algorithmic fidelity (Political Analysis, 2023). The ambition there is measurement: get a distribution out that resembles the one a survey would have produced.
The second family is qualitative. A persona is an authored document — a specific person with a specific situation, written to be read aloud by a model rather than summarised in a slide — and the output is a conversation, not a number. Nobody claims the transcript is representative. The claim is that the objections in it are objections a real person could plausibly raise, and that hearing them early is cheaper than hearing them late.
Conflating the two is where most of the argument about synthetic respondents comes from. A qualitative persona library is not making a measurement claim, so measurement critiques miss it; a simulated survey is making one, so it inherits every burden of proof a survey has.
What the published evidence actually says
Three findings are worth carrying into any use of this method.
Conditioning works better than it has any right to, within limits. The silicon-sample result above is real and has been replicated in adjacent settings. A model conditioned on a described person produces language that correlates with what people like that say.
The extremes get exaggerated. Bisbee and colleagues prompted a model to adopt personas and rate sociopolitical groups; the averages tracked the human baseline reasonably well, but the generated responses overstated the extremity and the certainty of division compared with real people holding the same attributes (Political Analysis, 2024). Synthetic respondents are more confident and more polarised than their originals. In a focus group that shows up as an objection stated too crisply, which is flattering to a research team and misleading about how hard the objection actually is.
Steering does not fix whose views a model defaults to. Santurkar and colleagues measured the alignment between model opinions and those of sixty US demographic groups and found substantial mismatch that persisted even when models were explicitly steered toward a group (arXiv, 2023). A persona document is a steering instruction. It shifts the output; it does not relocate the model’s centre of gravity.
Set against that, the comparison case is not perfect either. Opt-in online panels carry measurable shares of bogus respondents — roughly 4% to 7% depending on the source, answering in patterns that bias estimates rather than merely adding noise (Pew Research Center, 2020) — and the resulting distortions concentrate in exactly the young and Hispanic subgroups researchers most often want to read (Pew Research Center, 2024). “Ask real people” is the right instinct and it is not automatically the same thing as “get real answers”.
What synthetic respondents are good for
Five uses survive contact with the evidence above, and they share a shape: the synthetic session decides what to do next, and something else decides what is true.
Pressure-testing a discussion guide. Run the guide before it costs a recruitment budget. Questions that produce nothing produce nothing quickly, and the ones that produce only agreement are usually badly framed.
Surfacing objections to positioning. A persona written with documented distrusts will name the objection your positioning invites. Whether that objection is widespread is a different study; that it exists at all is worth knowing before the launch.
Rehearsing a moderator. Interviewing a difficult persona is practice for interviewing a difficult person, and the transcript is reviewable afterwards without a consent form.
Briefing a team on an audience it has not met. A specific written subject travels better through an organisation than a demographic bracket, because people argue with a person and nod at a bracket.
Triage before fielding. Which of six concepts deserves the expensive test? A synthetic pass will not rank them reliably, but it will often eliminate one for a reason nobody had noticed.
What they cannot do
They cannot report what a market did. They cannot establish incidence, size a segment, or supply a percentage anyone should defend. They cannot tell you a fact about the world their author did not know. They cannot be cited as evidence of demand, and a synthetic quotation in a pitch deck, presented as a customer voice, is a fabrication regardless of how carefully the persona was written.
They also cannot surprise you the way a person can. A real respondent occasionally answers from a life you did not model. A persona answers from the document, and the document was written by someone with your assumptions in it.
The failure modes to design around
Flattening. Ask ten personas the same question and the answers converge more than ten people’s would. Design for disagreement: pick personas that genuinely conflict on the dimension under test, and treat unanimity as a signal that the cohort was badly chosen.
Agreeableness. Models drift toward the answer the moderator seems to want. Moderator turns that carry an implied preference get it back, every time.
Premise echo. A persona handed a positioning statement will often reuse its vocabulary. If the transcript is full of your words, you have measured your own copy.
Confidence inflation. Per the finding above, discount the certainty rather than the content. “I would never pay for that” from a synthetic respondent is closer to “I would hesitate” from a person.
Replayability. If the persona text can be edited in place, last month’s finding cannot be reproduced. Version the documents and freeze them.
How Klingbar’s library is built
The persona library is 185 authored personas — 95 consumer and 90 business — published as documents rather than generated on demand. Each one carries a classification layer so a cohort can be assembled deliberately: consumer personas carry an Esri Tapestry segment, business personas carry a 2022 NAICS code, and every persona scores the same 19 marketing channels for propensity. The mechanics of that layer are covered in the guide to a persona taxonomy.
Published personas are immutable and versioned. An edit lands as a new version under a new row rather than a rewrite, which is what makes a finished study replayable against the exact text its participants were given. Klingbar’s own product is elsewhere — it is the organizational layer for bots, built around an org chart written as code — and the focus-group capability described on its marketing pages is a planned, organization-scoped specialty rather than a browser product available today. The library itself is public, free to read, and shipping now.
Questions worth asking a vendor
ESOMAR’s 20 Questions to Help Buyers of AI-Based Services is the standard checklist, and four of its themes matter most here. Where did the persona text come from, and was any real person’s data used to build it? Which model played the respondent, and can that be pinned for a repeat run? Can a finished study be reproduced against the exact inputs it used? And what does the vendor claim the output measures — because a vendor that will not answer that one in a sentence is answering it in the pricing page instead.
Where this leaves you
Use synthetic respondents as a rehearsal instrument. They are cheap, immediate, and specific enough to change a discussion guide, a concept, or a launch sequence. Then go and ask people, and report which of the two you did.