Picture the slide: purchase intent for a new snack, "n=2,000," turned around in four hours. The respondents are synthetic. The chart looks great. The problem is that nobody in the room can say what it's actually measuring.

That slide is showing up in more and more meetings.

I'm not against synthetic data. I use it. But it needs rules.

The short version

  • Synthetic respondents are model outputs, not people. Label them that way.
  • They're useful for early exploration, questionnaire testing and small, disclosed gap-filling.
  • They struggle with new products, emerging behaviors and anything the training data never saw.
  • Validate against real sample before anyone makes a decision on synthetic results.

What "synthetic respondents" actually means

The term covers a few different things. LLM personas answer questions "as" a 35-year-old mom in Ohio. Statistical augmentation generates extra records for a thin subgroup based on real survey data. And classic imputation fills in missing values. They carry very different risks, so the first governance step is naming which one you're using.

Where it genuinely helps

  • Stress-testing a questionnaire before field, to catch confusing wording or broken logic.
  • Brainstorming hypotheses so the real survey asks sharper questions.
  • Carefully boosting a small subgroup, with clear disclosure, when the base is already real and large.

Where it goes wrong

Models reproduce averages well and variance badly. Real people are inconsistent, surprising and sometimes contradictory, and that's exactly the signal research is supposed to catch. Synthetic answers also can't react to something genuinely new, because by definition there's no history to learn from.

There's a quieter risk too: synthetic data can make a thin sample look robust. A cell of 40 real people plus 160 modeled ones is still a cell of 40.

A governance checklist we use

  1. Disclose synthetic data in the methodology section, every time.
  2. Keep real, verified respondents as ground truth and benchmark against them.
  3. Report the real base size next to any augmented number.
  4. Never use synthetic data for final go/no-go decisions without a real-sample read.

The ESOMAR Code already separates a real individual from a synthetic persona, and I'd expect buyers to hold providers to that line.

Where real sample still wins

When the question is "will people actually buy this," you need people. That's why we keep investing in real respondents in 100+ countries rather than modeling them.

FAQ

Can synthetic data replace survey respondents?

Not for decisions. It can speed up early-stage work, but results should be validated against real, verified respondents before they inform strategy.

How do you validate synthetic respondents?

Run the same questions with a real sample and compare distributions, not just averages. Look at variance, subgroup differences and open-ended themes. If they diverge, trust the real data.