23/07/2026
Synthetic Respondents: How AI Is Mimicking Customers, Experts, and Employees
Synthetic respondents promise instant, cheap survey data, but for decisions that move capital, AI-generated answers carry a hidden integrity risk. Here's where they work, where they fail, and how to verify your respondents are real.
A vendor can now hand you a thousand "customer interviews" by Friday without a single human ever answering a question. The responses read well. They are internally consistent, fast, and roughly a tenth of the cost of the real thing. For a marketing team stress-testing ad copy, that trade can make sense. For an investment director signing off a deal thesis, it is a data-integrity problem wearing a convincing disguise.
Synthetic respondents have moved from research-lab curiosity to something a strategy team can buy this quarter. The question is no longer whether the technology works. It is where it holds up, where it quietly breaks, and how you tell the difference before the output reaches a decision that moves capital.
What are synthetic respondents?
Synthetic respondents are AI-generated personas that answer research questions as if they were real people. A large language model is prompted to adopt a profile, a 45-year-old procurement lead in Germany, a mid-market CFO, a lapsed subscriber, and then responds to surveys or interview questions in character. No human is in the loop. The output looks like primary data. It is not.
The category splits into terms worth keeping straight. Synthetic consumers simulate buyers for concept testing, pricing work, and segmentation. Synthetic users, the version Nielsen Norman Group has evaluated most closely, simulate people moving through a product or interface.
Some vendors package these as "digital twins" of a customer or persona, which is the same idea dressed in more confident language. Synthetic experts and synthetic employees, the harder claim, simulate specialists or internal staff whose judgement would normally cost real time and money to access. The further up that ladder you go, the thinner the ground gets.
Why the sudden interest, and who is actually using them
The pull is obvious: speed, cost, and scale. What normally takes weeks of recruiting and fieldwork can be generated in hours, at a fraction of the price, with no scheduling and no incentives to pay. For teams under deadline pressure, that promise lands hard.
Adoption is running ahead of comfort, though, and the practitioners closest to the work are the most cautious. In the Rival Group 2026 Market Research Trends Report, 42.75% of market researchers said they were "not excited" about using synthetic respondents, even as they welcomed other AI applications. That reluctance is not technophobia. It is the reflex of people who have watched clean-looking data ruin a recommendation and would rather not repeat the experience. The interest is real. So is the hesitation, and the hesitation is coming from the experts.
Where synthetic respondents genuinely work
For structured, low-stakes tasks, synthetic respondents can earn their place. Concept testing, early pricing reads, questionnaire piloting, and desk research to shape a hypothesis are all reasonable jobs for an AI panel: fast, cheap, and low-consequence. Even NielsenIQ, which has worked on the technology for years, frames the honest use case narrowly, calling synthetic respondents "a supplement to your ideation process when time is of the essence" rather than a replacement for real people. Nobody is deploying capital on that first pass, and the cost of being wrong is a revised draft.
The strongest evidence for the category sits exactly here. A 2025 study, LLMs Reproduce Human Purchase Intent, reported up to 90% alignment with human survey data across 57 surveys using a semantic similarity method. Read that ceiling carefully. It applies to structured, quantifiable purchase-intent tasks, and it comes from work associated with vendors who sell synthetic panels. Ninety percent on a clean, structured question is the best case. It is not the number you get when the question touches lived experience, emotion, or a niche market. Used for piloting rather than deciding, synthetic respondents are a legitimate tool. The trouble starts when teams promote them from rehearsal to evidence.
Where they fail, and why it matters for a decision, not a campaign
Synthetic respondents fail hardest on exactly the questions that justify commissioning primary research in the first place: the novel, the specific, and the uncomfortable. The failures are not random noise you can average out. They are systematic, and they push in predictable directions.
The empirical record is blunt, and it gets worse the closer you look at B2B. Emporia Research ran real professionals against synthetic ones on the same B2B questionnaire and found the AI systematically cheerful: asked about satisfaction in their current role, 47% of real respondents said they were "somewhat satisfied" against 69% of the synthetic ones. That gap is not rounding error. The model is manufacturing contentment that the real market does not feel. Cambridge University Press, in Political Analysis, found that 48% of coefficients from ChatGPT-generated survey responses differed significantly from human ones, with the sign flipping 32% of the time. A third of the time the model did not just get the magnitude wrong. It pointed the opposite way. The Nürnberg Institut für Marktentscheidungen (NIM), founded by GfK, ran 500 real respondents against 500 AI ones and found the AI diverged from humans on 75% of soft-drink questions and 80% of sportswear questions, while two separate human samples stayed consistent with each other.
Three limitations sit underneath those numbers. First, the model is trained on the past and blind to the genuinely new. Ask it about a product category or a market shift that post-dates its training, and it will invent a plausible answer rather than admit it does not know. Second, it has credibility without lived experience. Merrill Research put a concept to synthetic design engineers and got the textbook answer that sustainability matters; real engineers said sustainability matters "but not when we can't get the parts we need for months on end". The supply-chain reality that would actually shape the buying decision was invisible to the AI. Third, the model wants to please you. Nielsen Norman Group found synthetic users rated every concept favourably, claimed they would finish online courses they realistically would not, and praised discussion forums that real users found contrived. That sycophancy is a well-documented property of large language models, not a tuning quirk you can prompt your way around, and it means the tool is most agreeable precisely when you need it to push back.
The bias runs one direction, too. NIM's model favoured well-known brands and underestimated demand for lesser-known ones. If you are validating an incumbent, the flattery is comfortable. If you are underwriting a challenger or an early-stage category, the synthetic panel will quietly talk you out of the thing the deal is built on.
The obvious rebuttal is to run several models and average them, on the theory that different training data cancels out any single model's blind spots. It does not work. A 2026 controlled comparison of 117 real interviews against synthetic ones ran the same personas through three frontier models and found the variation was in style, not substance: the models diverged on sentence rhythm and metaphor while converging on identical themes, identical adoption postures, and identical modal answers. Averaging three models gives you three ways of phrasing the same consensus. It does not recover the disagreement, the refusal, or the outlier that made the research worth commissioning.
Why synthetic data breaks fastest outside English
The weakness compounds the moment you cross a language border, which is most of the world's deal flow. NIM traced part of GPT-4's divergence from real respondents to a model trained overwhelmingly on English and US data. For a US fund buying a European or Asian asset, that is not a footnote. It sits at the core of the exposure.
A synthetic panel asked about a German industrial buyer or a Japanese distributor is not reasoning from local reality. It is translating US-centric assumptions into a foreign context, then presenting the result with the same confidence it shows everywhere else. The further a market sits from the model's training data, the wider the gap, and the harder it is to catch, because the output still reads fluently.
This is why native-language conversation is not a nicety. Research conducted in the respondent's own language, by someone who can hear when an answer is rehearsed, is often the only way to surface the local truth a synthetic panel structurally cannot reach. For cross-border diligence, global coverage stops being a logistics convenience and becomes a condition of the data being real at all.
Synthetic experts and employees: where the simulation is thinnest
Simulating a customer is the easiest of the three roles in this article's title. Simulating an expert or an employee is the hardest, and it happens to be exactly where buy-side research earns its fee. Consumer behaviour is heavily represented in the data a model trained on. Specialist judgement and internal reality are not.
An expert's value is the judgement that is not written down anywhere a model could have read it: the regulatory workaround everyone in the sector uses but nobody publishes, the reason a category that looks like it is growing is actually stalling, the supplier whose quality slipped two quarters ago and has not recovered. That is the Merrill finding scaled up. The synthetic design engineers gave the answer any textbook would give. The real ones named the parts shortage that actually governs the purchase. A synthetic expert regresses to the published consensus, and in commercial due diligence the published consensus is precisely the thing you are paying to get behind. If the answer were public, you would not have commissioned the interview.
Synthetic employees fail for a second reason stacked on the first. When you run internal research, on culture, on why deals stall in the pipeline, on where an organisation is quietly breaking, the signal you need is dissent, discomfort, and the tacit knowledge people carry but rarely volunteer. A model tuned to be agreeable manufactures consensus exactly where the disagreement was the finding. You end up with a workforce that reads as aligned because the tool has no way of being anything else. For a strategy team assessing an acquisition's integration risk, a synthetic workforce is worse than no data. It is a blindfold with good production values.
The risk you did not choose
Here is the part that gets missed. The debate so far assumes you decide whether to use synthetic respondents. Increasingly, you do not. The more pressing exposure is synthetic or AI-assisted responses arriving inside a panel you commissioned in good faith, without disclosure.
This is where synthetic respondents stop being a methodology choice and become a fraud-and-provenance problem, the same problem that sits under panel fraud and fabricated survey data. NielsenIQ puts the stakes plainly: "producing convincing answers is different from providing accurate ones, especially when it comes to making business decisions that rely on data integrity." A synthetic response is engineered to be convincing. Whether it is accurate is a separate question, and by the time it reaches a deal model, nobody is asking it. A fieldwork vendor under margin pressure has every incentive to let AI fill quota. The responses pass surface checks. They arrive on time. And they land in your dataset next to the real ones, indistinguishable to anyone who is not looking hard.
Consider how this lands on a compressed deal timeline. A team commissions a quick read on customer retention in a niche category, clean responses come back inside a week, and the numbers support the thesis. If even a portion of those responses were AI-generated to hit quota, the retention signal is not weak, it is fictional, and it is now load-bearing in an investment committee pack. The cost of that is not a bad survey. It is a mispriced deal. We wrote about the wider version of this in When Surveys Lie, and synthetic contamination is its sharpest new edge. If you cannot prove where a response came from, you cannot defend the conclusion you built on it, and in due diligence an indefensible number is worse than a missing one.
How to know your respondents are real
The answer to synthetic contamination is provenance built in from the start, not a better detector bolted on after the fact: human-verified primary research, where every response traces back to a real, identifiable person who actually holds the experience they are speaking to. Verification cannot be inspected in at the end. It has to be the design of the process.
When you commission research, ask the questions a synthetic panel cannot survive:
- Who are the respondents, specifically? Can their identity and their relevance to your question be evidenced, not just asserted?
- Were they sourced for this project, or drawn from a standing pool? A pool optimised for throughput is the easiest place for synthetic responses to hide.
- Were they spoken to in their own language, by someone who can tell a real answer from a rehearsed one? Native-language conversation matters more than it looks. Much of the published weakness in synthetic data traces back to models trained overwhelmingly on English and US data, and that gap widens fast outside those markets.
- Can the vendor show its work? Recruitment logic, screening criteria, and a chain of custody from question to answer are things a real process can produce and a synthetic one cannot.
A verified-human process treats the respondent as the control group against which AI-generated data is measured, which is the logic we set out in human-verified primary research as the control group for AI. You can see the same standard applied under a compressed diligence timeline in this commercial due diligence engagement. The point is not to slow the work down. The point is that when someone in the investment committee asks where a number came from, you have an answer that survives the question.
When to use synthetic respondents, and when to insist on humans
Use synthetic respondents to rehearse, not to decide. They are a reasonable fit for questionnaire testing, early hypothesis shaping, and structured tasks where the cost of being wrong is a redraft and nobody is committing capital. The moment a finding is going to sit in an investment committee pack, a client deliverable, or a market-entry business case, the standard changes, and the burden shifts back to verified humans.
The rule of thumb is simple enough to apply on a Friday afternoon. If getting it wrong costs you a revision, synthetic data is a tool. If getting it wrong costs you the deal, insist on people you can name.

Synthetic respondents are a genuine addition to the toolkit. They are also a genuine risk to any decision that cannot afford a plausible-sounding guess. Knowing which situation you are in, and being able to prove which kind of data you are holding, is now part of the job.
If a finding is heading for an investment committee, a client deliverable, or a market-entry case, it is worth knowing where every response actually came from. Bell & Holmes runs a short data-provenance review: we look at how a panel was sourced, screened, and verified, and flag where synthetic or unattributable responses may have entered the set before the number becomes load-bearing.
If you have a live study or a vendor you want pressure-tested, start a scoping conversation. No pitch, just a straight read on whether the data behind your decision would survive the question.