Synthetic Respondents Need a Brief, Not Blind Trust
Ipsos’s UK publication on 10 September, *Concept Testing with Digital Twins*, makes a more disciplined case for synthetic respondents than much of the surrounding conversation. The proposition is not that brands can dispense with consumers. It is that properly built digital twins can help sift a large number of early product ideas, allowing human research to concentrate on the concepts worth taking seriously.
That distinction matters. Research teams are under familiar pressure: more innovation options, shorter planning cycles and little appetite for adding weeks to a launch timetable. If an AI model can rule out weak concepts before a conventional study begins, it could improve the allocation of both research budget and managerial attention.
But speed has a way of changing the question. A tool introduced to prioritise hypotheses can soon be treated as a substitute for measurement. Synthetic respondents may look like a sample, return structured answers and produce persuasive verbatims. None of that makes them evidence of what people currently think, feel or will do.
The commercial task is therefore not to decide whether digital twins are “good” or “bad”. It is to give them a precise job description — one that reflects the data they draw upon, the decision at stake and the consequences of being wrong.
A synthetic audience is most valuable when it narrows uncertainty without pretending to remove it.
Screening is a stronger use case than prediction
Ipsos’s work is notable because it is relatively specific about scope. Its food-sector twins were trained on structured and unstructured responses from real people, then assessed against an established early-stage trial-potential measure. The stated aim is to identify materially different concepts that deserve further attention, rather than to adjudicate between marginal copy variations.
This is a sensible boundary. At the front end of innovation, companies often face a surplus of ideas and a shortage of time. A synthetic model may be able to expose obvious mismatches between an idea and the known preferences, language and habits of a defined audience. It can create a first pass across a much wider option set than a team would normally put in front of respondents.
That is not the same as forecasting market success. An early product concept is necessarily thin: a proposition, an image, a price cue, perhaps a pack mock-up. Its usefulness lies in prompting questions. Is the benefit clear? Does the idea appear distinctive? Which assumptions should researchers investigate with actual buyers? Used in that way, a digital twin acts more like an informed filter than a virtual focus group.
The difference is commercially important. Filtering weak options can save money and reduce organisational noise. Declaring a winner, especially where investment, reputation or a major launch is concerned, requires more. It requires an encounter with a market that can surprise the business.
Human research does not merely return a score. It reveals confusion the brief did not anticipate, trade-offs respondents make when their own money is involved, and cultural signals that are only just emerging. Those are precisely the phenomena least likely to be faithfully represented by a model trained on yesterday’s evidence.
Similar answers are not independent answers
The core methodological risk is easy to miss because generated responses can appear richly individual. A thousand synthetic consumers may use different words, take different tones and produce tidy segment splits. Yet they are not a thousand independently observed people. They are outputs from a system shaped by common training data, modelling decisions and prompts.
That affects what variation means. A real sample contains disagreement, inconsistency, changing circumstances and occasionally inconvenient views. Some of that is error; some is the signal. A generated population may instead smooth its way towards what the model regards as the plausible answer. The result can be an attractively coherent story that understates uncertainty and overstates consensus.
A recent peer-reviewed study comparing LLM-generated respondents with 461 real US soft-drink buyers illustrates the point. The synthetic sample reproduced broad patterns in brand familiarity and selection, but it also rated established brands too positively and displayed less variation than the people it was intended to represent. It was directionally informative without being interchangeable with the underlying market.

This is a useful warning for brand teams. Familiar brands are often where models have the most available cultural material to draw on, which can create an illusion of confidence. But a launch decision depends on more than whether a concept resembles things people have seen before. It depends on whether it changes behaviour amid competing products, price pressure, retailer constraints and the social context of a real purchase.
Plausibility is not validation. Nor is a synthetic output made more robust simply by increasing the number of generated cases. Scaling a model’s answers can make a weak assumption look statistically impressive.
Validation must be tied to the decision
The right test for a digital-twin system is not whether it can produce human-sounding responses. It is whether it performs reliably on a clearly specified task, against data it has not already seen, at a level of accuracy that is useful for the decision being made.
That demands uncomfortable but necessary questions. What source data formed the model? Was it collected with consent for this use? Which categories, markets and audience groups does it cover? When was the behavioural evidence last refreshed? What outcome is it validated against: an attitudinal score, product trial, repeat purchase, creative comprehension or something else?
Most importantly, where does it fail? A supplier should be able to show error by segment and by type of question, not merely an overall result. A model that works well for established food categories may be much less dependable for a new service proposition, a sensitive issue or an unfamiliar market. A tool that predicts broad appeal may be poor at identifying the minority response that later becomes commercially decisive.
The Market Research Society has argued that synthetic data should enhance research, fill gaps and stress-test assumptions rather than replace direct interaction with people. Its guidance also highlights the absence of widely accepted quality benchmarks. That should lead buyers to insist on a validation plan before a platform is embedded in innovation or communications workflows.
For high-consequence decisions, the practical answer is a two-stage design. Use synthetic respondents to generate, organise and challenge hypotheses. Then place the strongest concepts into research with recruited people, using methods proportionate to the risk. The latter stage is not a ceremonial sign-off. It is where the business tests whether the model has missed the thing that matters.
This approach also changes what good reporting looks like. Instead of presenting a synthetic output as a definitive consumer finding, teams should distinguish between simulated indications and observed evidence. Senior stakeholders need to see the confidence range, assumptions and human validation alongside the recommendation.
Research leaders should protect the discovery function
The attraction of digital twins is understandable. They promise insight at the pace of internal debate, not fieldwork. Yet research’s strategic value has never been only speed. It is its capacity to put organisations in contact with evidence that disrupts their existing view of the customer.
A synthetic model is built from prior knowledge. It can combine, retrieve and simulate that knowledge at remarkable speed. It cannot reliably discover the consumer shift that has not yet entered the data, or the new language through which people are beginning to describe a category. This is particularly relevant when markets are unsettled by changes in price, regulation, culture or technology.
There is also a governance issue. If brand, product and agency teams begin using generated audiences informally, insights functions can lose sight of which decisions rest on real observation and which rest on simulation. The result is not simply a technical risk. It is a credibility risk for the organisation’s evidence base.
Research leaders should establish a small set of operating rules: classify synthetic work clearly; retain an audit trail of inputs and prompts; define approved use cases; require periodic testing against fresh human data; and prevent synthetic outputs from being reported as survey findings. These are not bureaucratic obstacles. They are the conditions that make fast experimentation defensible.
Digital twins may indeed earn an enduring role in concept development. But that role will be strongest where the industry resists its own temptation to oversell. The prize is not a cheaper imitation of customer research. It is a sharper research system: one that uses simulation to explore more possibilities, and real people to decide which possibilities deserve belief.



Comments