Skip to content

All resources

Explainer5 min

How to read a synthetic stakeholder panel responsibly

What a simulated panel can legitimately tell you, what it cannot, and the specific ways it goes wrong.

Full text

A synthetic stakeholder panel takes a scientific statement — a finding, an abstract, a draft summary — and generates how a range of constructed professional perspectives might read it. The output looks like qualitative research. It is arranged like interview transcripts, it has the texture of considered opinion, and that resemblance is the main hazard in using it.

What it can legitimately do

Surface ambiguity. If several constructed perspectives read the same sentence differently, that sentence is genuinely ambiguous. This is a property of the text, and it holds regardless of whether the simulated readers resemble real ones.

Generate hypotheses cheaply. Objections and questions worth preparing for can be enumerated in minutes rather than weeks. The panel is a checklist generator, and a checklist does not need to be representative to be useful.

Rehearse a difficult framing. Before a statement goes to people whose time is expensive, it is useful to see the obvious misreadings. Finding three of them in advance is worth something even if the fourth only emerges in a real conversation.

Expose a missing limitation. If no simulated reader mentions a caveat you consider central, that is a signal the caveat is not doing its work in the text — which is a fact about your writing, not about the simulation.

What it cannot do

It cannot tell you what proportion of clinicians would agree. There is no sampling frame, no response rate, and no population from which the respondents were drawn. Percentages computed from simulated output describe the generator, not any group of people.

It cannot substitute for human research where a decision depends on the answer. Advisory boards, payer research, and patient input exist because the answer matters enough to be obtained from the people themselves.

It cannot detect the objection nobody in the training data thought of. Simulation reproduces distributions of expressed views. The novel objection from the person with unusual experience is exactly the thing it will miss, and it is often the objection that matters most.

It cannot supply evidence for a claim. Simulated agreement is not evidence of clinician opinion, and presenting it as such — internally or externally — misrepresents both the method and the finding.

Specific failure modes to watch for

Fluency read as consensus. Generated text is uniformly articulate. Real experts hedge, contradict themselves, misunderstand the question, and change their minds mid-sentence. Uniform coherence across a panel is a property of the generator and should reduce your confidence, not raise it.

Persona collapse. Personas constructed from thin descriptions tend to converge on a similar voice. If the cardiologist and the health economist raise the same three concerns in the same order, the panel has less diversity than its labels suggest.

Prompt sensitivity. Small changes in how the statement is presented can move the output substantially. A result that does not survive rephrasing is a result about the prompt.

Agreeableness. Models tend toward accommodation. A panel that finds your framing broadly reasonable may be exhibiting a known model behaviour rather than telling you anything about your framing.

Laundering. The most damaging failure is organisational rather than technical: a simulated finding is quoted in a slide, the slide is quoted in a memo, and by the third document the qualifier is gone and the number reads as research. Every safeguard against this is procedural — labelling in the interface, labelling in every export, and a refusal to present simulated output in the visual language used for real data.

A working rule

Use a synthetic panel to decide what to ask, never to decide what is true. If a conclusion would change a decision, the panel has told you where to spend real research effort, not what that effort would have found. Read the disagreements rather than the totals, treat the transcript as a set of prompts for your own judgement, and carry the label with the finding every time it moves to a new document.

More guides