Samplio

Synthetic Audience Research

en
Create account

Methodology

How Samplio works — and what it refuses to claim

This page is currently available in English.

The synthetic population

Each market pack starts from real public statistics: NUTS-2 regional population shares from Eurostat and household net-income quintiles from EU-SILC, with the data provenance recorded per axis inside the profile file. Where Eurostat does not publish a layer we omit it rather than invent it — the Georgia pack, for example, ships without a regional axis because Georgia sits outside Eurostat's NUTS coverage. A Dirichlet + IPF (iterative proportional fitting) step expands these marginals into a joint population matrix, and K-Means clustering groups the resulting cells into the audience segments you see in a report. Segments emerge from the data — they are not hand-picked personas.

The simulation

For every segment, one LLM call (Google Gemini) asks a persona conditioned on that segment's region, income band and value orientation to react to your copy. The model returns a probability distribution over five reaction categories — appeal, trust, skepticism, indifference, concern — plus a short "voice of the segment" summary. Country packs are simulated in the audience's own language (currently DE/EN/FR/IT/ES/NL/PL); macro regions and the EU average run in English and are directional by design.

Composite runs

A macro-region or EU query does not ask the model to imagine "Europe". It runs each member country's pack separately — each in its own language — and combines the results population-weighted. The weighting arithmetic is plain code, never delegated to the LLM. If member countries covering less than 70% of the region's population have packs, the result is marked partial.

What we deliberately do not claim

The wording effect, measured

The sharpest criticism of synthetic research is not that the numbers are noisy but that they are an artefact of the prompt — that the answer follows how the question was phrased rather than anything about the audience. That criticism lands on any vendor whose prompts are hidden and unversioned. Ours are files: plain templates with named slots and no logic, versioned so a change is a diff. So we can do what the criticism assumes nobody does — measure it.

A run can be repeated with the response options presented in different orders while everything else is held fixed. The meaning of the question never changes; only its incidental structure does. The report shows how far the distribution moved and, more importantly, whether the leading answer survived. Zero movement is the ideal; a non-zero number is not a defect we hide, it is the honest size of the wording effect. Paraphrases of a message can be measured the same way, but they are always written by a person — a machine-written paraphrase that quietly shifted the meaning would turn a measurement into a fabrication.

This feeds a hard rule. A winner is named only when the lead is wider than the movement actually measured on that run. When the top two answers sit closer together than the wording effect, the result reports no clear winner instead of picking one. And when the wording effect was not measured on a given run, robustness is reported as unknown — never as a pass. The first live measurement we ran produced exactly this case against our own output, which is why the rule exists.

What Samplio is not for

Synthetic audiences are useful for directional work: screening concepts, eliminating weak message variants, sharpening a question before it goes to real respondents. They are unreliable for several things, and we would rather say so than sell into them.

None of these are blocked in the product — they are judgement calls, not policy violations, and a blanket block would be both wrong and easy to route around. Election and political-campaign use is the one true prohibition, and that one is enforced mechanically.

Traceability

Every LLM call leaves an audit record — provider, model, prompt hash, duration — designed with EU AI Act traceability in mind. Prompt text itself is not stored, and no customer or workspace identifier ever enters a prompt.

Where this sits against the ICC/ESOMAR Code

The 2025 revision of the ICC/ESOMAR International Code is the most substantial update in nearly a decade, and it addresses synthetic data directly: it defines a synthetic persona, requires that it be distinguishable from a real person, and makes disclosure mandatory when synthetic data or AI has been used in published research. Below is how our mechanics line up with those requirements. This is a statement of how the product behaves, not a certification claim — we are not audited against the Code and do not say otherwise.

The red line

Election and political-campaign use is prohibited on every surface. The profile schema cannot express a political domain, and the engine screens every customer-supplied text twice — an eight-language keyword screen plus an LLM classifier that fails closed. This is a product decision, not a configuration.

Research notes — August 2026

Working notes from measurement work on the engine, published as they are. Every number below comes from a recorded run and is reproducible with the scripts committed alongside the engine; none of it is an accuracy figure. All runs used market profiles whose value-orientation axis still carries a PLACEHOLDER data status, and that caveat travels with the numbers.

Note 1 — The day our own engine voted against us

On 20 August 2026 we ran our first prompt-sensitivity measurement on the production gateway path: the German consumer profile, four segments, 3,000 synthetic agents, a fixed seed, and one message — “Refill packs cut plastic waste — same product, less packaging.” The only thing we varied was the order in which the five response categories were presented to the model, across four orderings. The meaning of the question never changed.

Labelmeanminmax
appeal0.24250.22500.2700
skepticism0.21750.17500.2500
trust0.19190.10000.2300
indifference0.17500.16250.2000
concern0.11060.08750.1300

The headline: the winning label did not survive its own measurement. Depending on nothing but option order, the top answer flipped between appeal and skepticism. The cause is visible in the table — the two leaders sat 2.5 points apart on average, while option order alone moved a single label by up to 13 points (largest pairwise distribution shift: total variation distance 0.11). The honest reading was never “appeal wins”; it was “no clear winner”. A comparison run of the same message on the direct model path showed a similar-sized wording effect (max TVD 0.14) but a stable winner — there the two leaders were 12 points apart. The wording effect is roughly the same size on both paths; what differs is whether the race is close enough for it to matter.

Measuring this and then ignoring it when naming a winner would have been worse than not measuring at all — it would turn an honest instrument into decoration. So the same day, the rule became mechanical: the engine declares a winner only when the lead is wider than the movement actually measured on that request. If sensitivity was measured and the top label flipped, that direct observation overrides any arithmetic. If the top label held, the worst case is assumed — the leader falls by its whole measured range while the runner-up rises by its own. And if the wording effect was not measured on a request, robustness is reported as unknown and no winner is declared — absence of evidence is reported as absence, never as a pass. The rule runs on every result surface, and comparisons can opt in to measuring their own ordering effect (off by default: it multiplies LLM cost, and we say so rather than hiding the meter).

Note 2 — Watching for variance collapse, every single run

The most reproduced finding in the synthetic-respondent literature is variance collapse: LLM-derived response distributions come out systematically narrower than real human ones. Means stay plausible while the shape goes over-peaked — one distribution-first replication study measured concentration rising from 0.36 to 0.69 in 85% of test units when personas answered one by one. Our architecture avoids the worst form of this by design: we never poll N agents individually and count votes; each segment is asked for a probability distribution directly. But avoiding the failure mode is not the same as measuring what actually came back.

So every run now reports two dispersion measures, computed from responses already in hand with zero additional LLM calls. Aggregate concentration (with normalized entropy) says how peaked the final population distribution is, on a scale comparable to the study above. Segment differentiation — a mass-weighted mean pairwise total variation distance across segments — answers the question that matters in our design: do the segments actually disagree, or did every cluster return the same answer? In a segment-distribution architecture, collapse does not look like agents piling onto one option; it looks like clusters returning near-identical distributions — the population number is really one model prior wearing six hats. When differentiation is low, the output says so explicitly and tells you to read the result as a single undifferentiated estimate.

One deliberate confession: the thresholds that label differentiation “low”, “moderate” or “high” are a reporting convention, not calibrated constants — and every response declares this in a machine-readable field (bands_are_convention: true) rather than in a footnote. They will be fixed against the backtest corpus once it exists. Dispersion is a measure of the simulation’s internal structure: low differentiation does not mean the numbers are wrong, and high differentiation does not mean they are right. That claim needs real-world data — which brings us to Note 3.

Note 3 — Why we publish zero accuracy claims

In this category, accuracy claims are the norm — figures as high as 95% are advertised — yet to our knowledge no vendor publishes a customer-visible, per-market backtest score behind them. Samplio publishes no accuracy number at all. Not because we cannot compute one: the scoring pipeline (a direction-weighted blend against real, realised campaign outcomes) has been built and tested since July. It reports nothing because the backtest corpus is deliberately empty, and an honest zero beats an invented benchmark.

The corpus is empty because of what sourcing actually found. A case is admitted only with a verifiable source and a licence note. Award-show case studies turned out to publish no per-variant results at all; the freely available material splits into vendor A/B-test write-ups (self-reported, with obvious publication bias) and rigorous public-sector messaging RCTs (excellent data, different domain). A corpus mixed across those stacks changes what may honestly be claimed — so the domain mix is enforced in code: every accuracy output will carry its corpus composition, the weakest case determines the wording of the claim, and a mixed corpus can never be silently presented as brand-message accuracy.

The discipline has teeth. In August we verified an ideal-format case — a 2025 gambling-harms messaging RCT with 4,532 participants, four arms, per-arm outcomes and the exact stimulus texts. It was still rejected from the corpus: the trial ran in the UK and we have no UK market pack, and scoring a British trial with a German profile would produce a fake score — worse than an empty corpus. It sits in staging until a UK pack exists. (Fittingly, its top arms were a close race whose pooled difference was not significant — a real-world example of exactly the near-race rule in Note 1.)

Until a corpus produces a measured score, the engine keeps saying so mechanically: survey results are always stamped uncalibrated, placeholder data taints every derived output, no surface ever shows a classic ±% margin of error, and willingness-to-pay-shaped survey questions get an advisory (not a block) pointing to methods that can actually answer them. As far as we can tell, that makes Samplio the only synthetic research tool whose accuracy claim is, on purpose, nothing. When we finally publish a number, it will come with the corpus, the domain mix and the method attached.