Synthetic Consumer LabSynthetic Consumer Lab

Accuracy and limitations

.md

Every number in an SCL report comes from the responses of synthetic consumers (personas) played by language models. These responses are useful for seeing the direction and rough size of the difference between options; they are not a measurement of the real market.

Video in production
Product video3 min

Reading and reporting SCL results accurately

Shows where the warnings, confidence intervals and reasons sit in reports, and explains how to set up a second study with the same population and check the numbers in the Excel file before you share them.

What the results show and do not show

Shows Does not show
Which way preference leans between two products, prices or messages Sales, market share or revenue in the real market
Roughly how large the difference is A forecast precise to the decimal place
Which groups may respond differently The views of real people; personas are not real people
Hypotheses worth testing with real consumers Evidence that replaces research carried out with real consumers

Revenue values in price reports are calculated from the tested prices and the personas’ purchase decisions. Do not read them as a revenue forecast for your market.

What SCL does and does not do

What it does

  • Builds the population from population statistics and survey data. The personas’ age, education, province, income and household attributes are sampled from Turkish Statistical Institute (TurkStat) data and the TurkStat Household Budget Survey, and their values from the World Values Survey. If you apply no filters, the population’s distribution follows Türkiye as a whole.
  • Spreads responses across different language models. Each persona is assigned to one of several language models, so no single model’s habits determine all of the results.
  • Corrects known model habits. In survey questions, the order of the options is shuffled for each persona. Scale questions get an added statement reminding the persona that the middle option is only for genuine indecision. In eligible single-choice questions of a Multiple-Choice Survey, the answers of undecided personas are balanced according to the model’s own probabilities so that not every persona converges on the same answer; personas who are sure of their choice keep their answer. The responses you see in the report have been through these corrections.
  • Checks prices against the market. In pricing studies, SCL looks up the product’s current market price. If the prices you enter fall outside this reference, the Price might not be realistic dialog opens; if a market price was found, the dialog also shows the Estimated market price: and Typical range: values.
  • Shows uncertainty in the report. On the Insights tab of a Multiple-Choice Survey report, a warning line appears under findings that do not meet the reliability conditions. The Comparison report gives a 95% confidence interval. The Feature Price Impact report shows a 95% WTP Impact Interval for the feature’s effect on willingness to pay (WTP); if this interval does not include zero, the direction of the effect is more consistent.
  • Keeps a record of every study. For each study, an internal record keeps track of which population, questions, models and correction versions were used. This record is not shown in the app.

For details of these steps, see How synthetic consumers work.

What it does not do

If you have research on the same topic carried out with real consumers, download the report as an Excel file and put the two results side by side yourself. The download steps are in Exporting and downloading.

Known limitations

Agreeable and uniform responses

Language models tend to give positive, agreeable answers, drift towards the average and under-represent extreme views. The corrections above reduce these tendencies but do not remove them. In brand perception questions, if SCL has no background information about the brand, responses can come out more uniform and more positive. In a Multiple-Choice Survey, this information is matched to the brand name in the question text, so write your brand’s name clearly and correctly in the questions.

Small segments

  • The segment charts in the Price Analysis (WTP), Feature Price Impact and Competitive Analysis (Conjoint) reports do not show segments with fewer than 5 personas. If no segment reaches this threshold, some of these reports show the message “Insufficient segment data (at least 5 participants per segment required)”.
  • On the Insights tab of a Multiple-Choice Survey report, a segment finding counts as reliable only if the segment has at least 30 responses, the difference from the overall rate is at least 10 percentage points and the statistical test is significant (p<0.05). Under a finding that does not meet these conditions, a yellow warning line shows the condition it fails (e.g. “p>0.05” or “n=12<30”).
  • When you examine many questions and breakdowns at once, some “significant” results can appear by chance. On the Statistics tab, prioritise relationships whose Effect Size is “moderate” or “strong”.

For example, if you split a population of 100 personas by income group, which has ten bands, each group is left with 10 personas on average. If you plan to look at segments, enlarge the population or choose a coarser dimension, such as socio-economic status (SES) instead of income group. For details, see Segmentation.

New categories and narrow audiences

Language models know less about product categories that do not yet exist on the market and about very narrow audiences. Treat results in these areas as initial hypotheses and test them with real consumers.

Scenario and question wording

Results depend on how you describe the product, the price and the question; the same product can produce different results with a different description. In a Multiple-Choice Survey, each question is put to the persona separately and the persona does not see its earlier answers; only in a conditional question is it reminded of the answer that sets the condition. So write each question so that it makes sense on its own.

Personas say what they would do. Real purchases also depend on conditions such as shelf price, stock, sales channel and habit.

Variability between runs

Model responses are not the same every time. If you repeat the same study with the same population, the results change; a single headline figure, such as the optimal price in a pricing study, can shift noticeably between runs. Each new population is also sampled afresh, so two populations created separately with the same filters are not identical.

Personas that fail to produce a response are excluded from the results, so the participant count you see in the report can be lower than the population size.

Good practices

  1. Note the population and the scenario. Write down who you tested (population size, filters, vertical) and what you tested (product description, prices, question texts).
  2. Compare alternatives on the same population. When you compare two concepts, prices or messages, choose Select Existing Population in the population step of the second study (Use existing population in Fashion and Ad Testing). Otherwise part of the difference may come from differences between the populations.
  3. Repeat important results. Set up a second study with the same population and the same questions; in a Multiple-Choice Survey, you can bring in the questions with Import from Saved Survey. If the direction stays the same in both studies, you can trust the result more; if it changes, accept that there is no clear difference. The Refresh button on the report page does not ask the personas again; it only reloads the analysis.
  4. Look at the reasons alongside the numbers. The Themes and Responses sections of the Comparison report, focus group sessions and the persona comments in the Ad Test report show the personas’ reasoning. A Multiple-Choice Survey records only the choice; if you are looking for the answer to “why”, use one of these methods.
  5. Don’t skip the warnings. When you see a warning or a wide confidence interval, interpret the result more cautiously.
  6. Validate high-stakes decisions with real consumers. Use a small customer test or research with real consumers for this; the Comparison report makes the same recommendation.

Reporting findings inside your organisation

When you share SCL findings, your readers need to know that they come from a synthetic consumer study and under what conditions they were obtained. Follow these rules:

  • Name the source. Write “SCL synthetic consumer study” rather than “Consumer survey”.
  • Give the context. State the population size and filters, the vertical, the method, the study date and the question text.
  • Give the number with its denominator. Write “84 of 200 personas (42%)” rather than “42%”. Use the participant count shown in the report as the denominator, not the population size.
  • Pass on the uncertainty. Report confidence intervals and warnings together with the result; do not present differences from small segments as findings.
  • Report direction, not market share. Write “B was preferred over A in this population”, not “B takes 58% of the market”.
  • Check the PDF text. In most PDF reports, a language model writes the commentary and summary paragraphs based on the numbers in the report. Before you share it, compare this text with the numbers in the report.
  • Attach the source file. The Meta sheet of the Excel file contains the study name, the export date and the total numbers of personas and responses; in a Multiple-Choice Survey file, this sheet also holds the question texts and options. Share the Excel file along with the finding.
  • Don’t mix it with real data. If you show SCL results in the same table as real consumer research, state the source of each column separately.

Type a method, screen name or concept: for example “price”, “population”, “credits”.