# How synthetic consumers work

Explains which Turkish data sources persona attributes are drawn from, how a persona is built, how responses are generated and corrected, and how studies are recorded.

A synthetic consumer (persona) is a virtual consumer sampled statistically from Türkiye's population, income and survey data and played by a language model. SCL works in two stages: when a population is created, the personas' attributes are set by numerical sampling; when a study runs, each persona answers the questions in its own voice through a language model that is given its profile.

## Where persona information comes from

Persona attributes are drawn from probability tables based on the sources below. For attributes where you don't select a filter, the population follows the distribution of Türkiye's population aged 18 and over. When you set a filter or ratio for an attribute, your choice replaces the statistical default.

| Information | Source |
| --- | --- |
| Age, gender, education, marital status, household structure | TurkStat Address-Based Population Registration System (ABPRS) |
| Province distribution (81 provinces) | Province weights based on population data |
| Regional income differences | TurkStat regional (NUTS-2) household income data |
| Household income and budget | TurkStat income distribution statistics; indicators such as the consumer price index (CPI) and the minimum wage, used to adjust figures to current prices |
| Credit card ownership, online shopping, car and home ownership | TurkStat Household Budget Survey microdata |
| Values and lifestyle | World Values Survey data for Türkiye |
| Brand and market information (in some methods) | Dated briefs compiled from web and social media research |
| Market price (in pricing methods) | Online retail price listings |

## How a persona is built

A persona is built layer by layer, with each layer consistent with the ones before it; unrealistic combinations are discarded.

1. Demographics. Age, gender, education, marital status, work status, province, household size, news source, social media use and shopping habits are drawn in a fixed order, each conditioned on the ones before it. The persona is given a Turkish first name and surname.
2. Income group and SES. The persona is placed in one of ten income groups, from A1 to E2. The socio-economic status (SES) class (A, B, C1, C2, D, E) is not drawn directly; it is calculated from income group, education and work status.
3. Household budget. Household income is set from TurkStat income data according to the income group; essential expenses, discretionary spending, savings and instalment capacity are calculated from that income.
4. Values and lifestyle. Value scores are drawn together, consistent with one another, based on World Values Survey data. Based on these scores, the persona is placed in one of five lifestyle classes, such as **Çalışkan İyimser** (Industrious Optimist).
5. Personality. Big Five (OCEAN) personality scores (openness, conscientiousness, extraversion, agreeableness, neuroticism) are drawn on a 1–10 scale, in line with the lifestyle. Based on these scores, the persona is assigned one of ten consumer behaviour types (e.g. "Pragmatic Planners").
6. Behavioural tendencies. Six product-independent tendencies are scored: price-value consciousness, quality-brand orientation, deal proneness, novelty seeking, shopping hedonism and spending tendency. Scores are scaled so that the Türkiye average is 50.
7. Market segment. Each persona is assigned a market segment that fits its attributes, based on the segmentation you fetched with **Find segmentation** and accepted while creating the population. A persona that matches no segment may be left without one.
8. Custom feature. Features you add on the **Custom Feature** tab are given to personas at the share you specify. If **Smart Distribution** is on, a language model reviews the profiles and picks the personas the feature fits; if it is off, the feature is distributed at random.
9. Profile text. All of this information is turned into profile text using fixed templates. During a study, the language model receives the part of this information selected for the method.

Demographics, income, values, personality and behavioural tendencies are set by numerical sampling, without a language model. A language model is used in the **Smart Distribution** step, to assign market segments in some segmentations, and to assign vertical-specific archetypes in Fashion and Ad Testing (see [Verticals](/en/docs/research/verticals)). The [Persona profiles](/en/docs/populations/persona-profiles) page describes a profile's fields and where they appear.

## How responses are generated

### Each persona answers for itself

In a study, each persona's responses are generated separately: the language model is given the persona's profile and asked to answer the questions as that person. In most methods, personas are split across several language models, so that no single model's habits dominate the results.

What a persona is given depends on the method:

- In methods that ask about price, such as Price Analysis (WTP), Feature Price Impact and Competitive Analysis (Conjoint), the household budget also feeds into the decision. The budget is described in natural language with rounded amounts; no explicit spending cap is given, because a cap makes responses cluster at that amount.
- In methods that don't ask about price, such as Multiple-Choice Survey and Priority Ranking (MaxDiff), the detailed budget is not used; instead, the profile includes a summary of the income group, SES and monthly spending level.

In a Multiple-Choice Survey, each question is asked separately and the persona does not see its earlier answers; only for conditional questions is it reminded of the answer the condition depends on. If a question relies on another question, write the context it needs into the question text.

In this method, the persona only marks its choice and writes no reasoning. You can find answers with reasoning in the **Persona details** section of the Comparison report and in Real-Time Focus Group conversations.

### Supplementing with current information

Language models' knowledge of brands and prices can be out of date. SCL fills this gap in two ways.

Briefs: SCL's library holds dated briefs on some brands and markets, compiled from web and social media research. If a brand named in a survey question has a brief, the brief is matched automatically; in a focus group, matching is based on the product, the brand and the first message. Each persona receives a different part of the brief as background, depending on its media use and the channels it follows. You can't choose a brief yourself when setting up a study; Argus, however, can add relevant briefs to a population.

> [!TIP]
> Matching is based on the brand name in the question text. Write the brand's full name explicitly in the question text.

Market price: When you set up a pricing study, SCL finds a market reference for the product from online retail prices (or from a web search if the product is not in the listings). If the prices you enter look out of line with the market, the **Price might not be realistic** warning opens when you click **Next Step**; if a market price was found, the **Typical range:** and **Estimated market price:** lines are also shown. During the study, each persona is shown a market price range adjusted to its income group and price-quality tendencies, so that the persona judges prices against the current market rather than an outdated memory of prices.

### Combining and correcting responses

All personas' responses are combined into a single report. If no response can be obtained for a persona even after several attempts, that persona is left out of the results; this is why the number of personas in the report can be lower than the population size.

Language models have known habits: they are influenced by the order of options, drift towards the middle of scales, and pick the same, most likely answer for every persona even when they are undecided. In a Multiple-Choice Survey, SCL counters these with the following steps:

- The order of options is shuffled for each persona; if the survey has no conditional questions, the order of questions is shuffled too. Before analysis, each response is mapped back to its original option.
- Scale questions, such as agreement, satisfaction, likelihood and frequency, get an added sentence reminding the model that the middle option is only for genuine indecision and that the extremes are honest answers too.
- For single-choice questions with at most five options, the probability the model assigns to each option is read. A persona that is almost certain keeps its answer; an undecided persona's answer may be shifted to one of the options it finds plausible that has been chosen less often in the population than expected. Personas whose answers come from a model that does not return probability information are excluded from this step.
- If a brief containing positive and negative views of a brand named in the question was matched, the response distribution on five-point scale questions is moved, to a limited extent, towards the distribution of views in that brief. Here too, only undecided personas' answers change.

The report shows the responses after these corrections.

## How studies are recorded

For every study except focus groups, a record is opened before the first model call and completed when the study finishes. The record contains:

- the version of the software that ran the study,
- the population, questions, briefs and images used,
- data and settings versions, such as income data and segment rules,
- the models requested and how many times each model was actually called,
- errors that occurred during the study.

In addition, the input and response of every model call are stored separately. These records are not shown in the app; they are kept for SCL's internal audit.

## When comparing results

- In a reused population, the personas stay exactly the same. Compare two ideas, prices or messages on the same population.
- Each new population is generated with a new random draw; two populations created with the same filters do not contain the same personas.
- Language model responses can vary from call to call; if you run the same study again, the results will not be exactly the same.

> [!IMPORTANT]
> Results are language model output. SCL does not automatically compare study results with real consumer data. Read the results as an indication of direction and magnitude, and validate high-stakes decisions with real consumers. For details, see [Accuracy and limitations](/en/docs/getting-started/accuracy-and-limitations).

## Related pages

- [Persona profiles](/en/docs/populations/persona-profiles): the fields of a profile
- [Accuracy and limitations](/en/docs/getting-started/accuracy-and-limitations): the limits of the results
- [Creating a population](/en/docs/populations/creating-a-population): filters, ratios and segmentation
- [Custom features](/en/docs/populations/custom-features): adding your own features
- [Segmentation](/en/docs/results/segmentation): segments and breakdowns
