Synthetic Consumer Lab Synthetic Consumer Lab Try the demo
TR EN
Article — 2 min read

Can LLMs Capture Human Preferences?

The ability of LLMs to imitate human preferences and the effects of language and chain-of-thought methods.

Research Questions

  1. Can LLMs directly imitate human preferences and decision patterns?
  2. If perfect imitation is not feasible, can they reflect the diversity across customer segments (especially those differing by language)?
  3. Does the chain-of-thought conjoint approach make LLM decisions more human-like and help explain preference heterogeneity?

Results

  • LLMs (GPT-3.5 and GPT-4) exhibited greater impatience compared to humans; GPT-4’s discount rates were far above human levels.
  • GPT-3.5 displayed a lexicographic preference structure, producing outcomes inconsistent with human behavior.
  • The chain-of-thought method reduced GPT-4’s impatience, yet it remained more impatient than human respondents.
  • LLMs were able to reflect language-driven differences; models were more patient in weak-FTR (Future Time Reference) languages.
  • LLMs are not reliable for direct preference measurement, but they can be valuable tools for hypothesis generation and analyzing heterogeneity across segments.

Findings

  • Impatience:

    • GPT-4’s discount factor (δ) was significantly higher than human benchmarks; GPT-3.5 and GPT-4 rarely chose to delay the larger reward (22% and 16%, respectively).
  • Lexicographic Behavior (GPT-3.5):

    • Decisions were insensitive to interest rates and could not be rationalized by any plausible utility function—indicating a strictly lexicographic rule.
  • Effect of Chain-of-Thought:

    • Using chain-of-thought increased GPT-4’s likelihood of choosing the larger, later reward from 16% to 34.5%, and revealed thematic reasoning patterns (risk, uncertainty, opportunity cost).
  • Language Differences:

    • In weak-FTR languages (e.g., German, Mandarin), LLMs produced more patient choices—consistent with empirical findings in behavioral economics literature.
  • Thematic Analysis:

    • As delay length increased, references to “risk and uncertainty” rose systematically across model outputs.
  • LLM Models: 5

  • Synthetic Data: 2

  • Method: 5

  • Speed: 1

  • Ethics: 1

  • Accuracy: 4

  • Demographics: 3

If you would like to explore this research in more detail, click here to read the full paper.

Read the full text ← All articles
Design lab
Theme
Accent color
Web animation
Try the animation on the homepage hero by moving your mouse over it.