Trend Shift
Human Simulation Bets on Data Loops
Simile combines consented participant responses with generative agents for decision simulations at CVS Health. The key variable is shifting from whether agents can converse like people to whether data acquisition can support verifiable behavioral predictions.
The competition in human-behavior simulation may be moving from making agents sound human to acquiring consented, testable behavioral data continuously. Simile's CVS Health case suggests the approach is entering customer-journey and experiment-prioritization workflows. Its commercial importance lies not in replacing all research, but in whether exploratory decisions can be simulated faster and then repeatedly checked.
Commercial workflows are starting to absorb simulation
Simile says its work with CVS Health covers more than 400,000 participants, more than 200 behavioral scenarios, and 2.9 million consented responses. The system is used for customer journeys, NPS drivers, adherence, competitive positioning, and prioritizing follow-up tests. A new 20VC episode frames the company's edge as a distinctive data-acquisition strategy, while another describes its mission as simulating human behavior. The relevant change is that agent simulation is being placed inside decision workflows rather than remaining a demonstration of digital personas.
The mechanism connects interview memory to runnable agents
Earlier work by Joon Sung Park, Percy Liang, Michael S. Bernstein, and colleagues described generative agents using memory, reflection, and planning to maintain behavioral continuity, testing them in a town sandbox of 25 agents. A later study built 1,052 individual agents from qualitative interviews and reported that their accuracy in reproducing General Social Survey answers reached 85% of respondents' own two-week retest accuracy. In this framing, data acquisition is not merely training input: it may determine whether simulations retain enough contextual variation to be useful.
The consequence is a reordering of experiment budgets
If agents can screen customer-journey hypotheses, identify high-value segments, and rank offline tests, companies could reserve more budget for final validation instead of restarting research for every early question. The strongest countercase comes from PNAS: advanced LLM approaches largely failed to reproduce human behavioral distributions in 11- to 20-cent request games, with results changing unpredictably with wording, roles, and safety settings. Simulation is therefore better treated as a tool for prioritizing experiments than as a replacement for real participants.
What to watch next
The next observable evidence would be preregistered blind tests from Simile or its customers: agents predict real customer choices, retention, or NPS changes before subsequent experiments reveal actual outcomes. Repeated outperformance in ranking tests across industries would strengthen this claim; failure to transfer to new settings would weaken the commercial value of the data loop.
Sources
- The Twenty Minute VC (20VC) — The Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO
- The Twenty Minute VC (20VC) — The Company Simulating Human Behaviour
- Simile — CVS Health x Simile: Simulations for faster, safer decisions
- ACM — Generative Agents: Interactive Simulacra of Human Behavior
- arXiv — Generative Agent Simulations of 1,000 People
- Proceedings of the National Academy of Sciences — Take caution in using LLMs as human surrogates