Skip to content
The Twenty Minute VCThe Twenty Minute VC

The Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO

Joon Sung Park is the Founder and CEO of Simile, the AI simulation company building foundation models of human behaviour; allowing companies to test how real people may think, decide and act before making a decision in the real world. Simile has now raised $300 million in total, including a $200 million Series B announced last week at a $2 billion valuation, led by Greenoaks and Index Ventures. ----------------------------------------------- Timestamps: 00:00 Intro 01:26 The Valentine's Day Simulation That Put Joon on the Map 03:56 How Agents Got Memory, Planning & Reflection — The Origin Story 07:07 Simile vs. Frontier Models 07:59 Say vs Do: Why Behaviour Data Beats Survey Data 09:33 Prediction Is Overrated 12:31 Is Simile Just a Fancy Qualtrics? The TAM Question 15:48 How Many People Do You Need for Accurate Simulations? 16:56 The Data Flywheel: How the World Becomes the Ground Truth 19:59 Can Simile Simulate Elections, Democracies & Wicked Problems? 21:43 Who Should and Shouldn't Have Access to Simulation Technology? 33:07 How to Build a World-Class Research Team That Doesn't Quit 35:59 The Two Contradictory Superpowers That Make Great Founders 39:10 Competing for Research Talent in the Tens of Millions 43:28 From Researcher to CEO: What Joon Looks for in Academic Founders 44:33 Raising $300M in Six Months: How the Round Came Together 58:01 Is Simile the Future of Love and Matchmaking? 1:00:45 Quick Fire Round ---------------------------------------------------------------------------------------------- Subscribe on Spotify: https://open.spotify.com/show/3j2KMcZTtgTNBKwtZBMHvl?si=85bc9196860e4466 Subscribe on Apple Podcasts: https://podcasts.apple.com/us/podcast/the-twenty-minute-vc-20vc-venture-capital-startup/id958230465 Follow Harry Stebbings on X: https://twitter.com/HarryStebbings Follow Joon Sung Park on X: https://twitter.com/joon_s_pk Follow 20VC on Instagram: https://www.instagram.com/20vchq Follow 20VC on TikTok: https://www.tiktok.com/@20vc_tok Visit our Website: https://www.20vc.com Subscribe to our Newsletter: https://www.thetwentyminutevc.com/contact ----------------------------------------------- #20vc #harrystebbings #simulation #ai #ceo #founder

Joon Sung ParkguestHarry Stebbingshost
Aug 1, 20261h 5mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 1:27

    Simile’s big bet: foundation models of human behavior (and $100M simulation sessions)

    Joon and Harry frame Simile as a “simulation market” aiming to model and simulate human behavior at scale, not just build another AI app. Joon opens with the provocative claim that individual simulation runs could soon cost tens of millions to compute yet be valuable enough to sell for $100M.

    • Simile’s ambition: simulate future human behavior for decision-making
    • Thesis: AI winners need defensible, unique data acquisition strategies
    • Positioning: behavior foundation models vs general-purpose intelligence
    • Vision of high-stakes, extremely valuable simulation sessions
  2. 1:27 – 3:21

    The Valentine’s Day ‘Smallville’ demo: 25 agents that self-organized a party

    Joon explains the 2023 project that made him well-known: a game town populated by 25 LLM-driven agents with routines, jobs, relationships, and emergent social behavior. The simulation—set the day before Valentine’s Day—produced surprising self-organization like planning and decorating a party.

    • LLMs contain latent behavioral realism if “poked at the right angle”
    • Small town sandbox with 25 NPC agents living daily lives
    • Emergent coordination: agents planned parties and social events
    • Demo as an early proof that domain-agnostic human-like agents were possible
  3. 3:21 – 6:36

    How agents got memory, planning, and reflection (and why it matters)

    They dive into the technical origin of agent architectures: the need for agents to remember each other and sustain coherent long-horizon behavior. Joon describes early, simple memory storage plus a ‘reflection’ loop that synthesizes higher-level inferences—akin to shower thoughts—to form stable personalities and motivations.

    • Early memory hack: store experiences as natural language in markdown
    • Core limitation: context windows can’t hold a lifetime of experiences
    • Reflection mechanism: periodically summarize and infer higher-level goals/traits
    • Personality emerges by connecting repeated behaviors to deeper motivations
  4. 6:36 – 7:47

    Simile vs frontier LLMs: embracing human mistakes, bias, values, and taste

    Joon differentiates Simile’s objective from OpenAI/Anthropic-style frontier labs. Instead of optimizing for hyper-rational performance on math/coding, Simile wants models that replicate human irrationality, bias, and subjective preferences—so simulations mirror real people.

    • Frontier models: super-rational, “smart machine” optimization
    • Simile: represent human subjectivity (values, preferences, taste)
    • Goal: reproduce human errors and biases in-context
    • Simulation quality depends on psychological/behavioral fidelity
  5. 7:47 – 9:48

    ‘Say vs do’: why behavioral and experimental data beat surveys for causality

    Joon argues that web text is mostly ‘what people say,’ which diverges from what they do. Observational data helps correlation and prediction, but customers mostly need counterfactuals—what would happen if we changed X—requiring causal signals from experiments like RCTs and A/B tests.

    • Web data bias: captures stated opinions more than real behavior
    • Observational behavior data excels at correlations and forecasts
    • Customer need: shaping outcomes via counterfactual reasoning
    • Simile emphasizes experiments/RCTs/A-B tests as training signal
  6. 9:48 – 12:11

    Defensible data strategy: sourcing representative people and asking better questions

    They discuss why data acquisition is central and uniquely hard for human simulation models. Simile focuses on recruiting representative everyday people and eliciting deep context (life story, hard decisions) to capture the drivers behind behavior—not just surface preferences.

    • Defensibility in AI hinges on proprietary data strategy
    • Recruitment focus: representative populations, not elite experts
    • Data collection includes transactions, observation, and partner data
    • Deep elicitation: biographies and formative decision narratives
  7. 12:11 – 15:36

    Not ‘fancy Qualtrics’: simulations as interactive ecosystems and ‘wicked problem’ tools

    Harry challenges whether Simile is just next-gen survey tooling; Joon argues simulation goes beyond better surveys to multi-agent ecosystems and downstream consequence modeling. He extends the vision to societal-scale ‘wicked problems’ like climate coordination and democratic stability.

    • Core primitive: model a population, not just run a questionnaire
    • Future: simulate interacting populations and market dynamics
    • Applications: product launches, policy, systemic risk exploration
    • Potential to study equilibrium failures in collective-action problems
  8. 15:36 – 16:30

    Accuracy, sample sizes, and synthetic panels: scaling representation across segments

    Joon explains the practical question of how many people are needed for reliable simulations, referencing social science norms (often ~1,000 for narrow studies). The challenge is supporting arbitrary, on-the-fly filtering into subpopulations—pushing toward representing the entire population via synthetic panels.

    • More people enable more granular segmentation and filtering
    • Rule-of-thumb: ~1,000 participants can reach significance for narrow studies
    • Users demand dynamic filters, increasing coverage requirements
    • Trajectory: synthetic panels outgrowing traditional human panel markets
  9. 16:30 – 18:51

    The data flywheel: the world as ground truth (and why prediction alone is overrated)

    They unpack Simile’s learning loop: unlike tasks where rewards are immediate, simulation can generate huge numbers of testable hypotheses and validate them as real-world outcomes unfold. Joon argues this makes reality itself the evaluation stream, enabling compounding improvement over time.

    • Flywheel: compare simulated outputs to real-world outcomes
    • Analogy to coding agents: reward clarity differs, but validation still exists
    • Mechanism: generate many hypotheses and track which become verifiable
    • ‘World as ground truth’ becomes the continual supervision source
  10. 18:51 – 19:59

    Compute, unit economics, and the rise of ultra-expensive simulations

    Joon notes compute matters both for R&D exploration and for inference, but efficiency improvements can drastically reduce production costs (100x cheaper over time). Different simulations have different compute footprints, and he predicts a future where complex, multi-step simulations are extremely costly yet ROI-positive for major decisions.

    • Compute spend concentrates in exploration to find the right approach
    • Production model cost reduced ~100x through efficiency gains
    • Simulation complexity drives variable inference cost and pricing potential
    • Forecast: $10–$20M compute runs that customers pay $100M for
  11. 19:59 – 30:04

    Go-to-market reality: fast enterprise adoption, tight paid feedback loops, and PMF proof

    Joon explains why Simile started with enterprises: budgets, clear use cases, and—critically—tight feedback loops where paid customers validate accuracy. He shares how Fortune 500 leaders reacted to the Stanford demo, how Simile worked to prove accuracy, and how deals can close in ~3 months because the pain is acute.

    • Enterprise wedge: budget + validation + ‘paid feedback’ discipline
    • Fortune 500 pull after the Smallville demo signaled PMF
    • Validation claim: ~85% as accurate as people replicating their own attitudes/behaviors
    • Surprisingly fast sales cycles (as short as three months)
  12. 30:04 – 32:20

    Value capture and competitive framing: prevention, not just optimization (and vs prediction markets)

    Harry presses on pricing vs value created; Joon highlights disaster prevention as a major ROI driver, alongside optimization. They also compare Simile to prediction markets like Polymarket: Simile emphasizes ‘how and why’ outcomes happen, exposing intervention points rather than just a forecast.

    • Simulations can avert costly strategic mistakes (hundreds of millions)
    • Value is both prevention and optimization, with prevention as an easy ‘painkiller’ case
    • Differentiation from prediction markets: mechanism + steps + levers
    • Customer wow-moments: reproduce multi-month study results in minutes
  13. 32:20 – 38:51

    Building and retaining a world-class research team: balance, values, and ‘contradictory superpowers’

    Joon shares his hiring philosophy: build a complementary team that fills gaps (e.g., enterprise GTM leadership) while maintaining shared rigor and values. He looks for people who are ‘the common denominator of success’ and who combine rare, seemingly contradictory strengths—like paranoia today and deep long-term conviction.

    • Team balance: pair research excellence with product and GTM leadership
    • Hiring signal: candidates as repeat ‘common denominator’ of success
    • Retention focus: create a platform where individuals maximize their strengths
    • Archetype: contradictory superpowers (short-term paranoia + long-term faith)
  14. 38:51 – 43:14

    Research talent economics, retention in the Bay, and what makes academic founders investable

    They discuss the escalating cost of elite researchers (sometimes tens of millions in comp) and why vision and impact can still win recruits. For investing in academic spinouts, Joon’s key filter is whether founders are ‘married to impact’ (building something that reaches users and revenue) rather than ‘married to a problem.’

    • Reality: top research talent comp can reach tens of millions
    • Winning recruits with vision, impact, and belief the ambition can be real
    • Retention through trust and enabling people’s best work
    • Investor lens: choose founders driven by impact, not only intellectual fascination
  15. 43:14 – 53:52

    Raising $300M fast: insider preempts, Greenoaks process, and the froth vs fundamentals debate

    Joon walks through the rapid fundraising sequence: a $100M round, then an insider-driven preempt, then bringing in Greenoaks to raise $200M more—$300M total in ~6 months. He explains why they took capital they didn’t strictly need (compute/data acceleration), what he learned about VCs, and how he keeps focus amid frothy markets by anchoring on customers, demand, and technical progress.

    • Round dynamics: insider preempt after strong traction and tech progress
    • Greenoaks joined quickly due to prior market work and timing
    • Rationale: increase inputs (compute/data) to accelerate uncertain research outcomes
    • Awareness of froth; emphasis on fundamentals (customers, pull, model progress)
  16. 53:52 – 1:05:02

    Time-machine future: GPU-like collective intelligence, simulated twins, hedge funds—and love & matchmaking

    In a forward-looking segment, Joon frames frontier LLMs as ‘CPU intelligence’ and simulations as ‘GPU intelligence’—collective, diverse agents producing emergent societal behavior. They explore implications like personal simulated twins, market effects (even hedge fund ideas), and end with a discussion on whether simulation changes dating—where Joon argues shared lived experience will remain core.

    • CPU vs GPU analogy: single supermodel vs collective emergent intelligence
    • Representation at scale: simulated ‘twins’ as a new societal layer
    • Speculation: finance/markets could change under near-perfect simulation
    • Dating: efficiency vs romance; Joon emphasizes shared journey and memory

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.