The Twenty Minute VCThe Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO
CHAPTERS
- 0:00 – 1:27
Simile’s big bet: foundation models of human behavior (and $100M simulation sessions)
Joon and Harry frame Simile as a “simulation market” aiming to model and simulate human behavior at scale, not just build another AI app. Joon opens with the provocative claim that individual simulation runs could soon cost tens of millions to compute yet be valuable enough to sell for $100M.
- •Simile’s ambition: simulate future human behavior for decision-making
- •Thesis: AI winners need defensible, unique data acquisition strategies
- •Positioning: behavior foundation models vs general-purpose intelligence
- •Vision of high-stakes, extremely valuable simulation sessions
- 1:27 – 3:21
The Valentine’s Day ‘Smallville’ demo: 25 agents that self-organized a party
Joon explains the 2023 project that made him well-known: a game town populated by 25 LLM-driven agents with routines, jobs, relationships, and emergent social behavior. The simulation—set the day before Valentine’s Day—produced surprising self-organization like planning and decorating a party.
- •LLMs contain latent behavioral realism if “poked at the right angle”
- •Small town sandbox with 25 NPC agents living daily lives
- •Emergent coordination: agents planned parties and social events
- •Demo as an early proof that domain-agnostic human-like agents were possible
- 3:21 – 6:36
How agents got memory, planning, and reflection (and why it matters)
They dive into the technical origin of agent architectures: the need for agents to remember each other and sustain coherent long-horizon behavior. Joon describes early, simple memory storage plus a ‘reflection’ loop that synthesizes higher-level inferences—akin to shower thoughts—to form stable personalities and motivations.
- •Early memory hack: store experiences as natural language in markdown
- •Core limitation: context windows can’t hold a lifetime of experiences
- •Reflection mechanism: periodically summarize and infer higher-level goals/traits
- •Personality emerges by connecting repeated behaviors to deeper motivations
- 6:36 – 7:47
Simile vs frontier LLMs: embracing human mistakes, bias, values, and taste
Joon differentiates Simile’s objective from OpenAI/Anthropic-style frontier labs. Instead of optimizing for hyper-rational performance on math/coding, Simile wants models that replicate human irrationality, bias, and subjective preferences—so simulations mirror real people.
- •Frontier models: super-rational, “smart machine” optimization
- •Simile: represent human subjectivity (values, preferences, taste)
- •Goal: reproduce human errors and biases in-context
- •Simulation quality depends on psychological/behavioral fidelity
- 7:47 – 9:48
‘Say vs do’: why behavioral and experimental data beat surveys for causality
Joon argues that web text is mostly ‘what people say,’ which diverges from what they do. Observational data helps correlation and prediction, but customers mostly need counterfactuals—what would happen if we changed X—requiring causal signals from experiments like RCTs and A/B tests.
- •Web data bias: captures stated opinions more than real behavior
- •Observational behavior data excels at correlations and forecasts
- •Customer need: shaping outcomes via counterfactual reasoning
- •Simile emphasizes experiments/RCTs/A-B tests as training signal
- 9:48 – 12:11
Defensible data strategy: sourcing representative people and asking better questions
They discuss why data acquisition is central and uniquely hard for human simulation models. Simile focuses on recruiting representative everyday people and eliciting deep context (life story, hard decisions) to capture the drivers behind behavior—not just surface preferences.
- •Defensibility in AI hinges on proprietary data strategy
- •Recruitment focus: representative populations, not elite experts
- •Data collection includes transactions, observation, and partner data
- •Deep elicitation: biographies and formative decision narratives
- 12:11 – 15:36
Not ‘fancy Qualtrics’: simulations as interactive ecosystems and ‘wicked problem’ tools
Harry challenges whether Simile is just next-gen survey tooling; Joon argues simulation goes beyond better surveys to multi-agent ecosystems and downstream consequence modeling. He extends the vision to societal-scale ‘wicked problems’ like climate coordination and democratic stability.
- •Core primitive: model a population, not just run a questionnaire
- •Future: simulate interacting populations and market dynamics
- •Applications: product launches, policy, systemic risk exploration
- •Potential to study equilibrium failures in collective-action problems
- 15:36 – 16:30
Accuracy, sample sizes, and synthetic panels: scaling representation across segments
Joon explains the practical question of how many people are needed for reliable simulations, referencing social science norms (often ~1,000 for narrow studies). The challenge is supporting arbitrary, on-the-fly filtering into subpopulations—pushing toward representing the entire population via synthetic panels.
- •More people enable more granular segmentation and filtering
- •Rule-of-thumb: ~1,000 participants can reach significance for narrow studies
- •Users demand dynamic filters, increasing coverage requirements
- •Trajectory: synthetic panels outgrowing traditional human panel markets
- 16:30 – 18:51
The data flywheel: the world as ground truth (and why prediction alone is overrated)
They unpack Simile’s learning loop: unlike tasks where rewards are immediate, simulation can generate huge numbers of testable hypotheses and validate them as real-world outcomes unfold. Joon argues this makes reality itself the evaluation stream, enabling compounding improvement over time.
- •Flywheel: compare simulated outputs to real-world outcomes
- •Analogy to coding agents: reward clarity differs, but validation still exists
- •Mechanism: generate many hypotheses and track which become verifiable
- •‘World as ground truth’ becomes the continual supervision source
- 18:51 – 19:59
Compute, unit economics, and the rise of ultra-expensive simulations
Joon notes compute matters both for R&D exploration and for inference, but efficiency improvements can drastically reduce production costs (100x cheaper over time). Different simulations have different compute footprints, and he predicts a future where complex, multi-step simulations are extremely costly yet ROI-positive for major decisions.
- •Compute spend concentrates in exploration to find the right approach
- •Production model cost reduced ~100x through efficiency gains
- •Simulation complexity drives variable inference cost and pricing potential
- •Forecast: $10–$20M compute runs that customers pay $100M for
- 19:59 – 30:04
Go-to-market reality: fast enterprise adoption, tight paid feedback loops, and PMF proof
Joon explains why Simile started with enterprises: budgets, clear use cases, and—critically—tight feedback loops where paid customers validate accuracy. He shares how Fortune 500 leaders reacted to the Stanford demo, how Simile worked to prove accuracy, and how deals can close in ~3 months because the pain is acute.
- •Enterprise wedge: budget + validation + ‘paid feedback’ discipline
- •Fortune 500 pull after the Smallville demo signaled PMF
- •Validation claim: ~85% as accurate as people replicating their own attitudes/behaviors
- •Surprisingly fast sales cycles (as short as three months)
- 30:04 – 32:20
Value capture and competitive framing: prevention, not just optimization (and vs prediction markets)
Harry presses on pricing vs value created; Joon highlights disaster prevention as a major ROI driver, alongside optimization. They also compare Simile to prediction markets like Polymarket: Simile emphasizes ‘how and why’ outcomes happen, exposing intervention points rather than just a forecast.
- •Simulations can avert costly strategic mistakes (hundreds of millions)
- •Value is both prevention and optimization, with prevention as an easy ‘painkiller’ case
- •Differentiation from prediction markets: mechanism + steps + levers
- •Customer wow-moments: reproduce multi-month study results in minutes
- 32:20 – 38:51
Building and retaining a world-class research team: balance, values, and ‘contradictory superpowers’
Joon shares his hiring philosophy: build a complementary team that fills gaps (e.g., enterprise GTM leadership) while maintaining shared rigor and values. He looks for people who are ‘the common denominator of success’ and who combine rare, seemingly contradictory strengths—like paranoia today and deep long-term conviction.
- •Team balance: pair research excellence with product and GTM leadership
- •Hiring signal: candidates as repeat ‘common denominator’ of success
- •Retention focus: create a platform where individuals maximize their strengths
- •Archetype: contradictory superpowers (short-term paranoia + long-term faith)
- 38:51 – 43:14
Research talent economics, retention in the Bay, and what makes academic founders investable
They discuss the escalating cost of elite researchers (sometimes tens of millions in comp) and why vision and impact can still win recruits. For investing in academic spinouts, Joon’s key filter is whether founders are ‘married to impact’ (building something that reaches users and revenue) rather than ‘married to a problem.’
- •Reality: top research talent comp can reach tens of millions
- •Winning recruits with vision, impact, and belief the ambition can be real
- •Retention through trust and enabling people’s best work
- •Investor lens: choose founders driven by impact, not only intellectual fascination
- 43:14 – 53:52
Raising $300M fast: insider preempts, Greenoaks process, and the froth vs fundamentals debate
Joon walks through the rapid fundraising sequence: a $100M round, then an insider-driven preempt, then bringing in Greenoaks to raise $200M more—$300M total in ~6 months. He explains why they took capital they didn’t strictly need (compute/data acceleration), what he learned about VCs, and how he keeps focus amid frothy markets by anchoring on customers, demand, and technical progress.
- •Round dynamics: insider preempt after strong traction and tech progress
- •Greenoaks joined quickly due to prior market work and timing
- •Rationale: increase inputs (compute/data) to accelerate uncertain research outcomes
- •Awareness of froth; emphasis on fundamentals (customers, pull, model progress)
- 53:52 – 1:05:02
Time-machine future: GPU-like collective intelligence, simulated twins, hedge funds—and love & matchmaking
In a forward-looking segment, Joon frames frontier LLMs as ‘CPU intelligence’ and simulations as ‘GPU intelligence’—collective, diverse agents producing emergent societal behavior. They explore implications like personal simulated twins, market effects (even hedge fund ideas), and end with a discussion on whether simulation changes dating—where Joon argues shared lived experience will remain core.
- •CPU vs GPU analogy: single supermodel vs collective emergent intelligence
- •Representation at scale: simulated ‘twins’ as a new societal layer
- •Speculation: finance/markets could change under near-perfect simulation
- •Dating: efficiency vs romance; Joon emphasizes shared journey and memory