The Twenty Minute VCThe Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO
At a glance
WHAT IT’S REALLY ABOUT
Simile’s vision: defensible behavior data to run human simulations at scale
- Simile’s core bet is that the winning AI companies will have defensible data strategies, and for human-behavior modeling that means acquiring representative real-world behavioral data rather than relying on what people say online.
- Park argues prediction alone is less valuable than counterfactual, causal insight—customers want to know how to change outcomes—so Simile prioritizes experiments like A/B tests and randomized trials as training signal.
- Simile positions itself as complementary to frontier LLMs: instead of maximizing rational intelligence, it aims to replicate human mistakes, biases, preferences, and values to better mirror real behavior.
- The company sees a compounding “data flywheel” where the world itself becomes ground truth: Simile generates many hypotheses, observes what actually happens, and iteratively improves simulation accuracy and usefulness.
- Beyond market research, Park frames simulation as “representation at scale,” with potential applications from product launches to elections, democratic stability, and other “wicked problems,” raising serious access and misuse questions.
IDEAS WORTH REMEMBERING
5 ideasDefensible data is the moat for next-gen AI companies.
Park’s thesis is that model capabilities will commoditize, so durable advantage comes from uniquely collecting hard-to-get data—especially representative human behavioral and experimental data.
Simile optimizes for human likeness, not superhuman rationality.
Where frontier labs aim for better reasoning/coding/math, Simile wants models that err and bias like humans, because matching real preferences and decision patterns is what makes simulations actionable.
“Say vs. do” is a structural weakness of web-trained models.
Park emphasizes that online text captures stated beliefs more than revealed behavior, so Simile collects transaction/observational data and partners with customers to ground models in what people actually do.
Prediction is secondary to causality for business users.
Enterprises don’t just want to know sales will drop; they want levers to prevent it, which requires counterfactual reasoning and causal mechanisms rather than pure correlation.
Randomized experiments are a high-value training asset for behavior models.
Simile runs and leverages A/B tests and RCT-style data to learn how interventions change behavior, making the system more useful for “what should we do?” decisions.
WORDS WORTH SAVING
5 quotesMy fundamental thesis here is for AI companies of this generation, you need to have an interesting data strategy that's going to be defensible.
— Joon Sung Park
No one really cares about prediction. No one really cares about what's going to happen in the future unless you're trying to predict the stock market. What people actually care about is they want to shape the future.
— Joon Sung Park
We want our models to be biased in the same way humans are. In a way, we want to be a representation of people's values, preferences, and taste, sort of their subjective half of their brain.
— Joon Sung Park
I actually think simulation has even better mechanism, which is the world is our ground truth.
— Joon Sung Park
I think there's a world in which in about two, three years we're running a single simulation session that's going to take ten, $20 million to run. A single session, but it's going to be so valuable that people will pay $100 million for it.
— Joon Sung Park
High quality AI-generated summary created from speaker-labeled transcript.