Lenny's PodcastThe ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)
CHAPTERS
- 0:00 – 1:02
“Test everything”: why even tiny changes can have outsized impact
Ronny opens with his core philosophy: every code change and feature should be run as an experiment because small “obvious” tweaks can produce surprising results. He also frames experimentation as a portfolio—mostly incremental bets plus some high-risk, high-reward swings that are expected to fail often.
- •Experiment on every meaningful code/feature change
- •Small fixes can have unexpected negative or positive effects
- •Use a portfolio mindset: incremental + high-risk/high-reward ideas
- •Expect most bold ideas to fail; plan for it
- 1:02 – 4:56
Ronny’s background across Amazon, Microsoft, and Airbnb—and what this episode covers
Lenny introduces Ronny Kohavi’s career and why he’s considered a leading expert in A/B testing. They set the agenda: practical experimentation advice, organizational culture, common pitfalls, and key statistical concepts.
- •Ronny’s roles at Airbnb, Microsoft, and Amazon
- •The episode’s focus: tactical guidance for better experiments
- •Themes previewed: trust, validity checks, p-values, and culture change
- 4:56 – 9:01
The famous Bing “two-line” ad change that increased revenue ~12%
Ronny shares a legendary case study: a seemingly minor ad formatting tweak on Bing triggered an alarm because revenue jumped dramatically. The team assumed it must be a bug, replicated multiple times, and confirmed it was real—becoming one of Bing’s biggest wins.
- •A trivial UI change (promoting a line of ad text) produced ~12% revenue lift
- •Initial reaction: ‘too good to be true’—hunt for instrumentation/metrics bugs
- •Replication and follow-on experiments helped isolate what drove the lift
- •Illustrates how poor humans are at predicting experiment outcomes
- 9:01 – 17:01
Surprising UX wins: opening results in a new tab, plus reusable “rules of thumb”
They discuss another counterintuitive winner: opening clicked items in a new tab, debated heavily but repeatedly beneficial across products. Ronny points to ways teams can reuse learnings—papers and pattern libraries that aggregate experimentation results.
- •Opening links in a new tab: repeatedly wins despite strong design pushback
- •Institutional memory problem: big winners get forgotten and re-learned later
- •Microsoft ‘Rules of Thumb’ paper: extracting patterns from thousands of tests
- •goodui.org as a repository of experiment-backed UI patterns
- 17:01 – 20:45
Institutional learning: documenting results so the org compounds knowledge
Ronny explains why experimentation doesn’t automatically create learning unless results are captured and shared. He advocates for searchable experiment histories and periodic reviews of the most “surprising” experiments—wins and losses—to build a learning flywheel.
- •Define “surprising” as big gaps between predicted vs. actual outcomes
- •Capture both surprising winners and surprising losers
- •Create searchable archives of all experiments and outcomes
- •Run quarterly reviews of the most interesting experiments to drive adoption
- 20:45 – 24:48
Incremental optimization vs. big bets: failure rates and the social-integration flop
Ronny counters the idea that experimentation only produces micro-optimizations: you should test everything, but allocate capacity to ambitious bets that may fail. He shares how high failure rates are normal and recounts Bing’s costly, long-running social-integration effort that ultimately failed despite heavy investment.
- •‘Test everything’ doesn’t mean ‘only do small ideas’—run a portfolio
- •Typical failure rates: ~66% at Microsoft, ~85% at Bing, ~92% in Airbnb search
- •Important nuance: these are experiment-level failure rates (including aborts/bugs)
- •Bing social integration (Twitter/Facebook feeds): ~100 person-years, negative/flat results, eventually aborted
- 24:48 – 26:28
When not to A/B test—and when the platform cost should approach zero
They discuss cases where A/B tests aren’t feasible (e.g., one-time decisions) and the constraints for valid testing (sufficient units/users). Ronny emphasizes that once you invest in a strong experimentation platform, the marginal cost of running tests should become near-zero—making ‘test everything’ practical.
- •Not everything is testable (e.g., mergers & acquisitions)
- •Small companies may lack enough users to detect effects reliably
- •A mature platform reduces marginal experiment cost toward zero
- •Less mature platforms often require heavy analyst/data-science support
- 26:28 – 28:00
When startups should begin experimenting: practical user-count thresholds
Ronny offers rule-of-thumb thresholds for when A/B testing becomes statistically viable for most product metrics. He urges smaller teams to focus on detecting big effects and to start building experimentation culture and infrastructure before they reach “scale.”
- •Tens of thousands of users: start experimenting, expect only large effects detectable
- •~200,000 users: ‘magic’ zone where broader testing becomes practical
- •Startups should prioritize detecting 5–10% effects, not 1% effects
- •Build culture and platform early so the org is ready when traffic scales
- 28:00 – 32:43
Overall Evaluation Criterion (OEC): optimizing for long-term value, not short-term wins
Ronny introduces the OEC as the essential anchor for experimentation: teams must clearly define what they’re optimizing for, including guardrails that prevent short-term revenue hacks from harming user value. He frames OEC design as being causally predictive of lifetime value.
- •OEC answers ‘what are we optimizing for?’—harder than it sounds
- •Revenue alone is dangerous; add guardrails and constraints (e.g., ad “pixel budget”)
- •Use metrics like successful sessions/time-to-success as experience guardrails
- •Goal: an OEC that is causally predictive of lifetime value (LTV)
- 32:43 – 36:31
Long-term effects: long holdouts vs. models (Amazon email ‘spam’ and unsubscribe cost)
They explore how to reason about long-term impact through either long-running experiments or predictive models. Ronny shares an Amazon example where a simplistic attribution metric encouraged spamming users until unsubscribe cost was modeled—revealing many campaigns were net negative and inspiring better unsubscribe UX.
- •Two approaches: long-term experiments for learning vs. models using historical data
- •Email attribution without countervailing metrics incentivizes over-sending/spam
- •Model the long-term cost of unsubscribes to estimate LTV impact
- •Adding the cost metric revealed >50% of campaigns were negative
- •Product insight: unsubscribe at campaign level (e.g., ‘unsubscribe from this author’) reduces harm
- 36:31 – 45:25
Why redesigns so often fail—and how to de-risk them with incremental experiments
Lenny and Ronny unpack the common pattern that big redesigns underperform and become hard to roll back due to sunk cost and organizational momentum. Ronny advocates decomposing redesigns into testable steps (OFAT or small batches) so teams learn what works before committing fully.
- •Large redesigns frequently produce negative results
- •Sunk cost fallacy makes teams reluctant to revert after long build cycles
- •Decompose redesigns into smaller testable components (OFAT / limited factors)
- •‘Flat’ results should generally be no-ship due to maintenance overhead
- •Even under constraints (e.g., legal), test options and ship the least harmful
- 45:25 – 50:01
Airbnb case study: experimentation in search, top-down shifts, and COVID-era upheaval
Ronny shares what he can about Airbnb: in search relevance, everything was A/B tested, even if other parts of the org were less test-driven. He argues that during volatile periods like COVID, experimentation becomes more important—not less—because assumptions are even more likely to be wrong.
- •Airbnb search relevance: ‘everything was A/B tested’ within his scope
- •Skepticism about narratives of success without knowing the counterfactual
- •During upheaval (COVID), external validity is in question—tests clarify what works now
- •In crises, you’re not suddenly more likely to be right; failure rates remain high
- •Example of a big COVID-era bet (online experiences) that didn’t pan out
- 50:01 – 55:26
Trust as the foundation: experimentation platforms as safety nets (and the Optimizely cautionary tale)
Ronny argues that trust is the most important property of an experimentation system: it must reliably prevent bad launches and produce credible results. He describes how early misuse of “real-time p-values” in tools like Optimizely inflated false positives and eroded trust in experimentation.
- •Platforms serve two roles: safety net (fast aborts) and reliable measurement (scorecards)
- •Trust is slow to build and easy to lose; invalid results poison culture
- •Real-time ‘stop when p<0.05’ inflates Type I error dramatically
- •A/A tests are a practical way to detect inflated false positives
- •Example: public stories of tools ‘almost getting someone fired’ due to misleading stats
- 55:26 – 1:06:20
Invalid experiments, SRM detection, Twyman’s law, and what p-values really mean
Ronny shares the most common red flag of flawed experiments—sample ratio mismatch (SRM)—and how frequently it occurs even at mature organizations. He then explains Twyman’s law (interesting results are often wrong) and clarifies common p-value misunderstandings, including why replication matters when underlying success rates are low.
- •SRM: deviations from intended traffic split can indicate invalid randomization or logging
- •SRM causes: bots, data pipeline issues, and skewed filtering/traffic sources
- •Organizations sometimes ignore SRM warnings—platform UX must prevent misuse
- •Twyman’s law: big/‘too good’ lifts warrant investigation and replication
- •P-value is not ‘probability treatment wins’; it’s conditional on the null
- •Low base rates (e.g., 8% success) increase false positive risk; replicate marginal wins
- 1:06:20 – 1:14:11
Rolling out experimentation: build vs. buy, culture change, platform maturity, and speed wins
Ronny outlines how teams should get started: consult internal experts if available, decide how to balance vendor tools with in-house build, and pick a “beachhead” team that ships frequently. He closes with platform practices that increase velocity without sacrificing rigor, including variance reduction techniques.
- •Build vs. buy is usually a mix; vendors are stronger today than in the past
- •Culture shift strategy: start with a team that launches frequently and has a clear OEC
- •Beware ‘OEC confusion’—if directionality isn’t agreed (e.g., time on site), optimization fails
- •Invest in self-serve platforms to reduce reliance on data scientists for basic reads
- •Speed: fast scorecards, variance reduction (capping), and CUPED to reduce required sample size
- 1:14:11 – 1:23:07
Lightning round: book picks, Chernobyl, interview questions, Blink cameras, and life ‘experiments’
In the lightning round, Ronny shares favorite books centered on skepticism and evidence, a standout TV series, and a quirky technical interview question. He also recommends Blink cameras and wraps with a broader lesson: apply a hierarchy of evidence to everyday claims.
- •Recommended books: Calling Bullshit; Hard Facts; Mistakes Were Made (But Not By Me)
- •Favorite show: Chernobyl (and lessons about communicating evidence)
- •Technical interview prompt: what ‘static’ means in C++
- •Favorite product: Blink cameras for practical home/yard monitoring
- •Hierarchy of evidence: controlled experiments > observational > anecdotes