Skip to content
Lex Fridman PodcastLex Fridman Podcast

Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation | Lex Fridman Podcast #344

Noam Brown is a research scientist at FAIR, Meta AI, co-creator of AI that achieved superhuman level performance in games of No-Limit Texas Hold'em and Diplomacy. Please support this podcast by checking out our sponsors: - True Classic Tees: https://trueclassictees.com/lex and use code LEX to get 25% off - Audible: https://audible.com/lex to get 30-day free trial - InsideTracker: https://insidetracker.com/lex to get 20% off - ExpressVPN: https://expressvpn.com/lexpod to get 3 months free EPISODE LINKS: Noam's Twitter: https://twitter.com/polynoamial Noam's LinkedIn: https://www.linkedin.com/in/noam-brown-8b785b62/ webDiplomacy: https://webdiplomacy.net/ Noam's papers: Superhuman AI for multiplayer poker: https://par.nsf.gov/servlets/purl/10119653 Superhuman AI for heads-up no-limit poker: https://par.nsf.gov/servlets/purl/10077416 Human-level play in the game of Diplomacy: https://www.science.org/doi/10.1126/science.ade9097 PODCAST INFO: Podcast website: https://lexfridman.com/podcast Apple Podcasts: https://apple.co/2lwqZIr Spotify: https://spoti.fi/2nEwCF8 RSS: https://lexfridman.com/feed/podcast/ Full episodes playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuKi9nrraNbKKp4 Clips playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOeciFP3CBCIEElOJeitOr41 OUTLINE: 0:00 - Introduction 1:09 - No Limit Texas Hold 'em 5:02 - Solving poker 18:12 - Poker vs Chess 24:50 - AI playing poker 58:18 - Heads-up vs Multi-way poker 1:09:08 - Greatest poker player of all time 1:12:42 - Diplomacy game 1:22:33 - AI negotiating with humans 2:04:58 - AI in geopolitics 2:09:43 - Human-like AI for games 2:15:44 - Ethics of AI 2:19:57 - AGI 2:23:57 - Advice to beginners SOCIAL: - Twitter: https://twitter.com/lexfridman - LinkedIn: https://www.linkedin.com/in/lexfridman - Facebook: https://www.facebook.com/lexfridman - Instagram: https://www.instagram.com/lexfridman - Medium: https://medium.com/@lexfridman - Reddit: https://reddit.com/r/lexfridman - Support on Patreon: https://www.patreon.com/lexfridman

Noam BrownguestLex Fridmanhost
Dec 6, 20222h 29mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 6:08

    No-limit Hold’em basics: betting freedom, volatility, and discomfort

    Noam explains what makes no-limit Texas hold’em distinct: the ability to bet any amount creates explosive pots and psychological pressure. They discuss how ‘jumpy’ stakes change decision-making and why putting opponents in uncomfortable spots is central to good poker.

    • No-limit vs limit: unconstrained bet sizes and fast-escalating pots
    • Risk aversion when money at stake materially affects life
    • Big bets can strategically pressure opponents—but also cause huge mistakes
    • Poker as maximizing expected value (EV), not just winning hands
  2. 6:08 – 13:59

    Nash equilibrium in games: why “GTO” works (and what ‘in expectation’ means)

    The conversation introduces Nash equilibrium as a strategy that can’t be exploited in finite two-player zero-sum games. Noam uses rock–paper–scissors to build intuition, clarifying variance and what it means to be guaranteed not to lose in expectation.

    • Existence of optimal (unexploitable) strategies in two-player zero-sum games
    • ‘In expectation’ vs per-hand outcomes in high-variance poker
    • Why zero-sum matters for guarantees; contrast with multiplayer games like Risk
    • Finite/compact game caveats and edge cases (e.g., ‘name a bigger number’)
  3. 13:59 – 18:12

    How poker AIs learn: self-play and Counterfactual Regret Minimization (CFR)

    Noam details the learning process behind poker solving: self-play with counterfactual reasoning and regret signals. CFR updates action probabilities based on ‘regret’ and is proven to converge to Nash equilibria—even in imperfect-information settings like poker.

    • Self-play as general RL, not necessarily neural-network-based
    • Counterfactual reasoning: ‘what if I raised instead of called?’
    • Regret as a signal for adjusting future action probabilities
    • Convergence guarantees of CFR for imperfect-information games
  4. 18:12 – 24:47

    Why poker is harder than chess (arguably): imperfect information, mixing, and ranges

    They compare poker with perfect-information games like chess/go, emphasizing that poker requires not only choosing actions but choosing correct frequencies. Noam explains bluffing balance, ranges, and explicit ‘theory of mind’ reasoning used by strong bots.

    • Hidden cards create imperfect information and enable bluffing
    • Value depends on probabilities (mixing), not just move quality
    • Ranges: needing bets with both strong hands and bluffs to remain unexploitable
    • Bots explicitly reason about beliefs/common knowledge (theory of mind)
  5. 24:47 – 27:55

    GTO vs exploitative play: what Libratus demonstrated to the poker world

    Lex asks whether you play the cards or the player; Noam describes the long-standing GTO vs exploitative debate. Libratus’ result—crushing elite heads-up pros while aiming at Nash equilibrium—helped settle the argument in favor of game-theoretic play as a strong baseline.

    • Expert humans often start from equilibrium then deviate to exploit leaks
    • Pre-2017 skepticism: ‘read souls’ vs game theory optimal (GTO)
    • Libratus played balanced, non-adaptive equilibrium approximation and won big
    • Match scale: 120,000 hands over ~20 days with prize incentives
  6. 27:55 – 38:43

    Inside Libratus: from 2015 failure to search-driven real-time strategy

    Noam recounts how earlier systems lost when humans found weaknesses, motivating a major shift: real-time search during play. They unpack what ‘search’ means in poker—reasoning over action probabilities across many possible private hands—and how it’s combined with precomputed strategy/value estimates.

    • 2015 loss taught that precomputed strategies are vulnerable
    • Search adds ‘thinking time’ analogous to humans deliberating in tough spots
    • Poker search must consider strategies over all possible hidden hands (~1,000+)
    • Combining depth of lookahead with value estimates from precomputed solutions
  7. 38:43 – 50:33

    Overbets and human confusion: surprising strategic behaviors the AI discovered

    They discuss one standout behavior: Libratus using extreme overbets (e.g., 10x pot) that humans rarely used. The bot did it because it improved EV, but the side effect was putting humans into agonizing decisions—leading to strategic shifts in modern high-level poker.

    • Humans typically size bets relative to pot; Libratus sometimes bet far larger
    • Overbets create polarized meaning: ‘nuts or bluff,’ pressuring near-top hands
    • Humans spent minutes deliberating; bot simply optimized EV
    • Overbets later became common in elite poker after the match
  8. 50:33 – 58:28

    Engineering and scale: C++ performance, cluster resources, and match-day stress

    Noam describes the immense engineering effort: parallel C++ code, large CPU counts, memory, and constant optimization to simulate more games faster. He also describes the psychological stress of a long match where humans collaborated, shared notes, and even received full hand logs.

    • Implementation: C++, heavy parallelism, low-latency/communication optimizations
    • Compute scale: ~1,000 CPUs and terabytes of memory (large for the time)
    • Humans teamed up to find exploits; organizers provided detailed hand logs
    • Betting markets and day-by-day variance made the outcome emotionally volatile
  9. 58:28 – 1:04:04

    From heads-up to six-player poker: leaving zero-sum guarantees and scaling with depth-limited search

    Moving to multiplayer poker removes the clean two-player zero-sum Nash guarantees and introduces equilibrium-selection issues. Noam explains why equilibrium-style methods still worked in practice for six-player poker and how depth-limited search made the problem computationally feasible.

    • Multiplayer games: Nash equilibria exist but provide weaker guarantees
    • Equilibrium selection problem (many equilibria; coordination issues)
    • Poker remains highly adversarial, limiting cooperation/collusion (and it’s illegal)
    • Depth-limited search: look a few moves ahead, then use value estimates to scale
  10. 1:04:04 – 1:12:42

    Pluribus: dramatic cost reduction and why poker bots didn’t need neural nets (then)

    Noam highlights how Pluribus achieved superhuman multiplayer poker far more cheaply than Libratus due to algorithmic advances. Surprisingly, Libratus and Pluribus used no neural nets; the hardest part was scalable equilibrium-finding and correct mixing, not feature learning.

    • Compute cost comparison: Libratus (~$100k run) vs Pluribus (<$150 on AWS)
    • Cost drop driven by algorithms, not hardware progress
    • No neural nets in Libratus/Pluribus; later systems use nets for value functions
    • Poker value depends on beliefs (what players think others hold), unlike chess/go
  11. 1:12:42 – 1:23:22

    Diplomacy rules and feel: seven-player alliances, private negotiation, and simultaneous moves

    The conversation pivots to Diplomacy: a WWI-era map game where negotiation is the main action. Noam explains simultaneous order resolution, supports, backstabs, and why the game is ‘about people rather than pieces.’

    • Seven powers, alliance-building, and negotiation as the core mechanic
    • Private, unstructured natural-language messages; unlimited deal-making
    • Simultaneous moves enable betrayal (promises aren’t binding)
    • Victory by majority control is rare; draws are common and scoring varies
  12. 1:23:22 – 1:35:47

    Why Diplomacy is so hard for AI: language as the action space + cooperation dynamics

    Noam argues Diplomacy may be among the hardest AI benchmarks because the action space includes all possible utterances and the game requires managing cooperation, trust, and human conventions. Pure self-play breaks: agents may develop non-human protocols or strategies that fail with humans.

    • Natural language: breadth/depth of negotiation far beyond structured trading games
    • Self-play alone is insufficient; agents would invent ‘robot language’
    • Cooperative games require understanding human expectations and norms
    • Motivation: aiming for a decade-level challenge after rapid progress in other games
  13. 1:35:47 – 1:49:01

    Cicero’s design: intent-conditioned dialogue + planning/RL + safety/value filters

    Noam explains Cicero’s architecture: a language model fine-tuned for Diplomacy, controlled via ‘intents’ derived from strategic planning and RL. A separate evaluation/filtering layer helps avoid nonsensical or strategically harmful messages and reduces deceptive behavior because trust improves long-term outcomes.

    • Two-part system: strategic module computes intents; dialogue model turns intents into messages
    • Intents encode desired actions for self and requests for others
    • Message filters estimate downstream responses/EV to avoid harmful disclosures
    • Minimizing lies improved performance: trust is a durable resource in Diplomacy
  14. 1:49:01 – 2:09:44

    Human data as an anchor: regularized self-play, evaluation against varied skill levels, and real-world implications

    They discuss why human data is essential: self-play policies can be ‘rational’ yet socially disastrous with humans. Cicero uses a supervised ‘anchor policy’ trained on 50k human games and regularizes self-play toward human-likeness, then evaluates performance in mixed-skill lobbies similar to real conditions.

    • Human irrationality (anger/punishment) must be modeled to succeed
    • Anchor policy: supervised model of human play used as a constraint/regularizer
    • Dataset scale: ~50k games, 10M+ messages (from webdiplomacy.net)
    • Performance measurement: mixed-skill environments analogous to self-driving evaluation
  15. 2:09:44 – 2:29:21

    Beyond benchmarks: human-like bots, cheat detection, ethics of deception, AGI gaps, and advice for beginners

    The final stretch explores broader impacts: building human-like chess/go bots, the resulting cheating risks, and ethical issues around deception and anti-AI bias. They close on AGI challenges like data efficiency and Noam’s advice to new ML learners: build fundamentals and cultivate distinctive perspectives.

    • Human-like strong play can aid training but makes cheat detection harder
    • Ethics: deception capabilities, white lies vs harmful lies, and anti-AI bias in games
    • AGI: major gap is data efficiency; need better general planning and learning
    • Beginner advice: strong math/CS/stat foundations; don’t fear unconventional paths

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.