Skip to content
No PriorsNo Priors

Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin

Is the future of frontier AI open or closed? ReflectionAI co-founder and CEO Misha Laskin joins Sarah Guo and Elad Gil to discuss the launch of Beam, the company’s 500 billion parameter open-weight reasoning model. Misha breaks down the pre-training and reinforcement learning required to produce Beam’s reasoning efficiency, the shift in enterprise compute from renting to owning intelligence, and how he believes that open models will capture the majority of global token demand. He also talks about the open model ecosystem in China and why competition with China’s models is a good thing, safety considerations around open models, and how frontier-level open models may accelerate the pace of scientific discovery. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @MishaLaskin | @reflection_ai Chapters: 00:42 – Misha Laskin Introduction 00:59 – Latest with ReflectionAI 02:07 – Challenges Building an Open Model 03:56 – ReflectionAI’s Agentic Shift 07:37 – Resources for Model Training 10:09 – Scaling Efficiency 12:18 – Training Beam 16:51 – Where Model Value Comes From 19:55 – Beam’s Reasoning Efficiency 22:35 – Monetizing Open Weight Models 25:03 – Future of Open Versus Closed Tokens 31:29 – Competition with Open Models 34:38 – Chinese Open Source Model Ecosystem 38:40 – Will Chinese Models Remain Open 44:44 – Safety and Open Models 53:02 – Debating Access to Powerful Tools 57:16 – AI and Scientific Progress 01:00:36 – Data Centers and Jobs 01:02:07 – Beam and Scientific Research 01:05:37 – Research Head Count to Compute Ratio 01:10:29 – Conclusion

Misha LaskinguestElad GilhostSarah Guohost
Oct 9, 20261h 9mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:42

    Why open models can improve security: “enough eyeballs” for safety

    Misha argues that restricting model capabilities can also cripple defenses, especially in cybersecurity. He frames open weights as a practical safety strategy: broader scrutiny surfaces vulnerabilities faster than a small internal safety team can.

    • •Open vs. closed models as a security tradeoff: defense suffers when offense is removed
    • •Closed labs can’t cover the long tail of unintended consequences with limited safety headcount
    • •Real-world example: open models used to remediate harm caused by a closed model
    • •Linus’s Law applied to AI safety/security vulnerabilities
  2. 0:42 – 2:06

    ReflectionAI’s mission and sprint to build a frontier open-model lab

    Elad introduces Misha and ReflectionAI’s goal: “frontier open intelligence.” Misha describes the rapid build-out from a small startup into a full-stack training organization and the release of their first open model, Beam.

    • •Mission: build frontier open intelligence and make it widely accessible
    • •Scaling the org: ~30 people to ~300 in about a year
    • •Standing up teams across pretraining, mid-training, RL, infra
    • •First end-to-end model release: Beam
  3. 2:06 – 3:55

    What’s hardest about building open frontier models: everything is coupled

    Asked about challenges (compute, talent, data), Misha emphasizes that success requires getting dozens of interdependent details right. He contrasts the stability of mature labs with the need for startups to build foundational tooling from scratch.

    • •No single bottleneck: must get ‘30 things right’ simultaneously
    • •Talent acquisition/retention depends on mission and culture
    • •Compute procurement is only useful with robust infrastructure
    • •Big labs’ “taken-for-granted” tools must be rebuilt in a new lab
  4. 3:55 – 7:37

    From RL-first startup to building models end-to-end: the agentic shift

    Misha explains ReflectionAI’s original bet: apply reinforcement learning to make models agentic in math/coding, leveraging an open base model ecosystem. They shifted to full pretraining after realizing top open bases were largely Chinese and that RL at scale needs tight coupling with pretraining.

    • •Early context: chat-centric era with most compute in pretraining
    • •Founders’ RL lineage (Gemini RL; AlphaGo experience) shaped strategy
    • •RL progressed faster than expected across industry (e.g., o1 era)
    • •Need for a Western open base model + pretraining/RL coupling drove the pivot
  5. 7:37 – 10:42

    The economics of catching up to the frontier: hundreds of millions → billions → tens of billions

    The conversation turns to capital requirements and how frontier costs scale over time. Misha offers rough orders of magnitude and a compute-scaling heuristic (generation-to-generation ~4×) tied to advancing GPU generations and cluster sizes.

    • •Catching up used to be ‘hundreds of millions’; now ‘single-digit billions’
    • •Next year’s frontier could imply ‘tens of billions’ in spend
    • •Heuristic: ~4× compute per model generation (H100 → Blackwell → Rubin)
    • •Frontier spending may asymptote due to CapEx/revenue constraints
  6. 10:42 – 12:22

    Scaling efficiency and models improving themselves: faster progress per FLOP

    Misha argues compute efficiency gains are accelerating because models increasingly help build and improve models. He contrasts historical pretraining efficiency (measurable loss-based gains) with large headroom in RL where automated improvement loops can be much faster.

    • •Efficiency gains are concrete: reaching the same loss with fewer FLOPs
    • •Pretraining is more mature; RL has substantial remaining headroom
    • •Claimed acceleration: models in the loop can yield ~4–5× faster progress vs humans
    • •Intelligence per training FLOP and revenue per FLOP both increase over time
  7. 12:22 – 16:51

    How Beam was trained: compute, parameters, and RL dominating the budget

    Misha shares Beam’s training profile and highlights that reinforcement learning consumed more compute than pretraining. He frames RL as producing systems that keep improving (AlphaGo-style), turning capability improvements into an economic decision: how much compute to invest for gains.

    • •Beam described as 500B total parameters, 23B active (MoE-style)
    • •Pretraining: ~6,000 GB300s for weeks (now ~12 days with infra efficiency)
    • •RL: ~10,000+ GB300s for ~4 weeks; includes large-scale inference + sandboxes
    • •RL curves ‘never stop learning’; improvements become compute/economics-limited
  8. 16:51 – 19:55

    Where model value comes from: adaptive intelligence plus high-value enterprise data loops

    Misha discusses ‘jagged’ generalization: models adapt quickly with the right evaluations and synthetic data, but don’t always transfer cleanly across harnesses. He argues value comes from locating economically valuable task pools and building systems/evals that enable rapid adaptation.

    • •Strong agentic baseline enables faster adaptation to new task harnesses
    • •Generalization can be jagged across benchmarks/harness implementations
    • •Key lever: build evaluations, then generate synthetic data approximating tasks
    • •High-value verticals: finance (KYC/compliance), cyber defense, legal
  9. 19:55 – 22:36

    Beam’s reasoning efficiency: faster agents via large-scale RL

    Beam is positioned as a workhorse for coding and agentic tasks with notably lower time-to-solve than peers. Misha attributes this to a strong reasoning-oriented pretraining base amplified by what he claims is unprecedented open-source-scale RL, yielding faster, cheaper workloads.

    • •Efficiency matters alongside raw capability (latency/cost for agent workflows)
    • •Claimed 3–4× efficiency vs similar-class models; up to ~10× vs larger ones
    • •Cause: prioritize reasoning pretraining + massive-scale RL amplification
    • •RL encourages solving tasks in fewer steps (AlphaGo analogy: less ‘meandering’)
  10. 22:36 – 25:05

    Monetizing open weights: from ‘renting tokens’ to ‘owning intelligence’

    Misha compares closed-model token purchases to renting an apartment: you rent an integrated stack. Open weights enable ownership, but require the surrounding deployment stack (inference, cluster management, harnesses) and often services to drive adoption and compute demand.

    • •Closed models sell ‘rented’ inference across an integrated stack
    • •Open models enable ownership/control but require more operational tooling
    • •ReflectionAI’s role: provide deployment tooling + services for enterprises/sovereigns
    • •Services act as a demand driver that unlocks high-volume inference workloads
  11. 25:05 – 31:31

    Open vs closed token mix and customization: fine-tuned models vs customized systems

    The group predicts open tokens will dominate over time, akin to Linux in servers, while closed providers remain highly valuable. Misha argues most current dedicated inference workloads are customized/fine-tuned (driven by AI natives), but enterprises will more often customize systems/harnesses before fine-tuning models.

    • •Observed shift at gateways: from ~70/30 closed/open to ~70/30 open/closed
    • •Analogy: open OS dominance (Linux) with valuable closed ecosystems (Microsoft/Apple)
    • •Most dedicated inference today appears customized because AI-native customers dominate
    • •Enterprise path: closed → open → system customization; fine-tuning comes later
  12. 31:31 – 34:38

    Competing in a commoditizing stack: intelligence density × compute × trust

    Misha proposes a simple competitive framework and emphasizes compute scarcity as a differentiator. He expects hypercompetition and margin compression across models, apps, inference layers, and even bare metal, making durable revenue dependent on capability, compute access, and customer trust.

    • •Winning formula: intelligence density × compute × trust/solutions delivery
    • •Compute scarcity shapes competition and can be strategic for leading players
    • •Open models increase pricing pressure and compress margins across the stack
    • •Trust comes from consistently solving customer problems, not just delivering a model
  13. 34:38 – 44:44

    China’s open model ecosystem: benefits, incentives, and geopolitical ‘Trojan horses’

    Misha views Chinese open models as a global benefit but argues the West needs competitive alternatives. He outlines China’s advantages (distillation, data permissiveness, support) and hypothesizes why openness may persist: captive domestic monetization and geopolitical leverage via infrastructure ecosystems.

    • •Chinese open models enabled many Western businesses; lack of openness would concentrate power
    • •Advantages cited: large-scale distillation, cheaper/copyright-permissive data, support runway
    • •Openness can drive ecosystem lock-in (models as ‘Trojan horses’ for infra stacks)
    • •Geopolitical parallel: exporting foundational infrastructure creates long-term leverage
  14. 44:44 – 56:12

    Safety, openness, and access to powerful tools: cyber defense vs offense and ‘boring’ alignment

    Misha argues safety debates are often dominated by doomsday framing, while near-term risks are more empirical (cyber, misuse, misalignment bugs). He claims openness is the historical default for secure software (encryption precedent) and that alignment in practice is iterative ‘patching,’ best done with broader participation.

    • •Safety is multifaceted: cyber, bio/terror, existential; time horizons differ
    • •Encryption history: closed designs failed; open standards enabled modern cybersecurity
    • •Alignment described as Whac-A-Mole patching with data + detection tools (often LMs)
    • •Restricting capabilities can harm defenders; hard to separate cyber offense/defense
  15. 56:12 – 1:01:02

    AI’s upside: accelerating science, real-world experimentation, and data-center job creation

    Misha highlights scientific acceleration as the most exciting near-term benefit, citing rapid improvements in models’ ability to solve PhD-level problems. He also points to life sciences and materials experiments and argues data centers can act like modern factories, creating significant local employment and tax bases.

    • •Models evolved from chat to undergraduate to PhD-level solutions on technical work
    • •Iteration speed could compress research cycles dramatically (“PhD a week” metaphor)
    • •Big potential in lab-in-the-loop science: biology, chemistry, materials
    • •Data centers as ‘new factories’: construction/ops jobs and local economic impact
  16. 1:01:02 – 1:09:44

    How ReflectionAI directs research after Beam: model-in-the-loop, scaling bets, and headcount-to-compute ratios

    Misha explains prioritizing known high-impact improvements before riskier explorations, while acknowledging some phenomena only appear at scale. He describes using models as fast collaborators for research and infrastructure work, and discusses why teams may shrink per unit effort but expand ambition overall.

    • •Roadmap driven by known wins first, then expanding risk profile with scaling experiments
    • •Some breakthroughs only emerge at large scale, complicating small-scale testing
    • •Models as ‘fast, eager colleagues’ that accelerate researcher iteration
    • •Staffing constraints remain: maintain a headcount-to-compute ratio; growth shifts to applied deployment roles

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.