Skip to content
No PriorsNo Priors

Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin

Is the future of frontier AI open or closed? ReflectionAI co-founder and CEO Misha Laskin joins Sarah Guo and Elad Gil to discuss the launch of Beam, the company’s 500 billion parameter open-weight reasoning model. Misha breaks down the pre-training and reinforcement learning required to produce Beam’s reasoning efficiency, the shift in enterprise compute from renting to owning intelligence, and how he believes that open models will capture the majority of global token demand. He also talks about the open model ecosystem in China and why competition with China’s models is a good thing, safety considerations around open models, and how frontier-level open models may accelerate the pace of scientific discovery. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @MishaLaskin | @reflection_ai Chapters: 00:42 – Misha Laskin Introduction 00:59 – Latest with ReflectionAI 02:07 – Challenges Building an Open Model 03:56 – ReflectionAI’s Agentic Shift 07:37 – Resources for Model Training 10:09 – Scaling Efficiency 12:18 – Training Beam 16:51 – Where Model Value Comes From 19:55 – Beam’s Reasoning Efficiency 22:35 – Monetizing Open Weight Models 25:03 – Future of Open Versus Closed Tokens 31:29 – Competition with Open Models 34:38 – Chinese Open Source Model Ecosystem 38:40 – Will Chinese Models Remain Open 44:44 – Safety and Open Models 53:02 – Debating Access to Powerful Tools 57:16 – AI and Scientific Progress 01:00:36 – Data Centers and Jobs 01:02:07 – Beam and Scientific Research 01:05:37 – Research Head Count to Compute Ratio 01:10:29 – Conclusion

Misha LaskinguestElad GilhostSarah Guohost
Oct 9, 20261h 9mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

ReflectionAI’s Beam: scaling open-weight frontier models with reinforcement learning

  1. Misha Laskin explains why ReflectionAI shifted from betting on RL atop existing open models to building frontier open-weight models end-to-end, culminating in the release of Beam.
  2. He argues that training frontier models is a tightly coupled engineering systems problem—talent, infrastructure, data, compute, and evaluation all must be executed well, and the frontier’s cost is rising from hundreds of millions toward billions and beyond.
  3. Beam is positioned as an efficient coding/agentic model whose key advantage is faster reasoning per task, attributed to unusually large-scale reinforcement learning runs in addition to strong pre-training.
  4. ReflectionAI’s business thesis frames open weights as “owning intelligence,” requiring tooling, deployment support, and agentic systems around the model—especially for enterprises moving from closed-model spend to owned infrastructure.
  5. The conversation situates open models within geopolitics and safety debates, claiming Chinese openness has benefited the world but creates strategic dependence, while openness can also improve real-world cybersecurity and vulnerability discovery.

IDEAS WORTH REMEMBERING

5 ideas

Pre-training + RL are now inseparable if you want frontier-level agentic performance.

ReflectionAI found that scaling reinforcement learning depends tightly on having a strong pre-trained base, so they committed to training models end-to-end rather than relying on external open models (which were increasingly China-led). This coupling also shaped their “agentic shift,” where coding/agentic capability becomes a foundation that can be adapted quickly with the right evals and synthetic data.

The core challenge of building frontier models is systems integration, not a single constraint.

Laskin frames model building as a “30 things must go right” endeavor—talent, culture, data, compute, and the infra to make compute usable—many of which are invisible inside mature labs like DeepMind. The hardest part isn’t a single bottleneck; it’s orchestrating many interdependent systems under time pressure.

Frontier catch-up costs are exploding, but efficiency and economics may force an asymptote.

He estimates the capital to “catch up” to the moving frontier has risen rapidly: ~hundreds of millions (18 months ago), to single-digit billions (recently), to potentially ~$10B next year, driven by ~4× compute per generation. He also argues this won’t scale forever because CapEx must eventually align with revenue potential and because efficiency gains are accelerating.

Beam’s differentiation is not just capability, but speed-per-task driven by heavy RL.

Beam is described as a 500B-parameter MoE (23B active) optimized for coding and agentic tasks, with an emphasis on “reasoning efficiency” (achieving results faster/cheaper). They attribute much of this to very large-scale RL (10k GB300s for ~4 weeks), which trains the system to extract capability with fewer steps—analogous to AlphaGo becoming less “meandering” over time.

Open-model monetization is about packaging the missing enterprise stack around weights.

Commercially, he describes open weights as “ownership” versus closed APIs as “rental,” noting enterprises typically rent first until spend and strategic reliance justify owning. ReflectionAI positions itself as the layer that makes ownership workable—deployment tooling, cluster/inference management, agentic harnesses, and services to unlock high-value use cases that then drive inference demand.

WORDS WORTH SAVING

5 quotes

When you remove cyber offensive capabilities, you also remove cyber defensive capabilities.

— Misha Laskin

The state of the world today is that we have a few hundred safety researchers within closed labs that understand how these things work, and despite their best intentions, it's impossible to cover the long tail of unintended consequences that these systems might have.

— Misha Laskin

Linus's Law, with enough eyeballs, all bugs become shallow. I have the belief that with enough eyeballs, most security and safety vulnerabilities become shallow as well.

— Misha Laskin

You actually have to get 30 things right, and that's why everything is hard, because you have to get 30 things right.

— Misha Laskin

Open models are Trojan horses for the infrastructure that they bring with them.

— Misha Laskin

Building a frontier open-weight labPre-training vs. post-training/RL compute mixBeam model architecture and reasoning efficiencyCapital/compute scaling and efficiency gainsEnterprise “rent vs. own” AI adoptionMonetizing open weights via tooling and servicesChinese open-model ecosystem, incentives, and geopolitics

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.