No PriorsBeam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
At a glance
WHAT IT’S REALLY ABOUT
ReflectionAI’s Beam and the case for frontier open-weight models
- ReflectionAI’s Misha Laskin describes the company’s sprint from a ~30-person startup to a ~300-person lab capable of training and releasing Beam, its first frontier-class open-weight model.
- He explains why ReflectionAI shifted from an RL-on-top-of-open-models plan to building end-to-end pretrained foundations plus large-scale RL, citing coupling between pretraining and RL and a lack of top Western open bases.
- The episode details the resource realities of modern training—rapidly rising frontier costs, compute scarcity, and major efficiency gains from model-in-the-loop workflows—alongside concrete training scale numbers for Beam.
- Laskin argues open models can be commercialized by helping enterprises “own” intelligence via the surrounding deployment stack and solutions work, especially as enterprises move from renting closed tokens to on-prem/open deployments.
- The conversation contrasts open vs closed ecosystems across geopolitics and safety, claiming Chinese open models have helped the world but create strategic dependence, and asserting that openness can improve real-world security through broader auditing and defensive use.
IDEAS WORTH REMEMBERING
5 ideasFrontier model building is a tightly coupled engineering system, not one breakthrough.
Laskin argues model-building is a “30 things must go right” problem: talent, retention, compute procurement, infra reliability, data pipelines, evals, and training recipes all interact. Startups can’t rely on mature tooling the way big labs can, so simply making the cluster usable and repeatable is a major part of the work.
RL-only isn’t enough; pretraining and RL must be co-designed end-to-end.
ReflectionAI started with an RL-centric bet (make models agentic in math/coding), expecting to ride a strong open-model base. They shifted to full-stack pretraining + RL after realizing (1) top open bases were largely coming from China, and (2) RL at scale works best when you control the pretrained foundation it builds on.
The cost to reach the frontier rises fast, but efficiency gains partially offset CapEx.
He frames ‘catching up to frontier’ as more execution than exploration, which can be capital-efficient with top talent, but the moving frontier rapidly increases required spend. He estimates the frontier catch-up cost moved from hundreds of millions (18 months ago) to single-digit billions (recently) and plausibly toward ~$10B+ next year, tracking ~4× compute per generation.
Beam’s differentiation is “reasoning efficiency,” driven by unusually large-scale RL.
Beam is described as a 500B total parameter (23B active) model trained with ~6,000 GB300s for pretraining (weeks, trending toward ~12 days with infra efficiencies) and ~10,000 GB300s for ~4 weeks of RL—more FLOPs on RL than pretraining. Laskin claims large-scale RL yields agents that solve tasks faster and with less “meandering,” driving real customer cost/performance wins.
Open-weight businesses monetize the missing operational layer between weights and production.
Monetization is framed as enabling enterprises to “own” intelligence rather than “rent” it via tokens—analogous to renting an apartment vs owning a house. Open weights still require a full operational stack (inference software, cluster management, harnesses, services), and ReflectionAI positions itself as the partner that makes open deployment and scaling successful, especially as enterprises consider on-prem due to cost and compute scarcity.
WORDS WORTH SAVING
5 quotesEverything has been hard.
— Misha Laskin
You actually have to get 30 things right, and that's why everything is hard because you have to get 30 things right.
— Misha Laskin
So now you can— like, if you have the right question, you can just put it into, you know, a chat box and have it do, like, all the calculation and give you something really interesting.
— Misha Laskin
When you remove cyber offensive capabilities, you also remove cyber defensive capabilities.
— Misha Laskin
I have the belief that with enough eyeballs, most security and safety vulnerabilities become shallow as well.
— Misha Laskin
High quality AI-generated summary created from speaker-labeled transcript.