No PriorsBeam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
At a glance
WHAT IT’S REALLY ABOUT
ReflectionAI’s Beam: scaling open-weight frontier models with reinforcement learning
- Misha Laskin explains why ReflectionAI shifted from betting on RL atop existing open models to building frontier open-weight models end-to-end, culminating in the release of Beam.
- He argues that training frontier models is a tightly coupled engineering systems problem—talent, infrastructure, data, compute, and evaluation all must be executed well, and the frontier’s cost is rising from hundreds of millions toward billions and beyond.
- Beam is positioned as an efficient coding/agentic model whose key advantage is faster reasoning per task, attributed to unusually large-scale reinforcement learning runs in addition to strong pre-training.
- ReflectionAI’s business thesis frames open weights as “owning intelligence,” requiring tooling, deployment support, and agentic systems around the model—especially for enterprises moving from closed-model spend to owned infrastructure.
- The conversation situates open models within geopolitics and safety debates, claiming Chinese openness has benefited the world but creates strategic dependence, while openness can also improve real-world cybersecurity and vulnerability discovery.
IDEAS WORTH REMEMBERING
5 ideasPre-training + RL are now inseparable if you want frontier-level agentic performance.
ReflectionAI found that scaling reinforcement learning depends tightly on having a strong pre-trained base, so they committed to training models end-to-end rather than relying on external open models (which were increasingly China-led). This coupling also shaped their “agentic shift,” where coding/agentic capability becomes a foundation that can be adapted quickly with the right evals and synthetic data.
The core challenge of building frontier models is systems integration, not a single constraint.
Laskin frames model building as a “30 things must go right” endeavor—talent, culture, data, compute, and the infra to make compute usable—many of which are invisible inside mature labs like DeepMind. The hardest part isn’t a single bottleneck; it’s orchestrating many interdependent systems under time pressure.
Frontier catch-up costs are exploding, but efficiency and economics may force an asymptote.
He estimates the capital to “catch up” to the moving frontier has risen rapidly: ~hundreds of millions (18 months ago), to single-digit billions (recently), to potentially ~$10B next year, driven by ~4× compute per generation. He also argues this won’t scale forever because CapEx must eventually align with revenue potential and because efficiency gains are accelerating.
Beam’s differentiation is not just capability, but speed-per-task driven by heavy RL.
Beam is described as a 500B-parameter MoE (23B active) optimized for coding and agentic tasks, with an emphasis on “reasoning efficiency” (achieving results faster/cheaper). They attribute much of this to very large-scale RL (10k GB300s for ~4 weeks), which trains the system to extract capability with fewer steps—analogous to AlphaGo becoming less “meandering” over time.
Open-model monetization is about packaging the missing enterprise stack around weights.
Commercially, he describes open weights as “ownership” versus closed APIs as “rental,” noting enterprises typically rent first until spend and strategic reliance justify owning. ReflectionAI positions itself as the layer that makes ownership workable—deployment tooling, cluster/inference management, agentic harnesses, and services to unlock high-value use cases that then drive inference demand.
WORDS WORTH SAVING
5 quotesWhen you remove cyber offensive capabilities, you also remove cyber defensive capabilities.
— Misha Laskin
The state of the world today is that we have a few hundred safety researchers within closed labs that understand how these things work, and despite their best intentions, it's impossible to cover the long tail of unintended consequences that these systems might have.
— Misha Laskin
Linus's Law, with enough eyeballs, all bugs become shallow. I have the belief that with enough eyeballs, most security and safety vulnerabilities become shallow as well.
— Misha Laskin
You actually have to get 30 things right, and that's why everything is hard, because you have to get 30 things right.
— Misha Laskin
Open models are Trojan horses for the infrastructure that they bring with them.
— Misha Laskin
High quality AI-generated summary created from speaker-labeled transcript.