No PriorsBeam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
CHAPTERS
- 0:00 – 0:42
Why open models can improve security: “enough eyeballs” for safety
Misha argues that restricting model capabilities can also cripple defenses, especially in cybersecurity. He frames open weights as a practical safety strategy: broader scrutiny surfaces vulnerabilities faster than a small internal safety team can.
- •Open vs. closed models as a security tradeoff: defense suffers when offense is removed
- •Closed labs can’t cover the long tail of unintended consequences with limited safety headcount
- •Real-world example: open models used to remediate harm caused by a closed model
- •Linus’s Law applied to AI safety/security vulnerabilities
- 0:42 – 2:06
ReflectionAI’s mission and sprint to build a frontier open-model lab
Elad introduces Misha and ReflectionAI’s goal: “frontier open intelligence.” Misha describes the rapid build-out from a small startup into a full-stack training organization and the release of their first open model, Beam.
- •Mission: build frontier open intelligence and make it widely accessible
- •Scaling the org: ~30 people to ~300 in about a year
- •Standing up teams across pretraining, mid-training, RL, infra
- •First end-to-end model release: Beam
- 2:06 – 3:55
What’s hardest about building open frontier models: everything is coupled
Asked about challenges (compute, talent, data), Misha emphasizes that success requires getting dozens of interdependent details right. He contrasts the stability of mature labs with the need for startups to build foundational tooling from scratch.
- •No single bottleneck: must get ‘30 things right’ simultaneously
- •Talent acquisition/retention depends on mission and culture
- •Compute procurement is only useful with robust infrastructure
- •Big labs’ “taken-for-granted” tools must be rebuilt in a new lab
- 3:55 – 7:37
From RL-first startup to building models end-to-end: the agentic shift
Misha explains ReflectionAI’s original bet: apply reinforcement learning to make models agentic in math/coding, leveraging an open base model ecosystem. They shifted to full pretraining after realizing top open bases were largely Chinese and that RL at scale needs tight coupling with pretraining.
- •Early context: chat-centric era with most compute in pretraining
- •Founders’ RL lineage (Gemini RL; AlphaGo experience) shaped strategy
- •RL progressed faster than expected across industry (e.g., o1 era)
- •Need for a Western open base model + pretraining/RL coupling drove the pivot
- 7:37 – 10:42
The economics of catching up to the frontier: hundreds of millions → billions → tens of billions
The conversation turns to capital requirements and how frontier costs scale over time. Misha offers rough orders of magnitude and a compute-scaling heuristic (generation-to-generation ~4×) tied to advancing GPU generations and cluster sizes.
- •Catching up used to be ‘hundreds of millions’; now ‘single-digit billions’
- •Next year’s frontier could imply ‘tens of billions’ in spend
- •Heuristic: ~4× compute per model generation (H100 → Blackwell → Rubin)
- •Frontier spending may asymptote due to CapEx/revenue constraints
- 10:42 – 12:22
Scaling efficiency and models improving themselves: faster progress per FLOP
Misha argues compute efficiency gains are accelerating because models increasingly help build and improve models. He contrasts historical pretraining efficiency (measurable loss-based gains) with large headroom in RL where automated improvement loops can be much faster.
- •Efficiency gains are concrete: reaching the same loss with fewer FLOPs
- •Pretraining is more mature; RL has substantial remaining headroom
- •Claimed acceleration: models in the loop can yield ~4–5× faster progress vs humans
- •Intelligence per training FLOP and revenue per FLOP both increase over time
- 12:22 – 16:51
How Beam was trained: compute, parameters, and RL dominating the budget
Misha shares Beam’s training profile and highlights that reinforcement learning consumed more compute than pretraining. He frames RL as producing systems that keep improving (AlphaGo-style), turning capability improvements into an economic decision: how much compute to invest for gains.
- •Beam described as 500B total parameters, 23B active (MoE-style)
- •Pretraining: ~6,000 GB300s for weeks (now ~12 days with infra efficiency)
- •RL: ~10,000+ GB300s for ~4 weeks; includes large-scale inference + sandboxes
- •RL curves ‘never stop learning’; improvements become compute/economics-limited
- 16:51 – 19:55
Where model value comes from: adaptive intelligence plus high-value enterprise data loops
Misha discusses ‘jagged’ generalization: models adapt quickly with the right evaluations and synthetic data, but don’t always transfer cleanly across harnesses. He argues value comes from locating economically valuable task pools and building systems/evals that enable rapid adaptation.
- •Strong agentic baseline enables faster adaptation to new task harnesses
- •Generalization can be jagged across benchmarks/harness implementations
- •Key lever: build evaluations, then generate synthetic data approximating tasks
- •High-value verticals: finance (KYC/compliance), cyber defense, legal
- 19:55 – 22:36
Beam’s reasoning efficiency: faster agents via large-scale RL
Beam is positioned as a workhorse for coding and agentic tasks with notably lower time-to-solve than peers. Misha attributes this to a strong reasoning-oriented pretraining base amplified by what he claims is unprecedented open-source-scale RL, yielding faster, cheaper workloads.
- •Efficiency matters alongside raw capability (latency/cost for agent workflows)
- •Claimed 3–4× efficiency vs similar-class models; up to ~10× vs larger ones
- •Cause: prioritize reasoning pretraining + massive-scale RL amplification
- •RL encourages solving tasks in fewer steps (AlphaGo analogy: less ‘meandering’)
- 22:36 – 25:05
Monetizing open weights: from ‘renting tokens’ to ‘owning intelligence’
Misha compares closed-model token purchases to renting an apartment: you rent an integrated stack. Open weights enable ownership, but require the surrounding deployment stack (inference, cluster management, harnesses) and often services to drive adoption and compute demand.
- •Closed models sell ‘rented’ inference across an integrated stack
- •Open models enable ownership/control but require more operational tooling
- •ReflectionAI’s role: provide deployment tooling + services for enterprises/sovereigns
- •Services act as a demand driver that unlocks high-volume inference workloads
- 25:05 – 31:31
Open vs closed token mix and customization: fine-tuned models vs customized systems
The group predicts open tokens will dominate over time, akin to Linux in servers, while closed providers remain highly valuable. Misha argues most current dedicated inference workloads are customized/fine-tuned (driven by AI natives), but enterprises will more often customize systems/harnesses before fine-tuning models.
- •Observed shift at gateways: from ~70/30 closed/open to ~70/30 open/closed
- •Analogy: open OS dominance (Linux) with valuable closed ecosystems (Microsoft/Apple)
- •Most dedicated inference today appears customized because AI-native customers dominate
- •Enterprise path: closed → open → system customization; fine-tuning comes later
- 31:31 – 34:38
Competing in a commoditizing stack: intelligence density × compute × trust
Misha proposes a simple competitive framework and emphasizes compute scarcity as a differentiator. He expects hypercompetition and margin compression across models, apps, inference layers, and even bare metal, making durable revenue dependent on capability, compute access, and customer trust.
- •Winning formula: intelligence density × compute × trust/solutions delivery
- •Compute scarcity shapes competition and can be strategic for leading players
- •Open models increase pricing pressure and compress margins across the stack
- •Trust comes from consistently solving customer problems, not just delivering a model
- 34:38 – 44:44
China’s open model ecosystem: benefits, incentives, and geopolitical ‘Trojan horses’
Misha views Chinese open models as a global benefit but argues the West needs competitive alternatives. He outlines China’s advantages (distillation, data permissiveness, support) and hypothesizes why openness may persist: captive domestic monetization and geopolitical leverage via infrastructure ecosystems.
- •Chinese open models enabled many Western businesses; lack of openness would concentrate power
- •Advantages cited: large-scale distillation, cheaper/copyright-permissive data, support runway
- •Openness can drive ecosystem lock-in (models as ‘Trojan horses’ for infra stacks)
- •Geopolitical parallel: exporting foundational infrastructure creates long-term leverage
- 44:44 – 56:12
Safety, openness, and access to powerful tools: cyber defense vs offense and ‘boring’ alignment
Misha argues safety debates are often dominated by doomsday framing, while near-term risks are more empirical (cyber, misuse, misalignment bugs). He claims openness is the historical default for secure software (encryption precedent) and that alignment in practice is iterative ‘patching,’ best done with broader participation.
- •Safety is multifaceted: cyber, bio/terror, existential; time horizons differ
- •Encryption history: closed designs failed; open standards enabled modern cybersecurity
- •Alignment described as Whac-A-Mole patching with data + detection tools (often LMs)
- •Restricting capabilities can harm defenders; hard to separate cyber offense/defense
- 56:12 – 1:01:02
AI’s upside: accelerating science, real-world experimentation, and data-center job creation
Misha highlights scientific acceleration as the most exciting near-term benefit, citing rapid improvements in models’ ability to solve PhD-level problems. He also points to life sciences and materials experiments and argues data centers can act like modern factories, creating significant local employment and tax bases.
- •Models evolved from chat to undergraduate to PhD-level solutions on technical work
- •Iteration speed could compress research cycles dramatically (“PhD a week” metaphor)
- •Big potential in lab-in-the-loop science: biology, chemistry, materials
- •Data centers as ‘new factories’: construction/ops jobs and local economic impact
- 1:01:02 – 1:09:44
How ReflectionAI directs research after Beam: model-in-the-loop, scaling bets, and headcount-to-compute ratios
Misha explains prioritizing known high-impact improvements before riskier explorations, while acknowledging some phenomena only appear at scale. He describes using models as fast collaborators for research and infrastructure work, and discusses why teams may shrink per unit effort but expand ambition overall.
- •Roadmap driven by known wins first, then expanding risk profile with scaling experiments
- •Some breakthroughs only emerge at large scale, complicating small-scale testing
- •Models as ‘fast, eager colleagues’ that accelerate researcher iteration
- •Staffing constraints remain: maintain a headcount-to-compute ratio; growth shifts to applied deployment roles