No PriorsBeam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
CHAPTERS
- 0:00 – 0:42
Cold open: Why openness can improve AI security and safety
Misha argues that restricting models to reduce offensive cyber capabilities can also weaken defensive uses. He claims closed labs can’t cover the long tail of vulnerabilities, and that broad access—“enough eyeballs”—helps surface and fix safety issues faster.
- •Offense/defense are coupled in cyber; removing one harms the other
- •Closed labs have too few safety researchers to cover edge cases
- •Real-world example: open models used to remediate issues caused by closed models
- •Linus’s Law applied to AI safety/security vulnerabilities
- 0:42 – 2:06
Who Misha Laskin is and what ReflectionAI is building
Sarah introduces Misha and ReflectionAI’s mission: building frontier open-weight intelligence. Misha frames the last year as a sprint to assemble a full-stack frontier model lab and announces the release of Beam.
- •ReflectionAI’s goal: frontier open intelligence that’s widely accessible
- •Scaling the org from ~30 to ~300 people to build end-to-end models
- •Teams spanning pretraining, mid-training, RL, and infrastructure
- •Launch of Beam as Reflection’s first open model
- 2:06 – 3:55
What’s hardest about building an open frontier model lab
Asked what the single hardest constraint is—compute, talent, or data—Misha says it’s ‘everything.’ He emphasizes that success requires getting dozens of coupled details right, especially infrastructure that big labs have accumulated over years.
- •No single bottleneck; many interdependent systems must work
- •Hiring and retaining talent requires mission/culture coherence
- •Compute is only useful with robust tooling and operational stability
- •Big-lab ‘defaults’ (tools/processes) must be rebuilt from scratch
- 3:55 – 7:37
ReflectionAI’s shift: from RL-on-open-models to training models end-to-end
Misha explains the company’s original bet: use reinforcement learning to push chat models into agentic domains like math and coding, building on an external open base model. They later shifted to training their own base models because strong Western open models lagged and RL at scale required tight coupling with pretraining.
- •Early context: Gemini era was mostly chat; little agentic behavior
- •Founders’ RL lineage (AlphaGo-style thinking) shaped the thesis
- •RL progress accelerated faster than expected (industry + internal)
- •Need for a strong base model + pretraining/RL coupling drove end-to-end build
- 7:37 – 10:09
How much does it cost to catch the frontier? Capital and compute scaling curves
Misha gives order-of-magnitude estimates for “catching up” to the frontier and notes the bar rises rapidly each generation. He relates model generations to step-function compute increases and chip transitions (H100 → Blackwell → Vera Rubin).
- •18 months ago: hundreds of millions; now: single-digit billions; next: ~tens of billions (order estimates)
- •Heuristic: ~4× compute per model generation
- •Frontier scale framed as ~100k ‘frontier’ GPUs per generation
- •Economic pressure will eventually constrain runaway CapEx
- 10:09 – 12:22
Why scaling may still work: efficiency gains and model-in-the-loop acceleration
Even if CapEx has limits, Misha argues effective intelligence-per-FLOP is improving because models help build better models. He contrasts earlier steady efficiency gains with much larger headroom in RL, where better models can accelerate their own improvement loops.
- •Compute efficiency gains are measurable via loss/compute comparisons
- •Pretraining is more ‘hardened’; RL has larger remaining headroom
- •Claimed acceleration: model-assisted work outpaces human-only iteration
- •Economic framing: revenue-per-FLOP rises as intelligence density increases
- 12:22 – 16:51
Training Beam in numbers: architecture scale and RL-first investment
Misha shares concrete training scale for Beam: a large MoE-style footprint (500B total params, 23B active) and substantial GPU allocation. He highlights that RL consumed more compute than pretraining and describes the complexity of RL jobs (inference + sandboxes + training).
- •Beam: ~500B total parameters, ~23B active
- •Pretraining: ~6,000 GB300s for weeks (now ~12 days with infra efficiency)
- •RL: ~10,000+ GB300s for ~4 weeks; more FLOPs than pretraining
- •RL workloads combine massive inference, environments, and training loops
- 16:51 – 19:55
Where model value comes from: agentic capability, evaluations, and synthetic data
The discussion turns to where economic value appears: code and agentic workflows as a foundation, then enterprise-specific tasks. Misha argues models are ‘jagged’ across benchmarks, but can adapt quickly when you set up the right evaluations and generate synthetic data aligned to real workloads.
- •Code + agents are foundational; specialization follows data/evals
- •‘Jagged’ generalization across harnesses/benchmarks (e.g., Terminal Bench variants)
- •Value comes from finding economically valuable evaluation loops
- •Enterprise verticals mentioned: finance/KYC/compliance, cyber defense, legal
- 19:55 – 22:36
Beam’s reasoning efficiency: faster solutions via large-scale RL
Sarah probes why Beam is unusually reasoning-efficient—solving tasks in fewer steps/tokens. Misha attributes this to a strong reasoning pretraining base amplified by large-scale RL, which tends to produce more direct, efficient search and decision-making.
- •Efficiency matters: faster task completion lowers cost and latency
- •Claimed 3–4× efficiency vs similar-class models; up to ~10× vs larger models
- •RL training encourages extracting capability in minimal steps
- •AlphaGo analogy: RL leads from meandering play to efficient search
- 22:36 – 25:05
Monetizing open-weight models: ‘renting’ tokens vs ‘owning’ intelligence
Misha explains commercialization as maximizing inference demand, with open vs closed as an ownership model distinction. Open weights shift work to customers (deployment stack), so Reflection focuses on packaging the missing layers—infra, inference software, cluster management, and services to unlock high-value use cases.
- •Closed tokens as ‘renting’ the full stack; open weights enable ‘ownership’
- •Enterprises want control as spend scales and use cases mature
- •Open deployment still needs expensive surrounding software/ops layers
- •Services act as a demand driver to unlock durable inference consumption
- 25:05 – 31:31
Open vs closed token mix: why open usage is rising and what customization looks like
Elad asks where open vs closed tokens land over 1–2 years. Misha predicts open will dominate token demand (Linux analogy), while closed players remain highly valuable; he also distinguishes fine-tuned models from ‘customized systems’ in enterprise adoption paths.
- •Observed shift at gateways: from ~70/30 closed/open to ~70/30 open/closed
- •Linux analogy: open dominates infrastructure, closed can still capture huge value
- •Today: many dedicated inference workloads are fine-tuned (AI-native heavy)
- •Enterprise likely: more ‘customized systems’ (harnesses/workflows) before fine-tunes
- 31:31 – 34:38
Competing in open models: intelligence density × compute × trust
Competition is intense across the stack, and margins compress as open options strengthen buyers’ negotiating power. Misha proposes a simple durable-revenue formula: deliver high intelligence density, secure scarce compute, and earn trust by solving real customer problems.
- •Open ecosystems increase competition and compress margins across layers
- •Buyers negotiate harder with closed providers when open alternatives exist
- •Winning formula: intelligence density, compute access, and customer trust
- •Compute scarcity creates strategic advantage for high-revenue builders
- 34:38 – 44:44
China’s open-weight ecosystem: benefits, distillation, and why openness might persist
Sarah asks whether Chinese open models will stay open and what incentives drive them. Misha credits Chinese releases as a global benefit, argues multipolar open ecosystems are healthier, and suggests openness can be geopolitically advantageous by anchoring other countries to a broader hardware/software stack.
- •Chinese open models have enabled many Western businesses to exist/compete
- •Advantages cited: large-scale distillation, cheaper/looser data regime, state support dynamics
- •Hypothesis: China can monetize domestically while releasing openly abroad
- •‘Open models as Trojan horses’ for infrastructure lock-in (chips, software stacks)
- 44:44 – 57:17
Safety and open models: realism vs doomsday, and ‘patching bugs’ at scale
Misha calls safety concerns legitimate but critiques a dogmatic, doomsday-peaked discourse. He argues practical alignment today looks like mundane vulnerability discovery and patching, and that openness could scale the number of researchers finding issues—similar to how open cryptography strengthened security.
- •Safety spans empirical near-term risks to speculative long-tail scenarios
- •Alignment often resembles whack-a-mole patching with data + detection tools
- •Historical analogy: open cryptography and the rise of cybersecurity
- •Claim: more independent researchers (‘eyeballs’) improves robustness and defense
- 57:17 – 1:02:06
Looking ahead: AI accelerating science, and data centers as modern factories
Misha shares excitement about scientific acceleration: models went from useless to undergraduate-level to PhD-level help on his thesis-style prompts, increasing iteration speed dramatically. He also notes data centers’ economic impact, describing them as job- and tax-revenue engines akin to 20th-century factories.
- •Scientific workflows: faster iteration and higher-quality reasoning assistance
- •From ‘PhD takes years’ to ‘PhD a week’ framing (if the question is right)
- •Real-world lab science (bio/materials/chemistry) enables proprietary data flywheels
- •Data centers create substantial local employment and economic activity
- 1:02:06 – 1:10:48
How ReflectionAI directs research after Beam: scaling bets, model-in-the-loop, and headcount vs compute
The conversation closes on how a lab chooses what to do next: execute known high-impact improvements vs riskier bets, and when scale is required to see emergent effects. Misha explains how models help researchers move faster, but argues labs still need a stable researcher-to-compute ratio and will remain talent-hungry, especially for forward-deployed applied work.
- •Roadmap approach: exhaust known wins, then expand riskier exploration
- •Some behaviors only appear at scale; balancing small-scale tests vs big runs is an art
- •Model-in-the-loop boosts researcher throughput (hyperparams, infra, iteration)
- •Headcount likely stabilizes around O(100) for core work; big growth in applied/forward-deployed roles