a16zBuilding an AI Physicist: ChatGPT Co-Creator’s Next Venture
CHAPTERS
- 0:00 – 1:06
From ChatGPT to an “AI physicist”: why experiments must be in the loop
The conversation opens with the core thesis behind Periodic Labs: real scientific progress requires grounding AI optimization in physical experiments, not just digital reward functions. The guests frame the opportunity as building systems that can design and verify changes in the real world across materials, chemistry, and manufacturing.
- •Science ultimately advances through real-world experimental verification, not text-only reasoning
- •Periodic’s bet: couple LLMs + simulations + automated experiments into a closed loop
- •An “AI physicist” could impact any R&D process tied to the physical world
- •Superconductivity is introduced as an example of a measurable, high-impact target
- 1:06 – 3:57
Origin story: meeting at Google Brain and converging on physics + LLMs
Fedus and Cubuk recount how they met years earlier at Google Brain (via an amusing tire-flipping anecdote) and kept reconnecting through shared interest in quantum mechanics and superconductivity. As LLMs improved, they began asking whether these models could become first-class tools for physics research rather than just assistants for recall and coding.
- •Longstanding shared interest in quantum mechanics and superconductivity
- •LLMs became useful first for memory refresh and coding/simulation scripts
- •Key shift: using LLMs as primary actors in physics research workflows
- •Frontier ML progress (reasoning, RL) made the timing feel right to found Periodic
- 3:57 – 5:53
What Periodic Labs is building: a frontier lab for physics/chemistry with real rewards
Periodic Labs is described as a frontier AI research lab aiming to accelerate physics and chemistry by tightly integrating experiments, simulations, and LLM-driven iteration. The defining idea is to replace brittle, digital reward signals with physically grounded rewards where experiment is the ultimate ground truth.
- •High-throughput, high-quality experimental data generation is central
- •Simulations and LLMs are used together, but experiments provide the corrective truth
- •Goal: a physically grounded reward function analogous to unit tests in code
- •Nature/experiment becomes the RL environment—harder to ‘hack’ than preference rewards
- 5:53 – 8:40
How today’s LLM training differs: from RLHF preference rewards to physical verification
Fedus contrasts early ChatGPT’s RLHF pipeline with what Periodic needs for science. The chapter explains how reward functions evolved from “helpful assistant” preferences toward more verifiable correctness, and why scientific discovery demands experimental feedback rather than internet-only training distributions.
- •ChatGPT origins: supervised fine-tuning + RLHF on human preference comparisons
- •Early reward functions optimized friendliness/helpfulness, not mathematical correctness
- •Better reward design improved correctness, but remained fully digital
- •Periodic extends the toolset to physics simulators and lab automation, with experiment as truth
- 8:40 – 10:12
Why a real lab matters: quantum-scale focus and automated powder synthesis
Cubuk explains that after logic and math, physics—especially quantum mechanics at practical energy scales—is a natural next frontier for verification-driven AI. Periodic’s first lab emphasis includes solid-state physics/materials/chemistry and a powder synthesis setup amenable to low-cost robotics and high-throughput iteration.
- •Targeting quantum mechanical regimes relevant to materials, chemistry, biology
- •Solid-state physics/materials/chemistry chosen as a practical, high-impact focus
- •Powder synthesis highlighted as a foundational, automatable discovery method
- •Robotic automation can enable fast iteration for superconductors, magnets, and more
- 10:12 – 12:36
Why existing models can’t ‘do science’ yet: iteration, noisy literature, and missing negatives
The guests argue current models struggle because science requires iterative hypothesis-testing with action, not one-shot answers. They highlight limitations of published datasets: noise, wide variance in reported properties, and publication bias against negative results—making text-trained models reproduce uncertainty rather than reduce it.
- •Scientific progress requires iterative loops of sim → theory → experiment → revision
- •Models aren’t trained to act and learn from real-world feedback
- •Literature labels can vary by orders of magnitude, limiting learnability
- •Negative results are underreported but crucial as learning signals
- 12:36 – 17:47
Measuring success: superconducting temperature and ‘can you design the world?’
Progress is framed in concrete, hard-to-game metrics: discover materials with better measured properties. Superconductivity offers a simple scalar benchmark (critical temperature), while industrial relevance is framed as whether the system can deliver requested material/device properties reliably.
- •Superconductivity benchmark: push beyond today’s ambient-pressure ~135K Tc
- •Measure material properties directly (toughness, ductility, strength, etc.)
- •Physical measurements provide robust evaluation signals versus synthetic rewards
- •Ultimate test: can the system design and produce materials that meet specs?
- 17:47 – 23:12
Why scaling alone won’t crack physics: domain shift, missing data, and the wrong Y-axis
Fedus and Cubuk affirm belief in scaling laws but argue the evaluation axis and target distribution matter. They explain how in-domain gains may not translate to far out-of-domain physics tasks, and that some crucial experimental data simply doesn’t exist or is too noisy to learn from—necessitating new data generation via experiments.
- •Scaling laws hold, but the test distribution defines what improvements mean
- •In-domain vs out-of-domain generalization can have very different slopes
- •A code-optimized model won’t automatically discover cures or new materials
- •Key experimental datasets (e.g., superconductivity/synthesis) are noisy or incomplete; labs must create data
- 23:12 – 27:32
Why superconductivity (and magnetism) first: a North Star with many sub-goals
Superconductivity is presented as a motivating North Star that forces the team to solve many prerequisite capabilities—automation, characterization, simulation correctness, and iterative discovery workflows. It’s also philosophically compelling and technically attractive because phase transitions can be more robust to hard-to-model details than many other properties.
- •Superconductivity goal implies building autonomous synthesis + characterization pipelines
- •Discovery would be scientifically transformative even before commercialization
- •Phase transition properties can be more robust than defect-sensitive properties
- •Strategy: ‘make contact’ with a full closed loop in one domain, then generalize to others
- 27:32 – 28:49
Commercial path: ‘copilots’ for advanced-industry engineers and researchers
As a startup, Periodic pairs the long-term AI scientist ambition with nearer-term product value: copilots for R&D teams in space, defense, semiconductors, and advanced manufacturing. The opportunity is to reduce iteration time in materials/device development where current tooling and data are limited despite large budgets.
- •Medium-term product: copilots that accelerate physical R&D workflows
- •Target customers: industries with massive R&D budgets and physical iteration loops
- •Thesis: commercial success enables maximal scientific acceleration
- •Positioning as an intelligence layer that shortens experimentation and design cycles
- 28:49 – 32:49
Integrating ML and lab scientists: shared intuition, teaching sessions, and ‘API thinking’
The team-building focus shifts to cultural and communication bridges between ML researchers and experimental/simulation experts. They describe weekly cross-teaching, encouraging “no stupid questions,” and translating scientific problems into ML-evaluable inputs/outputs—supported by hybrid ‘bridge’ people spanning disciplines.
- •Physicists/chemists help design training steps that teach correct scientific reasoning
- •ML researchers learn domain goals, simulators, and experimental constraints
- •Weekly teaching sessions build shared vocabulary and intuition
- •Computer-science ‘API framing’ helps map science workflows into trainable/evaluable tasks
- 32:49 – 35:38
What they hire for: not degrees, but curiosity, rigor, mission, and urgency
Periodic explicitly does not require advanced physics degrees, emphasizing instead that modern science is too broad for any one person to know everything. They look for mission alignment, deep curiosity, pragmatic execution, world-class strength in a pillar (ML/experiment/simulation), and a strong sense of urgency to deliver impact soon.
- •Advanced degrees in physics/chemistry are not required
- •Modern discovery requires collaboration across many specialized subfields
- •Hiring bar: mission-driven, deeply curious, pragmatic, and world-class in some dimension
- •Emphasis on urgency: deliver scientific acceleration ASAP, not in a decade
- 35:38 – 40:58
Deployment into traditional industries: scoped wins, evaluations, and beyond retrieval
The discussion turns to deploying frontier AI into slower-moving, mission-critical industries. Periodic’s approach is ‘land and expand’ with tightly scoped, high-value problems and clear evaluation criteria, while addressing a key technical hurdle: moving beyond retrieval toward training/mid-training without violating internal access controls.
- •Many companies are actively seeking an AI strategy and ways to preserve expertise
- •Deployment plan: start with a well-scoped critical problem and measurable evals
- •Key pain points: automating simulations and integrating data into design pipelines
- •Beyond retrieval: encode knowledge via training/mid-training while managing permissioned data
- 40:58 – 45:19
Mid-training explained: injecting new physics/chemistry knowledge and connecting distributions
Fedus defines mid-training as continuing pre-training on new, domain-specific knowledge not present in the base model—distinct from post-training focused on preferences or behaviors. In Periodic’s case, this includes experimental and simulation data (from crystal structures to synthesis procedures) and the challenge of ensuring these added datasets improve generalization rather than remain isolated islands.
- •Mid-training = continue pre-training on newly acquired data/knowledge
- •Different from post-training (RL/SFT) focused on behavior shaping
- •Periodic mid-trains on experimental + simulation data that often doesn’t exist yet
- •Key ML challenge: connect distributions so added data improves performance across tasks
- 45:19 – 51:48
Partnering with academia and specialized tools: advisory board, grants, and geometric reasoning
The guests emphasize strong industry–academia synergy: academia produces foundational simulation tools and scientific thinking strategies, while industry scales computation and systems. Periodic plans an advisory board spanning superconductivity and synthesis, plus a grant program; they also discuss augmenting LLMs with geometric/neural tools (e.g., equivariant GNNs, diffusion models) for atomic and structural reasoning.
- •Academia contributes critical simulation tooling and scientific reasoning frameworks
- •Advisory board includes leading experimental/theory/synthesis experts
- •Grant program to fund academic work aligned with LLM/agent-driven discovery
- •Hybrid toolchains: LLMs + geometric models for atomic structures and physical reasoning