Y CombinatorWhy The Next AI Breakthroughs Will Be In Reasoning, Not Scaling
CHAPTERS
- 0:00 – 0:20
AGI feedback loops: AI-designed chips and removing compute bottlenecks
The conversation opens with the idea that sufficiently capable AI could start improving its own hardware constraints by designing better chips. The hosts frame o1’s progress in chip design as an early sign that this AGI-style feedback loop is becoming more plausible.
- •AGI speculation: AI designs chips better than humans
- •Chip design as a bottleneck to further AI capability
- •o1’s reported strength in chip design signals a new phase
- •Why this feels different than prior AGI talk
- 0:20 – 2:49
Sam Altman’s “intelligence age” timeline and techno-optimist vision
Garry and Jared discuss Sam Altman’s essay predicting AGI/ASI within “thousands of days,” with an explicit estimate of 4–15 years. They connect this optimism to a future of accelerated science, abundant energy, and major societal breakthroughs.
- •Sam’s 4–15 year AGI/ASI estimate
- •Techno-optimist outcomes: climate, energy, space, ‘intelligence on tap’
- •YC’s perspective from OpenAI’s early roots
- •Why the essay feels newly plausible in 2024
- 2:49 – 4:21
Why reasoning matters: o1 as the missing ingredient for doing real science
Jared argues OpenAI’s original motivation was using AGI to accelerate scientific discovery, and that requires better reasoning—not just fluent text. The group positions o1 as the step toward models that can think through complex problems and scientific workflows.
- •OpenAI’s early goal: accelerate science with AGI
- •Reasoning as the key limitation of earlier GPT models
- •o1 as a targeted leap in ‘thinking through’ problems
- •Reasoning capability as a precursor to engineering/science automation
- 4:21 – 4:51
YC x OpenAI o1 hackathon: real startups shipping real features
Diana describes a YC/OpenAI hackathon where startups used o1 to build production features, judged by Sam Altman. The point: o1 wasn’t just impressive in demos—it created immediate step-function upgrades for real businesses.
- •Hackathon format: funded startups building shippable product features
- •Sam Altman judging; direct exposure to o1 capabilities
- •Emphasis on practical, non-toy applications
- •Step-function improvements vs incremental gains
- 4:51 – 8:22
Demo: Diode Computer and end-to-end PCB design via component reasoning
Diana walks through Diode Computer’s workflow: using o1 to perform system design and component selection, then generating schematics and board layout as code. The chapter highlights why component choice and routing are hard—and why o1 changes what’s automatable.
- •PCB workflow stages: system design, schematics, routing (NP-complete)
- •o1 enables system design + component selection from datasheets
- •From high-level spec (wearable HR monitor) to PCB output
- •AutoPyL ‘electronics as code’ and auto-routing to a working board
- 8:22 – 10:25
A common product pattern: multiple models + RAG for extraction, o1 for reasoning
The hosts generalize from the Diode example: teams increasingly combine specialized models in a pipeline. Smaller models handle extraction/structuring, while o1 performs the heavy reasoning that was previously unreliable.
- •Using different models for different tasks in one workflow
- •RAG/structured data built from PDFs and unstructured docs
- •Small model for extraction → o1 for decision-making reasoning
- •Same prompts failing on GPT-4o but working on o1
- 10:25 – 11:52
Camfer ‘Devon for CAD’: natural language to CAD with physics-driven reasoning traces
Harj introduces Camfer, which generates CAD designs from natural language and can run simulations, effectively acting as a copilot to tools like SolidWorks. The group highlights o1’s ability to surface math/physics reasoning steps (e.g., PDEs) to solve engineering problems like airfoil optimization.
- •Natural language specs → CAD designs
- •Automation via UI interaction with existing CAD tools
- •Parallel simulations for engineering optimization tasks
- •o1 generating and solving equations/PDEs for fluid dynamics problems
- 11:52 – 14:26
Four orders of magnitude and compute-at-inference: why scaling isn’t the only lever
Garry returns to the idea of massive scaling—potentially four orders of magnitude—and connects it to solving harder engineering problems. The discussion reframes progress as not only bigger base models, but also more compute spent during inference to iteratively improve results.
- •Sam’s ambition: scaling spend toward trillion-dollar territory
- •Engineering problems (fusion, weather, physics) as ‘solvable’ with enough capability
- •Compute applied at inference enables iterative improvement
- •Analogy to human scientific organizations—more consistent, faster iteration
- 14:26 – 18:02
Architecture intuition for o1: RL roots (Dota), reward functions, and ‘teaching models to think’
Diana links o1’s approach to OpenAI’s reinforcement-learning heritage from Dota and AlphaGo-style techniques. The hosts speculate about reward functions and factually-grounded training data enabling better multi-step reasoning, plus the upcoming gap between o1-preview and the full o1 release.
- •OpenAI’s Dota era: RL competence before GPT fame
- •Bridging next-token prediction with RL-style training signals
- •Speculated data: math/science problems and factual corpora
- •Orthogonal research track to scaling: reasoning/RL; o1-preview vs full o1; o2/o3 coming
- 18:02 – 21:35
Evals as the moat: chain-of-thought replaces manual workflows, but testing becomes the advantage
Garry argues that as base models commoditize, competitive advantage shifts to proprietary evals and domain-specific test cases. The team connects this to prior advice (Jake Heller): break tasks into steps and build rigorous eval sets—while o1 may internalize the step breakdown, evals remain essential.
- •o1’s chain-of-thought may reduce need for manual step decomposition
- •Evals still critical to reliability and productization
- •Proprietary test cases as defensible advantage and ‘moat’
- •Distribution, UI, integrations, and switching costs remain classic moats
- 21:35 – 24:27
The ‘final 10–15%’ accuracy problem: who pays for perfection and why strong teams win
Harj suggests startups should target customers who will pay for near-perfect accuracy (e.g., aerospace) versus casual prototyping. The group argues AI doesn’t reduce the value of strong technical teams; instead, it raises the bar and rewards those who can capture the last-mile performance gains.
- •Segmenting customers by tolerance for error and value of accuracy
- •o1 makes reaching 80% easier; competitive edge is the last 10–15%
- •Strong technical teams capture disproportionate value
- •Product layer matters: UI, integrations, and workflow design
- 24:27 – 27:52
Case study pivoting in the o1 era: from fine-tuning services to AI customer support (GigaML)
The hosts tell a classic YC pivot story: a team moves from a niche non-AI idea to fine-tuning open-source models, then pivots again into AI customer support. They explain why fine-tuning-as-a-service got harder (cheaper/better base models), while customer support remains wide open due to low adoption and high edge-case complexity.
- •Pivot #1: non-AI idea → AI team; Pivot #2: fine-tuning → vertical application
- •Why fine-tuning businesses struggled as models improved and costs dropped
- •Customer support: huge edge-case space and trust gap vs rules-based systems
- •Market is still early despite many competitors—adoption hasn’t fully happened
- 27:52 – 31:50
o1 impact in production: Zepto scale, ticket automation, and an order-of-magnitude error reduction
They discuss GigaML’s real deployment metrics (tens of thousands of tickets/day) and the labor implications of automating high-turnover support roles. Crucially, o1 plus rigorous evals improved performance dramatically—from unusable error rates to near-production reliability during the hackathon.
- •Scale example: ~30,000 tickets/day; large support org replaced/augmented
- •Job impact framed as removing highly rote, high-turnover work
- •Before o1: ~70% error; with o1 + evals: ~5% error reported; big accuracy jump
- •o1-preview already delivers step change; full o1 expected to improve further
- 31:50 – 34:00
What gets disrupted next: coding agents, opacity of reasoning, and the next unlock (editable thoughts)
Diana asks which startups may be least helped—or even deprecated—by o1’s progress. The group flags coding-agent teams that built their own chain-of-thought tooling, then pivots to what’s missing: directable, interpretable reasoning where users can branch, rerun, and edit steps.
- •Potential pressure on AI coding-agent startups as o1 improves programming
- •Current limitation: opaque chain-of-thought and limited steering mid-flight
- •Need for interpretability/directability: ‘show steps,’ rerun/branch, edit plans
- •Speculation: future o2/o3 capabilities may unlock more controllable reasoning
- 34:00 – 35:16
Where new startups emerge: using o1 for the physical ‘atom world’ and an abundance agenda
The closing focuses on what o1 enables: startups in mechanical, electrical, chemical, bio, and other engineering fields where math/physics reasoning matters. Garry frames this as a race to deliver abundance and real-world improvements fast enough to outweigh public fear about AI.
- •Newly viable ideas: math/physics-heavy engineering domains
- •From ‘click faster’ software to real-world abundance creation
- •Technologists’ role in accelerating benefits to counterbalance fear
- •Wrap-up and sign-off