Skip to content
OpenAIOpenAI

Live from DevDay — the OpenAI Podcast Ep. 7

The OpenAI Podcast is live for the first time. Host Andrew Mayne sits down with startups Cursor, Abridge, SchoolAI, and Jam.dev—each reimagining how AI can transform their industries. From healthcare and education to coding and collaboration, we explore how these builders are putting AI to work in the real world. Subscribe to the OpenAI Podcast on Spotify and Apple Podcasts

Andrew Maynehost
Oct 6, 20251h 1mWatch on YouTube ↗

CHAPTERS

  1. 0:09 – 0:43

    DevDay kick-off: Building AI products with real-world stakes

    Andrew Mayne opens the live DevDay episode by framing the theme: developers using new OpenAI tooling to solve concrete problems. The first guest, Caleb Hicks of SchoolAI, sets the stage with education as a high-trust, high-impact domain.

    • Live recording from OpenAI DevDay
    • Focus on practical developer experiences, not just announcements
    • Introduction of SchoolAI and its classroom mission
  2. 0:43 – 1:06

    SchoolAI’s mission: Safe, managed AI tutors for students

    Caleb explains SchoolAI’s core approach—putting AI directly in students’ hands, but in a controlled, school-appropriate way. He emphasizes the role of orchestration and guardrails to make “one-time” personal tutoring safe and useful in classrooms.

    • Student-facing AI designed as managed, guardrailed tutoring
    • Orchestration of multiple agents/models for better outcomes
    • Safety and classroom constraints as first-class product requirements
  3. 1:06 – 2:35

    What changed in the last year: Model intelligence + cost improvements

    SchoolAI’s momentum is driven by two model-related shifts: noticeable gains in capability and meaningful reductions in cost. Caleb highlights why cost matters especially in education, where budgets are tight and usage can be massive.

    • Model intelligence leaps improve tutoring quality
    • Lower inference costs make classroom-scale deployment feasible
    • Education software pricing realities shape product strategy
  4. 2:35 – 4:02

    Educator adoption curve: From banning AI to teaching AI literacy

    Caleb outlines a common progression across schools: initial prohibition, then teacher productivity use, then recognition that students must learn AI to stay competitive. He argues the next step is deeply classroom-connected tutoring that understands what’s happening in class.

    • Early phase: widespread AI bans in schools
    • Middle phase: teacher productivity and admin acceptance
    • Emerging phase: AI literacy as a student necessity
    • Future: AI tutor connected to classroom goals and context
  5. 4:02 – 6:45

    SchoolAI product stack: Dot, teacher tools, and real-time student dashboards

    Caleb breaks SchoolAI into three layers: an assistant (“Dot”), structured teacher tools that generate artifacts (lesson plans, adapted content), and custom AI tutors teachers can deploy to students. The differentiator is visibility—teachers get live dashboards on student progress and needs.

    • Dot assistant: school-tuned GPT-style interface
    • Form-based teacher tools for common workflows (lesson plans, reading adaptations)
    • Custom, guardrailed tutors created lesson-by-lesson by teachers
    • Teacher dashboard surfaces student understanding and engagement signals
  6. 6:45 – 8:41

    Classroom reality and the “GPS for impact” idea

    Drawing on his teaching experience, Caleb describes the impossible tradeoffs teachers face when supporting many students. SchoolAI aims to guide attention—flagging which students need help right now—so teachers can intervene effectively across the middle majority, not just extremes.

    • Teacher workload scale: dozens of desks, many periods, hundreds of students
    • Tradeoff: focus on top performers, struggling students, or the middle 80%
    • AI as a prioritization system: identify who needs help today
    • Goal: actionable insights that teachers can immediately use
  7. 8:41 – 13:48

    DevDay tooling for education: Agent Builder, MCP, and evals at scale

    Caleb and Andrew connect DevDay announcements to real classroom constraints: permissions, file search, and non-technical users. They also stress evals—small error-rate improvements matter enormously when millions of students are using a system every day.

    • Agent Builder excitement: drag-and-drop workflows with robust permissions
    • Reducing custom infrastructure by adopting platform-native agent tooling
    • MCP servers as a standardized integration path for partners
    • Evals as essential when 2–3% error translates to huge real-world impact
  8. 13:48 – 15:32

    Jam.dev introduces “Please Fix”: editing websites like a doc, shipping PRs automatically

    Dani Grant describes a new product that lets non-engineers fix UI issues directly in the browser and submit clean pull requests. The goal is removing bottlenecks—small copy and design tweaks shouldn’t require tickets and engineering wait time.

    • Browser extension enables live editing of production-like UI
    • Non-engineers can make changes and generate a GitHub PR
    • PRs align with the team’s design system for engineer-friendly review
    • Targets the friction of tickets, prioritization, and minor UI fixes
  9. 15:32 – 18:35

    A new web paradigm: apps inside ChatGPT and “read, write, think”

    Dani reacts to DevDay as a shift in what browsing means—moving from static pages to agentic experiences. She links this to a future where teams iterate faster on “inside-ChatGPT” apps and where usability and polish become even more decisive.

    • Apps running inside ChatGPT reshape web interaction patterns
    • Usability becomes a competitive advantage as creation gets cheaper
    • Please Fix positioned to tweak UI even from ChatGPT-mediated experiences
    • Design details that don’t get prioritized today may become easier to ship
  10. 18:35 – 23:29

    Jam.dev’s product philosophy: optimize for user “wow,” not just features

    Dani explains Jam.dev’s iteration loop: constant user contact and a single north star—delivering an emotional “wow” in a painful workflow (bug fixing). She also shares the founder origin story from Cloudflare: the bottleneck was communicating bugs, not fixing them.

    • Primary metric: user delight and reduced pain in bug workflows
    • High-touch user feedback loop (co-founders/PMs contacting users)
    • Origin at Cloudflare: reproduction and communication were the bottleneck
    • “Please Fix” extends the thesis: eliminate engineering bottlenecks for small changes
  11. 23:29 – 25:59

    Future direction: self-improving agents via automated evals and prompt optimization

    Dani and Andrew discuss the desire for evals that write themselves using real company data, then automatically optimize prompts/agents. The broader advice: it’s an unusually fun moment to build, and many great startups begin as internal tools.

    • Engineers want evals generated from internal data automatically
    • Vision of agents improving themselves via eval-driven optimization
    • Founder advice: pick a problem you’ll love for a decade
    • Internal tools often become the most valuable products
  12. 25:59 – 31:36

    Abridge’s core value: ambient documentation that gives doctors time back

    Zach Lipton explains Abridge as an AI platform for doctor-patient conversations, reducing documentation burden created by electronic health records. Time savings show up both during the day and in reduced after-hours “pajama time,” improving patient focus and physician wellbeing.

    • Addresses EHR-driven clerical burden (paperwork vs patient care)
    • Automates visit documentation artifacts in the required clinical formats
    • Measured savings: up to ~an hour+ per day; large perceived burden reduction
    • Human impact stories: family time, reduced burnout, even relationship relief
  13. 31:36 – 44:07

    Healthcare-grade reliability: defining hallucinations, building eval pipelines, and earning trust

    Zach details how “hallucination” must be defined specifically for medical workflows—what’s unacceptable isn’t just falsity but unsupported content. He describes evaluation and remediation pipelines (including high recall targets) and emphasizes trust as ongoing delivery across product, security, and service.

    • Medical hallucination definition: “unlicensed” statements outside substantiating context
    • Models can often detect errors even if they sometimes generate them
    • Pipeline approach: sentence-level classification, remediation, and high-recall goals
    • Product expansion beyond scribing: point-of-care support across the full visit lifecycle
    • Trust is earned continuously via reliability, security, and enterprise delivery
  14. 44:07 – 46:22

    Cursor’s evolution: from autocomplete to autonomous coding agents (and why “vibes” matter)

    Lee Robinson describes how AI coding moved beyond simple completion into agents that refactor, self-correct, and use broader context. Cursor’s dogfooding culture blends formal evaluation with qualitative feel—sometimes “it didn’t feel right” is a key signal.

    • Shift from text autocomplete to autonomous, tool-using coding agents
    • Cursor team uses Cursor to build Cursor (continuous feedback loop)
    • Quality measured via both evals and day-to-day user experience
    • Models improving reduces need for overly long, brittle system prompts
  15. 46:22 – 49:49

    Shipping at Cursor: internal adoption as product-market fit + rapid model iteration

    Lee outlines how features incubate internally, gain adoption, then roll out through staged channels. He also explains Cursor’s “all of the above” model strategy: multiple foundation models plus custom models for autocomplete, updated frequently via reinforcement signals.

    • Anyone can ship features internally; adoption metrics determine viability
    • Staged rollout: internal → ambassadors/nightly → general release
    • Model evaluation in an IDE requires harnesses and real workflows
    • Custom autocomplete models with rapid online reinforcement learning updates
  16. 49:49 – 1:01:14

    New users, new workflows: agent-first interfaces, context practices, and the future of ‘vibe coding’

    Cursor is increasingly used by non-traditional developers, pushing the UI toward agent-centric modes that feel less like a classic IDE. Lee recommends different adoption paths for pros vs beginners and predicts “vibe coding” becomes normal prototyping—while deeper engineering still matters for production software.

    • Demographic broadening: PMs, designers, support teams, first-time builders
    • Agent-first UI reduces IDE overwhelm for newcomers
    • Best practices: pros ramp via tab completion → agent delegation; beginners start agent-first
    • Context quality (plans, agents.md) drives output quality
    • Future: more automation around bugs, on-call, and shipping—not just writing code

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.