Skip to content
Dwarkesh PodcastDwarkesh Podcast

AI researchers debate how close we are to recursive self-improvement

New episode with John Schulman, Charlie O’Neill, and Beren Millidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. π„ππˆπ’πŽπƒπ„ π‹πˆππŠπ’ * Transcript: https://www.dwarkesh.com/p/john-beren-charlie * Apple Podcasts: https://podcasts.apple.com/us/podcast/ai-researchers-debate-how-close-we-are-to-recursive/id1516093381?i=1000789067132 * Spotify: https://open.spotify.com/episode/0ePd4PUqCpN78hCjVRH0fr?si=wGvk7u5XQwaLfyrvysdIJQ π’ππŽππ’πŽπ‘π’ * Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street's tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to https://antithesis.com/dwarkesh * Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at https://janestreet.com/dwarkesh * Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at https://x.ai/bot To sponsor a future episode, visit https://dwarkesh.com/advertise. π“πˆπŒπ„π’π“π€πŒππ’ 00:00:00 – Steelmanning the case against RSI 00:18:39 – What’s driving the Chinese labs’ progress 00:28:06 – How will automated AI researchers be trained 00:33:51 – Will long-horizon RL elicit AGI? 00:45:24 – The sim-to-real gap 01:00:33 – How much progress is explained by data? 01:18:03 – Why is RL working so well? 01:24:54 – Move 37 and entropy collapse 01:28:32 – Rapid-fire timelines

Dwarkesh PatelhostBeren MillidgeguestJohn Schulmanguest
Sep 11, 20261h 37mWatch on YouTube β†—

Episode Details

EPISODE INFO

Released
September 11, 2026
Duration
1h 37m
Channel
Dwarkesh Podcast
Watch on YouTube
β–Ά Open β†—

EPISODE DESCRIPTION

New episode with John Schulman, Charlie O’Neill, and Beren Millidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. π„ππˆπ’πŽπƒπ„ π‹πˆππŠπ’

π’ππŽππ’πŽπ‘π’

β€’ Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street's tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to https://antithesis.com/dwarkesh

β€’ Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at https://janestreet.com/dwarkesh

β€’ Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at https://x.ai/bot To sponsor a future episode, visit https://dwarkesh.com/advertise. π“πˆπŒπ„π’π“π€πŒππ’ 00:00:00 – Steelmanning the case against RSI 00:18:39 – What’s driving the Chinese labs’ progress 00:28:06 – How will automated AI researchers be trained 00:33:51 – Will long-horizon RL elicit AGI? 00:45:24 – The sim-to-real gap 01:00:33 – How much progress is explained by data? 01:18:03 – Why is RL working so well? 01:24:54 – Move 37 and entropy collapse 01:28:32 – Rapid-fire timelines

SPEAKERS

  • Dwarkesh Patel

    host

    Host of the Dwarkesh Podcast covering AI, economics, and science.

  • Beren Millidge

    guest

    AI researcher and CTO of Zyphra.

  • John Schulman

    guest

    AI researcher known for work on reinforcement learning and large language model training (formerly at OpenAI).

EPISODE SUMMARY

In this episode of Dwarkesh Podcast, featuring Dwarkesh Patel and Beren Millidge, AI researchers debate how close we are to recursive self-improvement explores recursive self-improvement hinges on objectives, realism, and continual learning The guests steelman why recursive self-improvement might not happen soon: models could keep winning benchmarks while remaining bottlenecked by weak judgment, verification, continual learning, and sim-to-real transfer.

RELATED EPISODES

Sarah Paine - Why Putin and Xi can't escape geography

Sarah Paine - Why Putin and Xi can't escape geography

Dylan Patel – Two labs will soon control most of the world's workforce

Dylan Patel – Two labs will soon control most of the world's workforce

Ajeya Cotra – "This might be the clearest warning shot we ever get"

Ajeya Cotra – "This might be the clearest warning shot we ever get"

Ryan Greenblatt – What happens once AI can automate AI research?

Ryan Greenblatt – What happens once AI can automate AI research?

General relativity from first principles – Adam Brown

General relativity from first principles – Adam Brown

Grant Sanderson (@3blue1brown) – AI disproved a famous math conjecture. Now what?

Grant Sanderson (@3blue1brown) – AI disproved a famous math conjecture. Now what?

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.