Lenny's PodcastOpenAI's Sherwin Wu: How Codex reviews 100% of OpenAI's PRs
How OpenAI uses Codex on every code review, shrinking review time sharply; engineers now manage fleets of AI agents inside the company at scale.
CHAPTERS
- 3:21 – 6:54
How OpenAI engineers use Codex daily: AI-written code and universal PR review
Sherwin shares how deeply Codex is embedded in OpenAI’s engineering workflow, including near-universal daily usage and automated PR review. They discuss how AI changes throughput and why measuring “AI-written code” precisely is tricky, even if AI is the first author for most changes.
- •95% of engineers use Codex daily; 100% of PRs get a Codex review
- •Most code is generated by AI first, then humans steer and review
- •Codex-heavy engineers open ~70% more PRs, and the gap is widening
- •Trust in the model increases as capability improves (“worst the models will ever be”)
- 6:54 – 12:27
Engineers as “wizard tech leads”: managing fleets of agents and parallel work
The conversation shifts to how the software engineer role is evolving from writing code to orchestrating many concurrent agent threads. Sherwin uses the classic SICP “wizardry” metaphor to explain why prompting/steering is becoming the core skill.
- •Engineers increasingly behave like managers/tech leads of agent fleets
- •High parallelism: 10–20 threads of work at once, with periodic steering
- •SICP’s ‘incantations’ metaphor becomes literal with natural language coding
- •The Sorcerer’s Apprentice risk: power without control if you disengage too much
- 12:27 – 15:08
The new stress of agent management: when there’s no escape hatch
Lenny asks about the anxiety of running many agents and dealing with failures. Sherwin describes an internal experiment maintaining a 100% Codex-written codebase, revealing what breaks when humans can’t simply ‘take over’ and rewrite manually.
- •Agent stress comes from monitoring, debugging, and time lost when agents stall
- •OpenAI experiment: a fully Codex-written codebase with no manual fallback
- •Many failures trace back to missing/underspecified context rather than model ability
- •Best practice: encode tribal knowledge into repos via docs, structure, and guidance files
- 15:08 – 19:28
Scaling code review and CI with Codex: from bottleneck to automation
They dive into the practical implications of PR volume and review burden. Sherwin explains how Codex-assisted review and CI automation compress the time and friction between writing code and shipping it.
- •Codex reviews every PR; humans focus attention where risk is higher
- •Review time can drop from ~10–15 minutes to ~2–3 minutes for many PRs
- •Small PRs may need minimal human review when Codex provides strong coverage
- •Automation extends into linting/CI retries and deployment steps to reduce friction
- 19:28 – 31:39
Engineering managers in the AI era: leverage, bigger teams, and top-performer focus
Sherwin explains how AI changes management less directly than IC work but amplifies differences in individual productivity. He predicts managers will gain leverage through organizational-context tools and may manage larger teams than today’s norms.
- •AI ‘supercharges’ high-agency/top performers, widening productivity spread
- •Management philosophy: spend >50% of time with top ~10% performers to unblock/retain them
- •Org-context AI (GitHub/Notion/Docs) can aid reviews, summaries, and performance cycles
- •Managers may handle larger orgs as tools improve situational awareness and reporting
- 31:39 – 31:40
Core management philosophy: the “surgeon” model and looking around corners
Sherwin shares enduring lessons from classic management thinking—especially enabling ICs as if they are ‘surgeons’ supported by a team. They also riff on using AI to proactively detect blockers and anticipate future dependencies.
- •Goal: make engineers feel fully supported—tools, clarity, and obstacles removed
- •Managers add leverage by anticipating org/process blockers before they hit execution
- •Idea: use AI connected to internal systems to surface current blockers automatically
- •Extend to forecasting: ask AI to predict next month’s likely cross-team bottlenecks
- 31:40 – 37:39
The one-person billion-dollar startup—and the B2B SaaS explosion around it
Sherwin argues people underprice the downstream effects of extreme leverage. Even if one-person unicorns are rare, the enabling ecosystem could produce a boom in small, specialized software businesses and reshape venture economics.
- •One-person billion-dollar startup as a symbol of AI-driven leverage
- •Second-order effect: far more startups and verticalized bespoke software
- •Potential “golden age of B2B SaaS” with many $10M–$100M companies
- •Third-order effect: VC dynamics may shift if outcomes skew smaller but more numerous
- 37:39 – 41:56
Why many AI deployments have negative ROI: adoption realities outside the tech bubble
Moving from internal OpenAI practices to customer deployments, Sherwin explains why AI initiatives often disappoint. The biggest gap is organizational: many users lack fluency, and AI is mandated top-down without grassroots learning and iteration.
- •Silicon Valley power-user norms don’t match typical enterprise readiness
- •Negative ROI can come from low utilization, misuse, or misfit to workflows
- •Successful rollouts need both exec sponsorship and bottom-up adoption
- •Recommendation: create a tiger team of internal evangelists to explore and teach
- 41:56 – 43:57
Who belongs on the AI tiger team: technical-adjacent operators as champions
Sherwin details the kinds of employees most likely to drive adoption in non-tech-heavy organizations. Often it’s not engineers, but highly capable, systems-minded operators who love tools and can translate AI into day-to-day workflows.
- •Many companies don’t have many software engineers; champions come from elsewhere
- •Strong candidates: technical-adjacent ops leads, support leaders, ‘Excel wizards’
- •Tiger team role: experiment, document best practices, run trainings/hackathons
- •Frame: find AI ‘top performers’ and empower them to seed adoption across teams
- 43:57 – 49:07
Customer feedback vs model velocity: “the models will eat your scaffolding for breakfast”
Sherwin explains why blindly following customer requests can mislead AI builders. As models improve quickly, yesterday’s essential scaffolding (agent frameworks, vector stores, files-based prompting conventions) can become obsolete.
- •Models ‘self-disrupt’ the surrounding tooling as capability increases
- •Past waves: heavy agent frameworks and vector-store centric designs lost primacy
- •Advice: balance customer asks with a forward view of where models will be in 12–24 months
- •“Bitter lesson” applied to product: over-engineered logic often gets washed away by better models
- 49:07 – 53:35
Build for where models are going—and what’s next: longer tasks and better audio
Sherwin shares what he’s watching over the next 12–18 months: the length of coherent autonomous work and major improvements in multimodal audio. These shifts will change what products look like and which domains unlock next.
- •Key frontier: longer coherent task execution (multi-hour trending upward)
- •Product implication: designing feedback/monitoring loops for long-running agents
- •Audio is underrated: business is conducted via calls and speech-heavy workflows
- •Native speech-to-speech multimodal models expected to improve significantly
- 53:35 – 57:23
Business process automation: the underrated AI opportunity outside engineering
They discuss how most economic activity runs on repeatable procedures, not open-ended engineering work. Sherwin is especially bullish on AI embedded in deterministic, tool-integrated enterprise workflows that automate SOPs and operations.
- •Many roles are driven by repeatable SOPs with high determinism and constraints
- •AI can transform work by integrating with enterprise systems and data
- •Opportunity is larger than people on tech-centric social media assume
- •Automation can reshape how companies operate beyond software engineering gains
- 57:23 – 1:00:51
OpenAI’s ecosystem stance: why startups shouldn’t fear being ‘squashed’
Sherwin argues the market is vast and most failures aren’t due to OpenAI competing directly. He describes OpenAI as an ecosystem platform company—shipping capabilities to the API and keeping the platform neutral to foster builders.
- •Most startup failures come from weak customer resonance, not lab competition
- •OpenAI philosophy: platform-first; API is core and models generally reach the API
- •Commitment to ecosystem: neutrality, broad access, ‘rising tide lifts all boats’
- •ChatGPT scale and distribution create new opportunities for partners and builders
- 1:00:51 – 1:19:39
OpenAI’s mission, API stack overview, and lightning round wrap-up
Sherwin ties platform strategy to OpenAI’s mission of broadly distributing AI benefits, then outlines key platform primitives (Responses API, Agents SDK, UI components, evals). The episode ends with personal lightning-round picks and reflections on leaning into the AI wave.
- •Mission link: platform enables reaching ‘all of humanity’ via others building niche solutions
- •Platform layers: Responses API → Agents SDK → UI kits/widgets + evals tooling
- •Advice for not missing the wave: engage, experiment, and ignore most of the noise
- •Lightning round: books, media, products, life motto, and an Opendoor housing-model anecdote