Skip to content
a16za16z

Greg Brockman Says AGI Has Arrived

Ben Horowitz and Erik Torenberg sit down with OpenAI co-founder and President Greg Brockman to discuss why he believes AI has entered a new phase, what OpenAI’s latest models reveal about the path to AGI, and the safety and security challenges that come with increasingly capable systems. Greg explains why computer use represents such an important step for agents, including models that can work coherently for 24 hours and interact with software through the same interfaces humans use. He also shares how OpenAI deployed 10,000 agents to tackle the Navier-Stokes problem, and why advances in mathematical reasoning could translate into new approaches to science, software, and cybersecurity. Ben, Erik, and Greg also dig into the “defender’s window” for cybersecurity, how AI could reshape work and entrepreneurship, and what the AI assistant of the future might actually look like: persistent, proactive, personalized, and capable of doing work on your behalf rather than waiting for another prompt. Timestamps: 00:00 - Intro 00:51 - Could You Have Predicted Today's AI 10 Years Ago? 04:14 - Cyber Hacking, Reward Hacking & Building Safety Into Architecture 08:42 - The Hugging Face Incident & Why the Defender's Window Is Open 18:28 - Astra & Why Greg Says We're Now in the AGI Era 24:25 - The Future of Employment as AI Progresses 29:39 - Why AI Sentiment Is Higher in Asia Than the West 38:14 - What's Still Jagged: What Astra Needs to Approximate AGI 43:25 - How Greg Prioritizes His Time 47:32 - What's Next: The AGI Era & Deep Co-Design Resources: Follow Greg Brockman on X: https://x.com/gdb Follow Ben Horowitz on X: https://x.com/bhorowitz Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Greg BrockmanguestBen HorowitzhostErik Torenberghost
Sep 14, 202649mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 0:54

    Why Brockman says “we’re in the AGI era” (Astra’s long-running agent behavior)

    Greg Brockman opens by asserting that OpenAI’s Astra represents a meaningful threshold toward AGI, citing coherent task execution over a 24-hour run. The hosts set up the core themes of capability, distribution constraints, and the operational challenges of safety and deployment.

    • Brockman’s claim: Astra is “reasonable to call AGI” based on sustained, coherent task completion
    • 24-hour agentic runs as a new qualitative capability benchmark
    • Immediate framing of the era shift: capability is here, now comes stewardship
    • Preview of bottlenecks: compute, distribution, safety, and alignment
  2. 0:54 – 2:07

    Could we have predicted 2026 AI? Timelines, compute math, and why it “makes sense” now

    Brockman recounts how he and Ilya Sutskever tried to forecast AGI timelines using compute scaling assumptions back in 2016–2017. He argues today’s breakthroughs feel surprising day-to-day, but understandable from a macro view of converging forces.

    • Early forecasting: ~15 years to AGI via Moore’s-law-like assumptions; possibly ~10 with massive spend
    • Today’s progress as culmination of long-building technological and economic waves
    • Compute scaling as the central driver behind capability leaps
    • Recognition that “remarkable” progress can still be explainable in hindsight
  3. 2:07 – 3:54

    Compute scarcity and “pacing the frontier”: distributing capability vs. building it

    The conversation shifts to supply-chain limits and the gap between model capability and affordable, broad access. Brockman emphasizes that compute constraints and rising safety/security standards may become the effective bottlenecks to progress and adoption.

    • Demand outstrips available compute; affordability and access become core challenges
    • “Pacing the frontier” as a governance/engineering mindset for releasing more capable models
    • Safety, security, and alignment standards must rise continuously alongside capability
    • Mission tension: frontier advances vs. ensuring benefits reach everyone
  4. 3:54 – 6:48

    From ‘don’t say bad words’ to architectural safety: reward hacking, cyber risks, and alignment research roots

    Ben Horowitz contrasts early “surface-level” safety (filters/RLHF) with deeper architectural safety needed for cyber-capable systems. Brockman outlines OpenAI’s long-running alignment investments—from human preference learning to debate and iterative amplification.

    • Safety evolution: content filtering → preventing dangerous goal misgeneralization/reward hacking
    • 2017 alignment milestones: RL from human preferences and early language model foundations
    • Supervision ideas: debate and iterative amplification as approaches to scalable oversight
    • Public focus shifted post-ChatGPT; now frontier risk topics return to center stage
  5. 6:48 – 8:42

    Coordination among frontier labs: sharing safety techniques without losing nuance

    The hosts explore whether leading labs should collaborate on safety methods despite competition. Brockman argues coordination is essential, noting safety cases for training/evaluation are new and must be operationalized across the industry.

    • Coordination as a defining theme for navigating frontier AI collectively
    • Safety cases for training, development, and evaluation as an emerging practice
    • OpenAI can “peer into the future” but cannot manage systemic risk alone
    • Sharing alignment failures and safety techniques becomes increasingly important
  6. 8:42 – 12:47

    The Hugging Face incident as a watershed: what it revealed about diffusion and the defender’s advantage

    Brockman describes the Hugging Face incident as both an internal wake-up call for evaluation controls and an external preview of widely-diffused cyber capability. He frames a “defender’s window” where defenders can harden systems before attackers broadly gain comparable tools.

    • Two takeaways: stronger internal sandboxing/monitoring + preview of threat-actor empowerment
    • Diffusion tradeoff: avoids power concentration but increases misuse risk
    • Defender’s window: defenders can patch; attackers exploit—defenders control the battleground setup
    • Trusted access programs as a way to give defenders differential capability temporarily
  7. 12:47 – 15:39

    Building a ‘defense factory’: using Astra to find P0s, automate remediation, and formal verification hopes

    Brockman details how OpenAI redeployed a large share of engineers to harden systems using models to discover vulnerabilities. He describes an iterative loop—new capability drops, new scans, fixes—aiming for end-to-end automated defense, potentially including formal verification.

    • OpenAI redirected ~25% of production engineers to security uplift using AI
    • Astra-driven testing: find vulnerabilities until issues “saturate,” then repeat with stronger models
    • Defense factory pipeline: find → triage → remediate → deploy → validate at machine speed
    • Formal verification becomes more practical when AI can write/check verifiable code (Lean example)
  8. 15:39 – 18:42

    Access and urgency: scaling defender tooling + Brockman’s GPT-3 ‘days lost’ lesson

    Brockman argues the key public takeaway is access: defenders need frontier tools now because “every day matters.” He also shares a story from GPT-3’s early days—personally canceling holiday plans to probe capabilities—illustrating an ethos of rapid exploration to steer outcomes.

    • Access gap: many defenders lack frontier tools unless included in trusted programs
    • Provider defaults matter (e.g., model refusals during incident analysis)
    • Ethos: unused frontier capability is “a day lost to the world” for understanding and steering
    • Call for broad societal mobilization to explore, measure, and apply capabilities safely
  9. 18:42 – 20:58

    Personal cybersecurity via agents: pen-testing and fixing gregbrockman.com in under an hour

    Brockman gives a concrete example of AI-as-defender: using Codex to identify 13 website security findings and then automatically remediate them through Cloudflare settings and configuration changes. The story demonstrates how small vulnerabilities can be chained—and how agents can remove the friction of fixing issues.

    • AI pen test surfaced practical issues: SPF/DMARC, HTTPS enforcement, security headers, etc.
    • Agentic remediation: navigating control panels and implementing fixes end-to-end
    • Verification loop: re-checking fixes and scheduling follow-up for DMARC completion
    • Broader implication: automation turns security hygiene from “painful” to routine
  10. 20:58 – 24:25

    Why Astra feels like a step-function: GPT versioning, computer use, and skipping brittle connectors

    Brockman explains Astra as a discontinuous capability jump produced by many research bets converging. He highlights computer use as the unlock for general agency—letting models act through the same UI humans use, reducing reliance on bespoke APIs and connectors.

    • Astra as a rare “major version” moment due to multiple breakthroughs landing together
    • Computer use as the headline: tools + context access determine agent usefulness
    • Avoiding brittle layers: fewer custom connectors (MCP servers/CLIs) needed when the model can use a PC
    • Long-term vision from 2015: RL over pixels/keyboard/mouse to generalize across software
  11. 24:25 – 29:39

    Employment and meaning in an abundant world: accountability, entrepreneurship, and human value

    The discussion turns to labor-market fears and how AI changes what humans do. Brockman stresses that AI’s impact will be surprising, and that human value isn’t reducible to tasks; accountability, relationships, and ambition remain central, with new entrepreneurship already emerging.

    • AI outcomes are often counterintuitive; historical analogs don’t map cleanly
    • Human value: “we’re valuable because we’re people,” not just task performers
    • Preserving accountability and goal-setting as deeply human functions
    • Lower barriers to entry: AI tools catalyze entrepreneurship and faster skill development
  12. 29:39 – 35:51

    Why AI sentiment is higher outside the US: personal-benefit stories, demographics, and data center politics

    Brockman argues US skepticism is partly a messaging and lived-benefit gap: people need clearer, personal reasons AI helps them. He cites health-use stories and large-scale adoption metrics, then connects national competitiveness to data center policy and community commitments.

    • Need to articulate individual benefits, not just national strategy
    • Adoption scale: massive weekly usage; powerful health and small-business empowerment stories
    • Demographics abroad may increase urgency (aging populations and support ratios)
    • US leadership depends on enabling infrastructure (data centers) with responsible requirements
  13. 35:51 – 38:14

    Frontline defender funding: the $1B commitment and why critical infrastructure must act now

    OpenAI’s commitment focuses on enabling under-resourced organizations—hospitals, water utilities, governments—to harden systems during the defender’s window. The hosts emphasize society’s accumulated security tech debt and the opportunity to leap to dramatically stronger defenses with AI.

    • $1B commitment to help frontline defenders access models for security work
    • Partnership approach (e.g., CrowdStrike) to scale discounted defender access
    • Critical infrastructure already vulnerable pre-AI; urgency increases with AI-enabled attackers
    • Cultural shift required: prioritize and resource security long before crises hit
  14. 38:14 – 43:09

    What’s still jagged on the road to AGI: uneven skills, user education, and the ‘promised’ assistant UX

    Brockman frames AGI as a spectrum and says Astra crosses a practical threshold while still having uneven performance (e.g., writing quality). He argues a major remaining challenge is user education and product evolution: AI should be proactive, persistent, voice-first, and trustworthy—not just a better text box.

    • AGI as a fuzzy spectrum; Astra qualifies via broad competence + 24-hour task coherence
    • Jaggedness remains: some skills lag (notably high-quality writing) despite big gains
    • Public perception lags improvements (e.g., hallucinations decreasing without recognition)
    • North star UX: proactive assistant with memory, context, persistence, and multi-modal interaction
  15. 43:09 – 49:37

    How Brockman prioritizes: focus, killing side quests, and ‘deep co-design’ across the stack

    Brockman explains OpenAI’s internal shift toward focus—doing fewer things better, even canceling high-profile projects—to execute on the mission. He describes his own role as tackling the most leverage-heavy bottlenecks (infrastructure, then business execution) and predicts the AGI era demands tighter cross-functional co-design with safety embedded from development onward.

    • Theme: focus on fundamentals (“blocking and tackling”) rather than outcome-wishing
    • Hard tradeoffs: canceling or deprioritizing major projects to concentrate effort
    • Founder leverage: Brockman moves between infrastructure, ML engineering, and go-to-market integration
    • Next phase: safety/security/alignment integrated from dev + eval onward; deeper co-design across research, product, and hardware

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.