No PriorsNo Priors Ep. 105 | With Director of the Center of AI Safety Dan Hendrycks
CHAPTERS
- 0:00 – 1:24
Why Hendrycks chose AI safety: big stakes and neglected tail risks
Sarah opens by introducing Dan Hendrycks and his work on widely used evals. Dan explains he focused early on AI safety because AI seemed destined to be one of the century’s most important forces, yet key risks were being systematically under-addressed.
- •AI’s likely outsized impact motivated an early career focus on safety
- •AI safety framed as managing under-addressed tail risks
- •Perceived gap between AI’s importance and the attention it was receiving
- 1:24 – 3:12
What labs can (and can’t) do: racing dynamics and geopolitically driven outcomes
Dan contrasts the Center for AI Safety’s role with internal lab safety teams. He argues labs can implement basic misuse safeguards, but competitive and geopolitical pressures largely constrain their ability to meaningfully alter broader outcomes.
- •Labs can add basic refusals (e.g., obvious bio/terrorism prompts)
- •Market/geopolitical races limit willingness to slow down
- •Many impacts (labor disruption, power concentration) exceed company control
- •Safety is broader than technical fixes or prompt refusals
- 3:12 – 4:48
Alignment vs safety: aligned systems can still be geopolitically dangerous
Dan defines safety as a broad risk-management umbrella, with alignment as only one subset. Even perfectly aligned systems can yield dangerous strategic competition between states, pushing rapid military integration and higher risk tolerance.
- •Safety includes non-technical risks (power concentration, structural dynamics)
- •Alignment ≠ overall safety; obedient models can still enable conflict
- •US- and China-aligned AIs could intensify strategic competition
- •Geopolitical pressures can force risky deployment even with “aligned” models
- 4:48 – 6:51
National security relevance today vs soon: cyber, bio, and capability trajectories
Sarah asks why AI matters for national security, and Dan emphasizes current limits but fast-moving trajectories. He highlights bio/virology as becoming newly salient with reasoning models, while other domains remain more prospective but could shift quickly.
- •AI not uniformly decisive for national security yet, but timelines may compress
- •Cyber risk is important to prepare for, even if not yet catastrophic via AI
- •Bio/virology expertise is becoming more accessible through advanced models
- •AI could become a backbone of economic and military power, prospectively
- 6:51 – 10:00
Reducing misuse without killing benefits: gated access and enterprise pathways
Discussing defensive cyber and biotech benefits, Dan argues safety measures needn’t block legitimate use. He proposes simple gating—restricting sensitive expert capabilities for unknown users while enabling vetted access through enterprise channels.
- •Claims the safety/benefit trade-off is often overstated
- •Gating sensitive bio/cyber capabilities: refusals for unknown users
- •Vetted access via enterprise accounts for legitimate labs and startups
- •Harder domain: calibrating export controls without escalating conflict incentives
- 10:00 – 12:35
How AI might be weaponized: drones, EMP research, and destabilizing surveillance
Dan outlines pathways for AI-enabled weaponization beyond bio and cyber. He focuses on drones and on intelligence/situational awareness that could undermine nuclear second-strike stability by revealing submarines or hardened launchers.
- •Non-state actors are more plausible bio-weapon users than state actors
- •Cyber operations likely for both state and non-state actors
- •Drone swarms and autonomous systems as a default conventional AI weapon
- •AI-driven situational awareness could destabilize nuclear deterrence
- 12:35 – 14:43
Why voluntary pauses and treaties are fragile: verification, enforcement, and second-order effects
Dan argues “voluntary slowing” is often conflated with safety, but lacks teeth without verification and enforcement. He critiques strategies that assume cooperation or ignore how rivals would respond to visible bids for dominance.
- •Voluntary pauses can advantage worst actors unless enforced
- •Treaties need verification mechanisms and credible enforcement
- •Cyber norms and anti-espionage precedents show low compliance
- •“Race to superintelligence to dominate” ignores adversary responses and espionage
- 14:43 – 17:51
Immigration and talent: keep top researchers while accepting security realities
Sarah asks about immigration policy under espionage concerns. Dan argues AI-talent immigration should be separated from broader border politics and made easier for highly skilled researchers, noting the US depends heavily on multinational talent.
- •AI-talent policy should be distinct from southern border debates
- •Make it easier for top AI researchers to stay in the US
- •High multinational representation is strategically important to US capability
- •Security risks exist, but pushing talent out can backfire competitively
- 17:51 – 22:53
MAIM explained: “Mutually Assured AI Malfunction” as a deterrence regime
Sarah introduces Hendrycks’ proposal (with Schmidt and Wang) for MAIM, a deterrence concept inspired by nuclear strategy. Dan describes how states might deter destabilizing “superweapon” AI projects via credible threats of espionage and sabotage.
- •Analogy to MAD: shared vulnerability can deter extreme actions
- •As AI becomes pivotal, states may preemptively sabotage destabilizing projects
- •Espionage makes it hard to secretly pursue decisive dominance
- •Goal: deter “decisive edge” pursuits while accepting continued competition elsewhere
- 22:53 – 25:35
Policy toolkit: intelligence collection, cyber options, and chip non-proliferation
Asked for actionable steps for today’s administration, Dan proposes pragmatic statecraft. He emphasizes improved intelligence on rival AI programs, preparing cyber options to disable destabilizing projects, and tracking/inspecting AI chip flows like other strategic materials.
- •Expand intelligence efforts to monitor state AI programs (avoid surprise)
- •Prepare CYBERCOM-style options to disrupt truly destabilizing projects
- •Improve chip tracking via licensing regimes and end-use checks
- •Non-proliferation focus: prevent chips reaching rogue actors
- 25:35 – 30:37
Compute security in practice: DeepSeek, export-control limits, and deterring intent vs capability
Sarah challenges compute-security assumptions using China’s progress and compute-efficient training (e.g., DeepSeek). Dan argues this undermines “capability denial” strategies against major powers, making deterrence of intent more realistic while still restricting rogue actors’ access.
- •Compute efficiency reduces confidence in blocking great-power capability
- •Export controls still useful if enforced (end-use checks, follow-the-chips)
- •China may still acquire chips or steal weights; total denial is unlikely
- •Focus shifts: restrict rogue actors’ capabilities; deter state intent
- 30:37 – 36:24
Where evals go next: Humanity’s Last Exam, jagged intelligence, and agent benchmarks
In closing, Sarah pivots to evaluation: why new benchmarks matter and what they signal. Dan explains Humanity’s Last Exam as a “ceiling” for closed-ended academic questions, argues agentic task performance remains weak, and predicts a major perception shift once models gain reliable agent skills.
- •Humanity’s Last Exam: professor-sourced, closed-ended, high-difficulty STEM questions
- •Signals when traditional exam-style benchmarks are nearing exhaustion
- •Need agent evaluations to measure real task automation over time
- •Frontier is jagged: strong reasoning in spots, weak everyday agency elsewhere
- •Economic and societal “vibes shift” likely when agent skills arrive