Dwarkesh PodcastJoseph Carlsmith - Utopia, AI, & Infinite Ethics
CHAPTERS
- 0:00 – 0:55
Preview: Utopia, infinite ethics, and modeling the far future
A quick montage frames the episode’s three big themes: the possibility of profoundly better futures, the puzzles of ethics under infinity, and the limits of human forecasting. It sets an expectation that the conversation will blend philosophy with AI-relevant implications.
- •Utopia as a genuinely possible, radically better future
- •“Infinite ethics” as a challenge for moral reasoning
- •A middle ground between ignoring weird implications and overreacting to them
- •Why thinking about the future requires lossy abstractions
- 0:55 – 3:40
Joe Carlsmith’s work: AI existential risk, timelines, and takeoff
Dwarkesh introduces Joe’s background at Open Philanthropy and in philosophy, then they discuss what Joe focuses on in AI safety work. The emphasis is on how timelines and takeoff speed estimates feed into prioritization and perceived catastrophe risk.
- •Joe’s role at Open Philanthropy and philosophy background
- •Current focus: AI timelines and takeoff speeds
- •Why probability estimates (1% vs 10% vs 90%) matter for priorities
- •Planning implications when timelines are shorter
- 3:40 – 7:56
Defining a “better future” when the future might feel alien
Dwarkesh raises a longtermist worry: future societies might be disturbing or unintuitive from today’s standpoint. Joe argues the relevant standard is what you’d think with deep understanding, not initial culture shock.
- •Past humans would likely prefer modern life after adjustment
- •Future-to-present gap may be far larger than past-to-present
- •Initial alienation vs endorsement after fuller understanding
- •Distinguishing inside-experience from holistic understanding
- 7:56 – 10:06
Will the future regret our era’s existential risk-taking?
They explore whether the 21st century will be viewed as a “cool, formative” adolescence or a horrifying near-miss. Joe argues ex post nostalgia is misleading for death risks—survivorship bias hides the probability mass where everything ends.
- •Teenage-analogy limits when stakes include extinction
- •Survivorship bias in looking back at risky periods
- •Comparisons to WWII and the Cuban Missile Crisis
- •Why the ex post perspective is the wrong calibration tool for existential risks
- 10:06 – 11:48
Utopia: why it matters, and why our visions are too status-quo anchored
Joe defines utopia as a profoundly better future and argues we underestimate how large the value gap could be. He suggests common “utopias” are shallow extensions of current life rather than transformations of the human condition.
- •Utopia as ‘profoundly better,’ not necessarily perfect
- •We underestimate the magnitude of possible improvement
- •Common visions are overly anchored to today’s constraints
- •The improvement might be more like asleep→awake than bad job→vacation
- 11:48 – 15:47
Getting to utopia requires wisdom and capability growth, not a simple ‘hedonium’ target
Dwarkesh presses on whether utopia is recognizable or categorically different; Joe reconciles this via a long process of becoming wiser. He rejects simplistic hedonium tiling as a default and emphasizes value complexity and cognitive enhancement before irreversible choices.
- •Tension: recognizable utopia vs radically alien future
- •Utopia as an outcome of gradual species-level maturation
- •Critique of ‘hedonium’ framing (sterility/uniformity vs real pleasure)
- •Value complexity and uncertainty under cognitive enhancement
- •Need to chart what minds can be before locking in civilization design
- 15:47 – 26:01
Utopian thinking’s dangers: ideology, rigidity, and uncooperative action
Dwarkesh notes historical atrocities tied to utopian projects and asks if EA/longtermism shares that risk. Joe treats the danger as broader: intense conviction can drive rigid, destructive behavior unless tempered by caution and coordination norms.
- •Utopian projects as historical red flags
- •Risk comes from conviction and rigidity, not just EA/utilitarianism
- •Importance of how beliefs are acted on (cooperative vs destructive)
- •Many people believe the world can be much better, but it’s often not operative day-to-day
- 26:01 – 28:13
Constraints and dystopian equilibria: Robin Hanson’s EMs vs coordination
Robin Hanson’s emulation (‘EM’) scenario raises a competitive equilibrium where digital labor is driven toward subsistence. Joe agrees this is not utopia and suggests coordination and preemptive governance could, in principle, avert worst competitive dynamics—though complexities remain.
- •Hanson’s EM world as mediocre or worse, not utopian
- •Competitive pressures can push futures in bad directions
- •Coordination/preemptive action as the main counterweight
- •Uncertainty about which future-quality distribution is realistic
- 28:13 – 34:52
How much compute is the human brain? Why FLOPs estimates matter
Joe explains his neuroscience-grounded approach to estimating FLOPs needed to emulate task-relevant human cognition, and how Open Phil uses this as an input to broader AI timeline models. They discuss why training compute estimates hinge on assumptions about scaling and learning efficiency.
- •Goal: estimate FLOPs sufficient for task-relevant cognition emulation
- •Brain compute as an input into ‘training cost to human-level AI’ methods
- •Ajeya Cotra’s training-compute distribution and key uncertainties
- •Horizon length: how often you get learning signal per experience/token
- •Dependence on scaling hypothesis vs algorithmic breakthroughs
- 34:52 – 41:03
Methods, uncertainty, and what neuroscience can (and can’t) tell us
Joe outlines multiple triangulation methods: mechanistic brain modeling, comparisons to AI (especially vision), physical energy limits, and communication-to-computation extrapolations. He emphasizes expert disagreement and large uncertainty, while downplaying quantum-mind hypotheses.
- •Triangulation across multiple weak evidence sources
- •Mechanistic decomposition and additivity vs multiplicative interactions
- •Neuroscience lacks algorithm-level understanding; experts disagree
- •Why some emphasize biophysical detail vs others emphasize mechanistic sufficiency
- •Very low credence in Penrose-style quantum cognition claims
- 41:03 – 45:12
Infinite ethics: why standard moral theories break under infinity
Joe introduces infinite ethics as the attempt to rank and act in infinite worlds, where many ethical principles yield paradox or undefined comparisons. He argues it matters theoretically (stress-testing moral theories) and practically (non-zero credence we live in, or can affect, infinite settings).
- •Infinite worlds can make common ethical theories ‘break’
- •We still have intuitions about infinite heaven vs infinite hell
- •Non-zero credence that the universe is infinite (or influence could be)
- •Even tiny probabilities of infinite impact can dominate EV reasoning
- •Decision theory can create ‘acausal’ influence scenarios relevant to infinity
- 45:12 – 50:58
Acausal influence and evidential decision theory in an infinite cosmos
They discuss how correlated copies (or similar agents) across an infinite universe could create something like influence without causal contact, especially under evidential decision theory. This leads to deep questions about what “difference you make” even means in infinite sets and how to compare infinite outcomes.
- •Evidential decision theory and treating correlated agents as ‘under your control’
- •Infinite universe hypothesis with far-away copies and variations
- •Empirical question: what changes under acausal reasoning?
- •Normative question: how to rank infinite outcomes; impossibility results
- •Why infinity breaks both anthropics and ethics in analogous ways
- 50:58 – 1:01:38
Living with destabilizing implications: longtermism, humility, and ‘what feels real’
Dwarkesh asks how Joe stays committed to EA/longtermism amid unresolved infinite-ethics problems. Joe recommends caution about radical life changes from confusing abstractions, while arguing survival and wisdom-building is a convergent priority that keeps options open for better future reasoning.
- •No fully principled stopping point; it’s a real art to respond to weird ideas
- •Caution about wholesale ethical upheaval from destabilizing arguments
- •Strategy: survive, become wiser, keep options open, then act better
- •Using ‘does it feel real?’ as a signal after serious engagement
- •Ant ethics as an example of acknowledging uncertainty without becoming extreme
- 1:01:38 – 1:18:32
Anthropic reasoning: SIA vs SSA, doomsday argument, and where both break
Joe explains self-indication (SIA) and self-sampling (SSA) through God/white-room cases and the doomsday argument. He argues SSA yields implausible ‘telekinetic’ prediction patterns, while SIA has its own pathologies (e.g., pushing toward huge or infinite populations), suggesting infinities break anthropics much like they break ethics.
- •SIA: more observers with your evidence → higher probability of that world
- •SSA: you’re a random sample from a reference class → doomsday-style updates
- •SIA’s handling of ‘surprised to be early’ and room-number conditioning
- •SSA problems: apparent telekinesis and predicting fair coin outcomes pre-toss
- •SIA problems: ‘presumptuous’ updates toward enormous/infinite universes
- •Both struggle with infinities and measure problems in cosmology
- 1:18:32 – 1:24:02
Futurism’s ‘unreality’: abstraction, social dynamics, and keeping models grounded
Joe critiques futurist discourse as often losing contact with reality because the mind must use extremely lossy abstractions. He suggests the challenge is to preserve concreteness (like vivid history does) while knowing any specific imagined detail is likely wrong.
- •Imagination is limited; future-modeling needs gappy abstractions
- •Futurism can become a ‘say anything’ zone without constraints
- •Status/social dynamics can pull discourse away from truth-tracking
- •Contrast with history: concreteness shifts moral/emotional understanding
- •A ‘delicate dance’: imagine concrete scenes, admit they’re wrong, keep the concreteness anyway
- 1:24:02 – 1:32:09
Blogging, productivity, and recommendations to close
They end with a practical discussion of Joe’s writing process, trade-offs between editing and output, and blogging as an anti-perfectionism exercise. Joe shares book recommendations and plugs where to find his work.
- •Long posts as a deliberate trade-off: less editing, more output
- •Blogging as practice in shipping ideas vs perfectionism
- •Balancing reader-attentiveness with productivity constraints
- •Recommendations: The Precipice; Angels in America; Housekeeping; engaging Nick Bostrom
- •Where to find Joe’s blog, website, and AI-related work links