Skip to content
Dwarkesh PodcastDwarkesh Podcast

Michael Nielsen on Dwarkesh Patel: Why Ether Died Slowly

Lorentz fit Einstein equations while keeping the ether ontology; Michelson-Morley only ruled out ether wind, so a single result cannot force a paradigm shift.

Dwarkesh PatelhostMichael Nielsenguest
Apr 7, 20262h 3mWatch on YouTube ↗

CHAPTERS

  1. 0:00 – 7:04

    Michelson–Morley: why the famous ‘ether-killer’ story is misleading

    Dwarkesh and Michael use Michelson–Morley to show how scientific progress often outruns clean, textbook narratives. Rather than simply “disproving the ether,” the experiment primarily discriminated among competing ether theories—and even its interpreters (including Michelson) didn’t converge quickly on a single conclusion.

    • Michelson–Morley as a test between multiple ether models, not a simple refutation
    • Einstein likely wasn’t decisively guided by the experiment when forming special relativity
    • Why naive Popperian falsification is hard to apply to real episodes in science
    • The role of interpretation: what an experiment ‘means’ depends on the surrounding theory set
  2. 7:04 – 15:11

    Lorentz, Poincaré, and Einstein: math vs interpretation, and expertise as a trap

    They dig into how Lorentz had much of the right mathematics while clinging to an ether-based interpretation, and how Poincaré grasped key conceptual issues yet missed the final reframing. The episode illustrates how brilliant scientists can stall for decades when a conceptual shift requires abandoning deeply internalized intuitions.

    • Lorentz transformations: correct structure, ether framing, and ‘local time’ as a mathematical convenience
    • Poincaré’s near-miss: relativity principle and light-speed invariance, but a dynamical (not kinematic) view of contraction
    • Muon time dilation as later evidence that ‘local time’ is physically real time
    • No centralized method: great scientists can remain unconvinced long after community consensus shifts
  3. 15:11 – 23:20

    Verification loops can be centuries long: heliocentrism, parallax, and what counts as ‘better’

    Dwarkesh argues that key ideas can be accepted well before decisive experimental confirmation—sometimes by centuries. Copernicus wasn’t initially more accurate (or even simpler in epicycles), yet the scientific community still found reasons to prefer the heliocentric direction, raising the question of what ‘progress’ means absent tight feedback loops.

    • Aristarchus vs. stellar parallax: 2nd century BC idea, 1838 measurement
    • Copernicus initially loses on accuracy and sometimes simplicity to refined Ptolemaic models
    • Why ‘verification’ often lags ‘adoption’—and why that gap is philosophically important
    • Newton’s later unification (celestial + terrestrial + tides) as a retrospective explanation of why heliocentrism mattered
  4. 23:20 – 29:59

    Why natural selection wasn’t obvious earlier: the idea existed, but not the full explanatory program

    They contrast the apparent conceptual simplicity of Darwin’s natural selection with Newtonian gravity, asking why Darwinism arrived so late. Michael emphasizes that Darwin’s genius wasn’t the bare mechanism (which breeders partly knew), but the comprehensive argument tying enormous swaths of biology and geology into one explanatory web.

    • Huxley’s ‘how stupid not to think of it’ reaction vs. Newton’s Principia (never feels obvious)
    • Breeders knew components; Darwin made the case for centrality across the biosphere
    • Lucretius-like proto-ideas differed crucially (one-time filter vs. ongoing process + tree of life)
    • Prerequisites for Darwin/Wallace: deep time (Lyell), paleontology, biogeography, and accumulated anomalies
    • Missing mechanism of heredity (genes) as a major gap in early Darwinism
  5. 29:59 – 35:54

    AI ‘science’ and AlphaFold: prediction without classic explanation

    The conversation turns to modern AI successes and whether they resemble traditional scientific theories. AlphaFold becomes the focal example: an extraordinary predictive artifact built atop decades of expensive data acquisition, raising questions about explanation, interpretability, and whether new kinds of ‘theories’ are emerging.

    • AlphaFold’s success depends heavily on the Protein Data Bank and experimental infrastructure
    • Classic ideal: few parameters + deep principles vs. neural nets as high-dimensional fits
    • Three stances: (1) not explanation, (2) contains extractable micro-explanations, (3) a new kind of scientific object
    • Analogy to Mathematica: previously-unusable complexity becoming a workable intermediate object
    • Human extraction of model insights (e.g., chess strategies from AlphaZero) as a template
  6. 35:54 – 42:46

    Could gradient descent discover general relativity? Distillation, regularizers, and ‘global’ theory shifts

    Dwarkesh challenges the idea that brute optimization over observations can yield deep conceptual flips like Copernicus or Einstein. Michael suggests we may lack the right ‘verbs’ (operations) over models—distillation, simplification, constraint-setting—but agrees the hardest progress often requires non-local jumps and new framing constraints.

    • Why distilling a Ptolemaic-style fit may not trigger the crucial ‘swap’ to a new worldview
    • Theory change as a forcing function: special relativity makes Newtonian instantaneous gravity untenable
    • Einstein’s path: ugly intermediate stages before a simple final formulation
    • Need for many parallel research programs with different starting biases and heuristics
    • Scientific progress as maintaining diversity long enough for rare conceptual breakthroughs
  7. 42:46 – 50:39

    Anomalies everywhere: Uranus vs Mercury, Prout’s hypothesis, and selection bias in scientific legends

    They explore how the same style of anomaly-handling can succeed brilliantly (Neptune) or fail (Vulcan), undermining simplistic falsification stories. A deeper theme emerges: most anomalies are mundane, but a few are revolutionary—yet there’s no reliable ex ante rule to tell which is which.

    • Neptune prediction as a triumph of Newtonian patching; Mercury precession as a GR opening
    • Why ‘just add an auxiliary hypothesis’ is both indispensable and often misleading
    • Pioneer anomaly resolving into thermal asymmetry (most anomalies are like this)
    • Prout’s whole-number atomic weights and chlorine’s 35.5: 85-year hostile loop until isotopes
    • Implication for AI: tight experimental loops don’t automatically translate into easy theory selection
  8. 50:39 – 58:56

    Why aliens may have a different tech stack: the tech tree is vastly larger than we assume

    Michael argues that “science” isn’t a short phase civilizations quickly complete; instead, the space of discoverable ideas and engineered possibilities may be enormous and only partially explored. Different perceptual biases and historical contingencies could steer civilizations into distinct regions of that space, producing divergent ‘stacks.’

    • Analogy: computer science got a ‘theory of everything’ early (Turing/Church), yet decades of deep discoveries followed
    • Phases of matter as an expanding frontier (superconductors, fractional quantum Hall, etc.)
    • Programming and information manipulation likely contain many undiscovered fundamental primitives
    • Path dependence: cognition and perception (visual vs auditory biases) could shape exploration routes
    • ‘GitHub for aliens’ thought experiment: vast libraries of ideas may be legible only with immense effort
  9. 58:56 – 1:15:24

    Diminishing returns vs ‘restocked dessert buffets’: new fields reset the frontier (and enable gains from trade)

    Dwarkesh presses on the empirical ‘ideas getting harder’ phenomenon, while Michael offers a counter-model: diminishing returns only apply in a static menu, but new fields continually open up and replenish the buffet. They also explore implications of divergent paths: long-run comparative advantage and gains from trade between civilizations—tempered by power and transaction costs.

    • Why ‘low-hanging fruit’ arguments can fail when new fields appear unexpectedly (e.g., computer science)
    • Attention centralization and ‘fashion’ effects in what gets seen as the big breakthrough
    • Trade implications of divergent stacks: friendliness and exchange can be deeply valuable
    • Limits: power imbalances, transaction costs, and the difficulty of translating capacities (not just ideas)
    • Manufacturing futures: commoditized fabrication vs entrenched process knowledge (3D printing vs ribosome analogy)
  10. 1:15:24 – 1:26:25

    Are there infinitely many deep principles? Noether, Church–Turing, crypto, and the expanding primitive set

    They question whether foundational principles are a small, nearly-complete set, or whether we’ll continue discovering new ‘deep invariants’ that reorganize wide areas of thought. Michael’s instinct is that new primitives keep appearing (especially in computation), suggesting the space of deep principles may be far from exhausted.

    • Examples of deep unifiers: Noether’s theorem and universality in computation
    • Computation’s ‘early foundation’ still hid later primitives: public-key crypto, consensus/ledgers/crypto-systems
    • Deep principles as discoveries inside a fixed formal universe (like computation)
    • Bloom et al.-style evidence: maintaining progress requires rapidly growing R&D effort
    • Michael’s critique: narrow metrics miss externalities and field-creating shifts (e.g., GPUs enabling new trajectories)
  11. 1:26:25 – 1:35:28

    What drew Michael Nielsen to quantum computing early: contingent timing, tools, and ‘market for follow-ups’

    Michael recounts how quantum computing could have emerged earlier (von Neumann plausibly could have invented it), but didn’t—highlighting how salience and enabling instrumentation matter. He then explains his own entry point: being handed a stack of foundational papers and recognizing a tractable frontier with profound unanswered questions.

    • Why quantum computing wasn’t ‘inevitable’ in the 1950s despite the right minds existing
    • 1980s convergence: computation becomes culturally salient + improved ability to manipulate individual quantum systems
    • Feynman/Deutsch as foundational; importance of having (or lacking) real devices as a bottleneck
    • Michael’s 1992 exposure via Gerard Milburn and immediate sense of tractable, foundational open problems
    • Following up as a career strategy: choosing where to ‘pick up the shovel’ given one’s skills and the field’s gaps
  12. 1:35:28 – 1:43:57

    Does science need a new way to assign credit? Open science as political economy, not just access

    Michael reframes open science as an evolution of the credit-and-attribution machinery that makes large-scale knowledge production possible. As sharing expands from papers to code, data, and in-progress work, the core challenge becomes how to price and reward contributions so that the system incentivizes the right kinds of disclosure.

    • Open science’s practical wins: open access, open code, open data becoming default expectations
    • Historical analogy: moving from secrecy/anagrams to journals + reputation-based careers
    • The modern mismatch: valuable outputs (code/data/partial results) often lack clear credit mechanisms
    • Physics vs biology preprints paradox: both claim competitiveness as the reason for opposite behaviors
    • Collective science example: LHC-scale collaboration where no single person can deeply grasp every layer
  13. 1:43:57 – 1:49:17

    Prolificness versus depth: routine throughput, high-variance bets, and the psychology of non-release

    They debate whether impact comes from high output (equal-odds productivity) or from long gestation on a few deep problems. Michael proposes a hybrid: ruthlessly accelerate routine work while making space for high-variance exploration—while noting a common failure mode where talented people never ship anything due to aversion to judgment.

    • Two modes: routine execution vs high-variance exploration where outcomes are unclear
    • Einstein’s 1905 as an extreme case; Darwin as prolific in correspondence and evidence-building
    • Equal-odds rule (Simonton): impact correlates with output volume—tempered by counterexamples like Gödel
    • Non-release failure mode: brilliance + obsession with ‘the great project’ leading to zero publications
    • Desire for biographies of the ‘gifted who missed’ as a way to learn what blocks real contribution
  14. 1:49:17 – 2:03:03

    What it takes to internalize what you learn: forcing functions, creative artifacts, and AI as an escape hatch

    Dwarkesh describes a podcaster’s dilemma: you can produce good episodes without truly internalizing knowledge, so learning doesn’t compound. Michael argues that deep learning usually requires demanding forcing functions (problems, artifacts, public stakes) and time spent stuck—while warning that AI chat can substitute for the aversive, growth-producing parts of understanding.

    • The ‘no clamp’ problem: superficial understanding decays without exercises or artifacts
    • Forcing functions: problem sets, building implementations, writing essays/books, or other creative outputs
    • Importance of being stuck; long-duration projects yield durable internalization
    • Raising stakes (‘jugular’ tasks) changes effort and depth—analogous to athletic training vs flow
    • AI tools: helpful for routine work, but seductive as a way to avoid the hardest thinking

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.