Lex Fridman PodcastSertac Karaman: Robots That Fly and Robots That Drive | Lex Fridman Podcast #97
CHAPTERS
- 0:00 – 5:25
Autonomous flying vs. driving: what’s harder at scale?
Sertac compares consumer drone autonomy with the broader challenge of deploying autonomous systems at massive scale. The key difficulty isn’t just control or perception, but operating densely in human-centered environments with safety, legal, and social constraints.
- •Consumer drone autonomy is comparatively easier than widespread autonomous transportation
- •True difficulty is large-scale deployment among humans
- •Robots historically succeed in isolated/controlled settings (factories, Mars)
- •Dense autonomy (thousands of vehicles) is unprecedented
- •Ground vehicles may reach high-density deployment before air vehicles
- 5:25 – 6:37
Sky filled with drones? Delivery vs. passenger air mobility
Lex and Sertac explore whether thousands of drones could occupy the sky, and what high-value use cases might justify it. Sertac highlights building-to-building regional travel as a compelling opportunity beyond delivery.
- •Large-scale delivery drones are plausible, but not the only target
- •Passenger transport: rooftop-to-rooftop trips (e.g., Boston to NYC)
- •Small group air-taxi concepts (4–8 passengers)
- •Value proposition: time savings vs. cost premium
- •Air mobility may complement rather than replace airports/airplanes
- 6:37 – 10:27
Flying cars and the “agile airspace” problem (and safety certification)
Sertac frames feasibility around an underused band of airspace: high enough to avoid simple hazards, low enough to avoid commercial aviation. The core technical challenge is building far more complex autonomy software while meeting aircraft-level safety assurance.
- •Long-horizon forecasting is highly uncertain (even exotic propulsion may emerge)
- •“Agile airspace” is underutilized and potentially valuable
- •Key barrier: safe autonomy + certification-grade reliability
- •Software complexity likely orders of magnitude beyond today’s aircraft systems
- •Machine learning may drive intelligence, but safety assurance is its own challenge
- 10:27 – 13:05
Simulation’s real role: development first, then training
The conversation shifts to simulation as a tool not just for ML training but for engineering development and validation. Sertac distinguishes what’s easy to simulate (dynamics, internal sensors) from what remains hard (cameras, radar, and especially humans).
- •Simulation is crucial for system development, not only ML training
- •Dynamics and inertial sensors are relatively straightforward to simulate
- •Exteroceptive sensors are harder: lidar easier than cameras/radar
- •Camera simulation is improving rapidly (rendering/ray-tracing realism)
- •Next major hurdle: simulating human behavior realistically
- 13:05 – 17:35
The hardest simulation target: humans (and why prediction is still missing)
Sertac explains why humans are still the “tell” in rendering and behavior models, and why this matters for autonomy. He outlines the evolution from localization/mapping to perception, and the remaining challenge of predicting what humans will do next.
- •Humans are easy for our brains to spot as ‘off’ in rendered scenes
- •Physics-based simulation works when rules are known; humans require data-driven modeling
- •Human labeling is a bottleneck for realism and behavior modeling
- •AV progress stages: (1) where am I? (2) where is everyone? (3) what will they do next?
- •Prediction of other agents remains a central unsolved difficulty
- 17:35 – 24:30
Game theory on the road: interaction, “abuse,” and social trade-offs
Lex probes the game-theoretic nature of driving—how the ego vehicle shapes outcomes through signaling and assertiveness. Sertac discusses how empty vehicles are treated as objects, the need to reason about aggressiveness, and the broader efficiency–livability trade-off.
- •Interaction effects are often ignored in current AV prediction stacks
- •People treat unmanned vehicles differently than human-driven cars
- •Robots may need to infer driver types (aggressive vs cooperative) from behavior
- •Design choices force explicit trade-offs: efficiency vs livability/sustainability
- •Testing in human environments introduces unavoidable innovation–risk tension
- 24:30 – 29:45
Waymo vs Tesla: not just safety vs speed, but different goals
Sertac reframes the popular comparison: Waymo and Tesla optimize along different axes. Waymo is positioned as building a deep “AI engine” with a research mindset, while Tesla pairs incremental automation with an immediate consumer product and market feedback loop.
- •Public understanding and informed consent matter across strategies
- •Waymo: long-term engine-building, strong research orientation
- •Tesla: consumer product integration + incremental capability rollout
- •Both may converge on similar underlying autonomy ‘engines’
- •Use cases and product incentives shape engineering decisions
- 29:45 – 36:10
Starting Optimus Ride: realistic scope, real deployments, human-in-the-loop operations
Sertac recounts the motivations behind founding Optimus Ride and skepticism about near-term full autonomy. The company’s approach emphasizes a systems view: vehicles are autonomous, but humans supervise at a higher level to improve efficiency rather than safety.
- •Early industry timelines for full autonomy felt unrealistic to the founders
- •Urban Challenge experience revealed real-world limitations and complexity
- •Optimus Ride focuses on market-constrained deployments instead of universal autonomy
- •Humans can supervise fleets (many vehicles per operator) rather than tele-drive
- •Goal: safe autonomy by default; human input boosts throughput/efficiency
- 36:10 – 42:37
Geofenced mobility that people actually like: replacing the “shuttle experience”
Optimus Ride targets transportation-deprived campuses/industrial sites where traditional shuttles are unpopular and costly. Smaller, agile vehicles and better routing reduce wait times and can reclaim land currently devoted to parking, improving economics and livability.
- •Initial markets: ~2x2 mile geofenced environments (yards, campuses, districts)
- •These areas are often underserved because conventional ride-hail economics don’t fit
- •Reducing parking demand can unlock major land/value and urban design benefits
- •Shuttles are disliked partly because driver cost pushes them to be large and infrequent
- •Smaller vehicles improve agility, routing flexibility, and user experience
- 42:37 – 48:17
Scaling to city-level density: data velocity, fleet economics, and ultra-low-cost rides
They discuss what it takes to reach the ‘you see one, you see another’ density that makes autonomous mobility feel ubiquitous. Sertac emphasizes data density (repeatedly seeing the same intersections) and how shifting labor costs changes vehicle design and ride economics.
- •City-scale service likely requires tens of thousands of vehicles; taxis can be 10–20% of traffic in some cities
- •Data density/velocity beats sparse nationwide miles for learning local interactions
- •People tolerate waiting for goods more than for personal rides—affects operations strategy
- •Driver labor dominates ride-hail cost; supervision-at-scale changes the unit economics
- •Purpose-built lightweight urban vehicles could drive per-trip costs extremely low
- 48:17 – 58:39
Timelines and ‘time dilation’ in tech prediction; why iteration beats forecasts
Lex asks for deployment timelines; Sertac argues that prediction is fundamentally unreliable due to unknown technology gaps and poor metrics for AI progress. He introduces ‘time dilation’ in forecasting and emphasizes iterative learning cycles as the practical path forward.
- •AI progress lacks a clear Moore’s-law-like metric, making forecasts shaky
- •‘Time dilation’: far-off capabilities feel near; near-term capabilities feel far
- •Technology gaps can be unknown in difficulty (possibly ‘AI-complete’)
- •Practical planning works at 1–2 year horizons; beyond that becomes vision-setting
- •Iterative experimentation in real contexts accelerates learning more than speculation
- 58:39 – 1:03:21
Is lidar a crutch? Cameras, fusion, cost trade-offs, and certification
Sertac supports the idea that camera-only autonomy can eventually work, but argues lidar may remain attractive as it becomes cheap and simplifies safety cases. He explains why lidar accelerates demos today, and why early deployments likely include some form of lidar.
- •Camera-only autonomy is plausible long-term, but may demand heavy compute and complexity
- •Lidar can simplify engineering and demos—hence its popularity in AV startups
- •Future may favor lidar if it becomes cheap and reduces compute/certification burden
- •Optimus Ride philosophy: strong computer vision + minimal but strategic lidar + sensor fusion
- •Short-range/low-res solid-state lidar is feasible; fundamental limits (e.g., speed of light) constrain performance
- 1:03:21 – 1:12:48
AlphaPilot drone racing: aggressive autonomy, hardware/software co-design, and human limits
Sertac describes the AlphaPilot challenge and why drone racing is an ideal testbed for pushing autonomy beyond human reaction speed. The discussion expands into compute bottlenecks, high-frame-rate sensing, data transfer limits, and why the future needs co-designed hardware and algorithms.
- •AlphaPilot builds on existing human FPV drone racing by adding full autonomy
- •Goal: autonomy that can fly faster than humans can process/react
- •Bottlenecks include both hardware and software; co-design is essential
- •Camera frame rate and data links (copper channel limits) constrain perception pipelines
- •Physical limits show up in chips (clock distribution, heat, speed-of-light constraints)
- 1:12:48 – 1:18:06
Crashes, VR testbeds, and beating humans: repeatability vs strategy
They discuss what it means to ‘finish’ and ultimately outperform humans, echoing lessons from DARPA challenges. Sertac highlights the importance of safe experimentation infrastructure (VR/AR + mocap), and argues machines win via repeatability and endurance while strategy remains harder.
- •Whether teams ‘finish’ depends on course difficulty and evaluation design
- •Iterated learning is accelerated when crashes are acceptable and cheap
- •VR/AR and motion-capture setups enable rapid testing without physical crashes
- •Machines can beat humans through consistency across repeated runs (no fatigue)
- •Humans retain advantages in higher-level strategy; machines excel in precision/repeatability
- 1:18:06 – 1:22:50
The most beautiful idea in robotics: Bellman’s equation and the curse of dimensionality
In closing, Sertac names Bellman’s equation as a foundational, beautiful concept in decision-making. He explains how it captures both elegance and the computational explosion that makes optimal planning hard, while still being surprisingly useful in practice.
- •Robotics splits into perception, decision-making/control, and interaction; Sertac focuses on decisions
- •Bellman’s equation underpins dynamic programming and reinforcement learning
- •It embodies the curse of dimensionality and computational hardness
- •Despite worst-case complexity, practical systems often work well
- •Ends with reflections on how much we still don’t know in math and the universe