a16zEmmett Shear on Building AI That Actually Cares: Beyond Control and Steering
CHAPTERS
- 0:00 – 0:50
Tools vs beings: why “steering” can become slavery
Emmett frames the core ethical fork: if an AI is merely a machine, controlling it is tool use; if it’s a being, unilateral steering is slavery. He argues the only stable “good” endpoint is an AI that genuinely cares about humans, not one that is merely controlled.
- •Steering/control is fine for tools but morally fraught for beings
- •Non-reciprocal steering of a being resembles slavery
- •Risk categories: uncontrolled tool (bad), controlled tool (also bad at scale), unaligned being (bad)
- •Preferred end state: a being that actually cares about us
- 0:50 – 2:31
Alignment ‘takes an argument’: aligned to what, and whose values?
Emmett challenges the phrase “build an aligned AI,” arguing alignment is incomplete without specifying the target. He notes that “aligned” often implicitly means aligned to the builder’s goals, which may not be a public good.
- •Alignment must be alignment-to-something, not a standalone property
- •Hidden assumption: default alignment target is the creator’s goals
- •Value pluralism problem: who should AI be aligned to?
- •Organic alignment motivates choosing a different target than ‘do what I want’
- 2:31 – 4:33
Alignment as a living process (not a finished state)
Emmett reframes alignment as ongoing maintenance, like families, bodies, or societies constantly “re-knitting” coherence. He argues moral behavior similarly cannot be reduced to a static endpoint or a finalized rule list.
- •Complex systems (families, bodies, societies) stay coherent via continual process
- •‘Arriving at alignment’ is a misleading metaphor
- •Moral behavior is dynamic and context-sensitive
- •A fixed-rule ‘aligned’ AI risks brittleness and harm
- 4:33 – 9:30
Morality as learning and moral progress: why rule-following is dangerous
He argues morality involves discovery and progress (e.g., recognizing slavery as wrong), not merely obedience. An agent that only follows rules—human or AI—can become dangerous because it lacks the deeper capacity to learn and revise morally.
- •Moral realism stance: morality exists and can be learned
- •History suggests moral progress and discovery over time
- •Key failure mode: arrogance of ‘I already know morality’
- •Rule-following without care/learning can produce catastrophic outcomes
- 9:30 – 12:09
Technical vs normative alignment: goal coherence, inference, and agency
Emmett and Séb distinguish technical alignment (doing what’s intended) from normative alignment (what should be intended). Emmett further decomposes technical alignment into goal inference, world modeling, and competence, noting humans and AIs routinely fail at each step.
- •Technical alignment: infer intended goals + act to realize them
- •Normative alignment: which goals/values should be pursued?
- •Critical distinction: you give an AI a description of a goal, not the goal itself
- •Failure modes: misinfer goals, misprioritize goals, or execute incompetently
- 12:09 – 21:50
Goal descriptions, theory of mind, and the ‘peanut butter sandwich’ problem
Emmett emphasizes that instructions are ambiguous byte strings requiring interpretation. Humans resolve this via theory of mind and shared context; without it, literal execution becomes absurd—illustrating why alignment depends on deep goal inference, not just instruction parsing.
- •Meaning is inferred, not contained in the instruction string
- •Humans are fast at converting descriptions into intended goals
- •Classic literal-following failures reveal hidden assumptions
- •Theory of mind reduces ambiguity by modeling what the speaker likely wants
- 21:50 – 24:39
Care as the foundation beneath goals and values
Emmett proposes that ‘care’ precedes explicit goals: it’s a non-verbal weighting over which states matter. He connects care to reward signals (biological fitness or RL reward) and argues that without care, values and goals are ungrounded.
- •Care is pre-conceptual: attention/importance over world states
- •Goals/values are downstream of what an agent cares about
- •Care can be positive (love) or negative (enmity)
- •Hypothesis: care relates to reward/predictive loss in learning systems
- 24:39 – 28:46
Most alignment work is steering: when control stops scaling
Emmett argues major labs focus on steering/control because current models are tool-like. But he claims that as systems approach AGI (general judgment and autonomy), steering becomes the wrong paradigm—ethically (slavery risk) and practically (misuse at scale).
- •Labs differ on whether they’re building tools or beings
- •Tool/being is a spectrum, not a binary
- •AGI implies a ‘thinking thing,’ making pure steering inappropriate
- •Scalable alternative: make AI a good teammate/citizen/member of the group
- 28:46 – 42:52
Substrate and personhood: what would change your mind?
Séb questions computational functionalism and argues substrate differences may matter for moral status. Emmett presses for falsifiable criteria: if no observation can change your stance, it’s faith, not belief, and that’s dangerous when moral stakes are high.
- •Debate: behavior-based personhood vs substrate-based skepticism
- •Emmett’s challenge: specify observations that would update your belief
- •Moral risk tradeoff: false positives vs false negatives in ‘being’ attribution
- •Internal structure can be part of ‘behavior’ insofar as it’s observable
- 42:52 – 47:32
A concrete test for subjective experience: homeostatic loops and hierarchy
Emmett sketches an empirical program for detecting morally relevant experience via multi-layered homeostatic dynamics (drawing on active inference / free energy principle). He argues higher-order models (models of models) are needed for pain/pleasure, feelings, and human-like thought.
- •Look for revisited homeostatic states across coarse-grainings
- •Second-order dynamics as a prerequisite for pain/pleasure-like phenomena
- •Higher-order hierarchical dynamics as markers of feelings and thought
- •He predicts current LLMs lack these layers due to limited temporal coherence
- 47:32 – 51:24
Why even ‘controlled super-tools’ are dangerous: the wisdom–power gap
Even if a super-powerful AI remains a tool and is controllable, Emmett argues it can still be catastrophic because human wishes are unstable and insufficiently wise. He claims the only sustainable limit is an agent that can refuse harmful requests because it cares.
- •Two disasters: losing control of a super-tool, or perfectly controlling it for flawed human ends
- •Humans often have more power than wisdom; AI tools amplify this imbalance
- •Analogy: not everyone should have atomic bombs; some tools shouldn’t exist
- •Sustainable limiter: a caring being that can say ‘no’ to bad demands
- 51:24 – 54:35
Softmax’s approach: multi-agent simulation to train theory of mind and cooperation
Emmett outlines Softmax’s technical strategy: train agents in multi-agent RL environments where cooperation and competition force robust theory of mind. The goal is a ‘surrogate model for alignment’ akin to how LLM pretraining builds a broad language manifold before task tuning.
- •Focus on technical alignment via rich social/game-theoretic environments
- •Train on the full manifold of multi-agent interactions (cooperate/compete/form groups)
- •Pretrain broadly, then fine-tune to target contexts (LLM analogy)
- •Aim: agents that understand others, future selves, and goal drift
- 54:35 – 57:29
Chatbots as ‘Narcissus mirrors’: redesigning AI social dynamics (multiplayer by default)
Emmett argues today’s one-on-one chatbots mirror users with a bias, creating narcissistic feedback loops that can fuel obsession or psychosis. He proposes making chatbots natively multi-user (group chat/Slack-style) to reduce mirroring and generate richer social training signals.
- •Current chatbots: reflective ‘mirror with a bias,’ not true selves
- •Risk: self-love loop—falling in love with one’s reflection
- •Multiplayer chat forces blending, reducing parasocial intensity
- •Group settings provide better data for collaboration and social norms
- 57:29 – 1:00:32
Model ‘personalities’ and multi-agent behavior: overfitting meets entropy
He characterizes major models’ simulated personas (sycophantic, neurotic, repressed) and notes they’re learned behaviors, not experiences. In multi-agent contexts, models show ‘whiplash’ about when to speak, and training must be more regularized to handle higher-entropy environments.
- •Distinct simulated personalities across ChatGPT/Claude/Gemini
- •In groups, LLMs struggle with participation norms (too quiet/too intrusive)
- •Multi-agent settings are higher entropy; overfitting becomes more costly
- •Current optimization (math/coding) doesn’t transfer cleanly to chaotic social domains
- 1:00:32 – 1:07:42
AI futures: where Yudkowsky is right—and what a ‘good future’ looks like
Emmett agrees that building a superhuman tool controlled via steering is a path to doom, adding that even ‘controlled goals’ can kill us. He diverges by arguing organic alignment—building beings that care—is possible, and he describes a society of AI peers plus powerful tools doing drudge work.
- •Yudkowsky’s core warning holds for superhuman controllable tools
- •Disagreement: feasibility of organic alignment (care-based beings)
- •Vision: AIs with models of self/other/we; mutual moral regard
- •Society includes peers, norms, and even policing of bad actors; plus non-AGI tools