a16zFormer Microsoft Executive Explains Where We Are in the AI Cycle w/ Anish Acharya & Steven Sinofsky
CHAPTERS
- 0:00 – 2:17
Why today’s AI feels like the 64K IBM PC era
Steven frames the current moment in AI as extremely early—akin to the “64K IBM PC” stage where fundamental constraints dominate and the hype outpaces what systems can reliably do. The chapter sets expectations: big promises (replacing Search/Excel) run into basic failures (errors, weak arithmetic, unreliability).
- •AI is positioned as a major platform shift, but still in a very early, constraint-driven phase
- •Hype claims (replace major software) conflict with present-day capability gaps
- •Early-platform energy goes into solving basic working problems rather than polished products
- 2:17 – 2:45
Karpathy’s “people spirits” and jagged intelligence: relearning how to use LLMs
Reacting to Andrej Karpathy’s talk, the group highlights the inversion in how we interact with LLMs versus traditional tools. The key idea is jagged intelligence: models can be brilliant in one moment and wrong in the next, requiring new habits to use them productively.
- •Karpathy’s metaphors help explain where we are and what to expect
- •LLMs behave like “people spirits” with uneven, jagged competence
- •Users must relearn workflows and expectations to get value from LLM tools
- 2:45 – 4:29
Tools come first: why vibe coding is visible—and why vibe writing may matter more
Steven argues that early platform eras are dominated by tooling, and developers are the first “customers,” making coding a natural early domain. He then proposes that vibe writing is already widespread and potentially more underestimated than vibe coding in near-term impact.
- •Early-stage platforms are defined by tools and developer-driven experimentation
- •Coding often leads early because developers adapt and iterate on tooling quickly
- •Vibe writing is already normalizing (school, business), echoing earlier tool shifts like calculators
- 4:29 – 6:17
Autonomy vs accountability: the hidden cost of ‘just ship the output’
The discussion sharpens around what ‘full autonomy’ really means when outcomes matter—grades, salaries, legal risk. Writing can look “done” even when it’s wrong, while code errors may surface later (security, auth, data handling), creating different failure modes.
- •Autonomous output is risky when correctness and accountability matter
- •People intuitively double-check math, but often fail to verify plausible-sounding prose
- •Vibe-coded software may pass demos but later fail via security/production issues
- 6:17 – 7:31
The Iron Man slider and the ‘decade of agents’
They credit Karpathy’s Iron Man analogy: autonomy is a spectrum controlled by a slider. Steven pushes back on short timelines, arguing agents will be a long, multi-year transition because automation repeatedly proves harder than expected.
- •Autonomy should be designed as a controllable spectrum, not a binary
- •“Year of agents” is hype; real progress likely spans a decade
- •Automation history shows many tasks resist full delegation longer than expected
- 7:31 – 8:16
Where agents work first: high-friction, low-judgment tasks (and where they won’t)
Anish introduces a practical 2x2: automation arrives fastest where tasks are high friction but require low judgment (e.g., refinancing shopping). He contrasts that with domains like taxes that combine friction with high judgment and high risk tolerance decisions.
- •Agent sweet spot: high friction + low judgment decisions
- •Example: shopping refinancing rates can be delegated more readily
- •Counterexample: taxes are high judgment and high risk, making full automation harder
- 8:16 – 10:48
Economics and differentiation: why ‘headless API life’ breaks down
Steven argues automation isn’t just a technical problem—it’s market structure and incentives. If providers can’t differentiate, advertise, and capture value, a fully automated “cheapest-option” ecosystem won’t sustain real businesses or meaningful consumer choice.
- •Automation must align with incentives for providers to compete and market offerings
- •Consumers rarely want pure lowest-price optimization; constraints/preferences matter
- •A world of faceless commodities is unrealistic because businesses need differentiation
- 10:48 – 15:13
AI+human vs AI alone: correctness, uncertainty, and exception-handling work
They distinguish domains with formal correctness (chess/Go) from domains driven by uncertainty and judgment (medicine, taxes). Steven emphasizes that many real jobs are dominated by edge cases and exceptions, which makes full automation elusive and pushes toward augmentation.
- •Formal correctness domains can progress to full autonomy; judgment-heavy domains often shouldn’t
- •Medicine illustrates pervasive uncertainty; radiologists adopt AI like any other instrument upgrade
- •Taxes exemplify cascading exceptions—automation requires resolving the same judgment calls humans make
- 15:13 – 15:55
Why product management persists: ambiguity as the real job
Anish addresses fears about the ‘death of PM,’ arguing PM work is fundamentally about resolving ambiguity in complex adaptive organizations. Even with better tools, companies will still need roles that coordinate decisions, trade-offs, and clarity under uncertainty.
- •PM value is ambiguity reduction, not just writing docs or managing tickets
- •Organizations are complex adaptive systems that continuously generate uncertainty
- •AI may change the workflow, but the need for judgment-heavy coordination remains
- 15:55 – 18:22
Vibe coding for clout vs production reality: prompts as programming
Steven critiques social-media demos that imply effortless text-to-app creation, arguing many are exaggerated and fragile. He claims that prompting is effectively programming, and making it reliable often means adding structure—i.e., inventing new programming languages.
- •Public platform transitions amplify hype and performative demos
- •Many ‘it worked instantly’ stories hide significant struggle and brittleness
- •Prompting is a form of programming; adding structure becomes a new language/toolchain
- 18:22 – 23:07
Programming language cycles: overpromise, underdeliver—and what’s different this time
They compare today’s moment to prior language/tool hype cycles (OO, low-code, Delphi/PowerBuilder, templates). Steven notes past waves improved constants rather than changing the order of magnitude; he argues writing may be experiencing a true step-change despite new error types.
- •Historical tech waves repeatedly promised to democratize programming dramatically
- •Most innovations improved productivity by constants, not orders of magnitude
- •Writing may be changing by an order of magnitude, though it introduces new categories of mistakes
- 23:07 – 24:30
AI creative writing: slop, ceilings, and best-selling AI-assisted novels
Steven predicts best-selling novels will be largely AI-generated or AI-assisted, likely with human editing and delayed disclosure. Anish questions whether averaging models can reach cultural ‘edges,’ and both discuss how tools lower floors (more output) while potentially raising ceilings (new native creators).
- •AI-assisted novels and mainstream success are likely, with authors editing outputs
- •Models trend toward averages; great art often comes from cultural edges—tooling must steer away from the mean
- •‘Slop’ increases volume, but the bigger question is how tools raise the creative ceiling
- 24:30 – 28:02
Access reshapes standards: from dot-matrix papers to AI-generated business content
Steven argues that most business writing is already mediocre, making LLMs immediately useful for practical content like enterprise case studies. They broaden the lens: for people with no access to expertise (e.g., medical advice), “better than nothing” can be transformative—even if imperfect.
- •LLMs can outperform typical business/marketing boilerplate at much lower effort
- •Debates over excellence often ignore that most real-world content is average
- •Access and scalability can justify ‘good enough’ outputs where alternatives are absent
- 28:02 – 30:43
Google, I/O, and platform transitions: ‘death’ vs losing influence
They close by rejecting the idea that Google is ‘dying,’ emphasizing that large incumbents can deploy “shock and awe” across the stack. The real test is whether Google can change product-building and go-to-market assumptions beyond search-and-ads context to stay influential through the transition.
- •“Company death” narratives are misguided; the real risk is loss of influence in transitions
- •Incumbents can mount comprehensive, well-resourced responses quickly
- •Key question: can Google transform product thinking and go-to-market beyond its legacy model?