a16zHow Jev Turns AI Into Software That Gets Things Done
CHAPTERS
- 0:00 – 1:17
“Where the automation?”—why AI feels smart but doesn’t run the world
Diogo and Martin open with a blunt frustration: despite rapid progress in AI, day-to-day business workflows remain largely unautomated. They contrast impressive demos (chatbots, coding assistants) with the lack of dependable systems that can take real action without constant supervision.
- •AI appears highly intelligent yet underdelivers on practical automation
- •Coding agents may speed output but don’t necessarily improve software quality
- •The real gap is turning AI capability into trustworthy, action-taking systems
- •Motivating question: why aren’t routine tasks (like support) truly automated yet
- 1:17 – 3:00
Meet Diogo, TypeSafe, and Jev: “AI for software,” not humans-in-the-loop
Ben and Martin introduce Diogo and ask for the Jev/TypeSafe pitch. Diogo frames TypeSafe’s goal as making AI a native building block inside software so automation becomes feasible, rather than building yet another assistant that only helps a human operator.
- •TypeSafe aims to make AI usable as a software primitive
- •Jev is presented as a first model/product in this direction
- •Focus shifts from “AI for people” to “AI inside programs”
- •Automation requires systems you can run in the background safely
- 3:00 – 4:25
Just-in-time code vs “smart software”: expanding what software can express
Diogo distinguishes tools like Codex/Cloud Code/Cursor—which generate conventional code—from an approach that increases what software itself can do. The ambition is to represent intent and higher-level meaning so programs can handle tasks that are hard to fully specify up front.
- •Coding tools generate traditional code; Jev targets new software capabilities
- •Goal: express intent and extend the vocabulary of programming
- •Programming as ‘hyper-specification’ is powerful but limited for messy tasks
- •Vision: software that can reliably interpret and act on higher-level goals
- 4:25 – 6:15
A new primitive inside code: natural language + state machines + confidence
Martin reframes Jev as an “intelligent layer” programmers embed into systems—like a library that takes natural-language intent and interacts with structured program state. The conversation highlights that this introduces probabilistic behavior in a way typical software engineering hasn’t had at scale.
- •Jev as a code-embedded primitive, not a code generator
- •Combines natural-language descriptions with structured state machine outputs
- •Supports confidence/probability-aware decisions
- •Represents a different mental model for programmers than deterministic APIs
- 6:15 – 7:41
Pragmatic ML interface design: “Jev is absolutely a classifier”
Diogo embraces the idea that Jev is a classifier and argues that ‘useful’ ML interfaces are what matter for software. He describes the craft of operating in the overlap between what AI is good at and what code needs, emphasizing practical constraints over hype.
- •Machine-native outputs must align with useful software abstractions
- •Probabilities/classification are not a weakness; they’re practical primitives
- •Jev aims to outperform the need for bespoke 2019-style MLE teams
- •Early framing: “diamond in the rough” intelligence needs productization
- 7:41 – 9:01
Design space and trade-offs: slider between language-in/out and programs
Martin asks whether the future is a continuum between natural language and imperative programming. Diogo argues it’s a slider governed by constraints like cost, speed, and “intelligence per dollar,” with Jev intentionally designed to sit inside program state.
- •Future likely spans a continuum from natural language to programmatic control
- •TypeSafe optimizes for intelligence-per-dollar (with speed as a competing axis)
- •Interface choices (e.g., “input state”) signal intended embedding in programs
- •Near-term reality: AI may behave more like a ‘smart database’ than a library
- 9:01 – 12:18
Diogo’s background: systems thinking meets AI research
Diogo recounts an unconventional path—from math competitions to CS and a Kaggle win driven by heavy automation—into the AI research world. He describes moving through NeurIPS exposure, startups, Google Brain, and eventually OpenAI, retaining a systems-builder mindset throughout.
- •Mathlete roots; preference for CS as ‘useful, fun math’
- •Kaggle success came from systems-style automation rather than theory
- •NeurIPS exposure and mentorship pulled him into mainstream AI circles
- •Career stops: startup (Jeremy Howard), Google Brain, then OpenAI
- 12:18 – 21:10
“Build prod, not God”: optimism, developer reality, and the anti-hype stance
Ben highlights TypeSafe’s joyful, pro-progress posture versus broader AI doom narratives. Diogo argues many critiques come from “monomodel” thinking and miss the developer perspective: the core problem isn’t superintelligence, it’s that basic automation still doesn’t work reliably.
- •TypeSafe ethos: focus on production utility over grandiose AGI theater
- •Skepticism toward ‘one big brain to rule them all’ framing
- •Developer experience reveals a disconnect between hype and real automation
- •Concern: industry optimizes for impressive demos rather than dependable work
- 21:10 – 25:32
Why automation lags: long tail, but also misplaced evaluation targets
The discussion drills into why customer support and internal workflows remain stubbornly manual. Diogo downplays ‘it’s just data’ as the main blocker and argues the industry optimized for outputs humans judge (chat quality) rather than automation metrics (task completion reliably).
- •Support automation looks good on volume but fails on uniqueness/edge cases
- •Diogo rejects a pure ‘data distribution’ explanation as sufficient
- •Benchmark should be pragmatic: automate the obvious, high-ROI tasks first
- •Core critique: optimizing human judges (RLHF) over end-to-end automation
- 25:32 – 28:07
Reliability as the product: robustness beyond determinism
Diogo explains that TypeSafe’s differentiator is reliability—defined not as uptime, but as consistent competence. He distinguishes determinism, robustness to irrelevant prompt changes, and a deeper ‘always smart’ behavior developers can program against without endless examples.
- •Reliability ≠ uptime; closer to consistent ‘intelligence’ under variation
- •Determinism helps tests, but real systems need robustness
- •Goal: developers program against Jev without prompt babysitting
- •Each additional ‘nine’ of reliability unlocks new applications
- 28:07 – 30:00
Coding agents vs Jev: syntax help is easy; architecture is the hard part
Martin and Diogo compare Jev’s embedded primitive with coding agents. Diogo argues agents are strong at syntax but weak at semantics and architecture, and the right trade-off depends on whether speed matters more than design quality for a given project.
- •Coding agents excel at code generation mechanics, struggle with architecture
- •Architecture remains a key human creative advantage (today)
- •Trade-offs: accept ‘median’ architecture for speed in some contexts
- •Jev complements agents by expanding what software can do, not just how fast it’s written
- 30:00 – 32:30
From “SaaSpocalypse” to “inverse SaaSpocalypse”: why SaaS wins with new primitives
Ben notes the market narrative flip: coding agents threatened SaaS values, but Jev excited SaaS builders. Diogo argues SaaS companies are positioned to win because they know workflows deeply and can invest capex to embed new capabilities across large user bases.
- •SaaS isn’t easily replicable; value lives ‘beneath the hood’
- •SaaS companies best understand real workflows worth automating
- •AI enables SaaS to become dramatically more useful (beyond bolted-on chatbots)
- •Vision: AI primitives amplify distribution advantages and product depth
- 32:30 – 38:20
New capabilities and interfaces: “do what I mean,” disappearing forms, and voice control
They explore how Jev-like primitives could transform user interfaces, reducing rigid multi-choice forms and enabling systems that infer intent. Diogo shares an example of voice-driven computing where the model decides whether speech is a command or text entry—hinting at a shift toward DWIM software.
- •Potential decline of rigid UI patterns like legacy forms/menus
- •‘Do what I mean’ as a next-level interaction contract
- •Example: voice control that classifies intent (command vs dictation) dynamically
- •Future may require improvements in cost/latency for real-time interfaces
- 38:20 – 40:53
Apps vs the guts of systems: intelligence will live deep in infrastructure
Diogo predicts most AI calls in a mature world won’t be for human-facing prose but for deep system behavior—routing, decisions, and internal automation. Martin describes how AI previously didn’t ‘fit’ into software (schema-following failures), and they frame Jev’s approach as a bridge between AI and stateful systems.
- •Most future AI usage likely happens ‘in the guts,’ not the UI layer
- •Economic revolution implies massive machine-to-machine AI calls
- •Past failure mode: prompts/schemas don’t reliably bind model output to programs
- •Jev’s state-machine mapping offers a more software-native integration path
- 40:53 – 42:24
Closing: the anti-frustration machine and the fight for trustworthy DWIM software
They end by returning to the central aspiration: software that reliably does what you mean. Diogo emphasizes restraint about overpromising, while committing to the long-term reliability work required for real-world automation and composable systems.
- •Positioning: Jev as an ‘anti-frustration’ primitive for builders
- •DWIM as a practical north star rather than sci-fi speculation
- •Acknowledgment: not ready for every overpromised use case—yet
- •Commitment: reliability-first iteration to unlock real automation