a16zWhy Building an AI Agent Is Easier Than Deploying One
At a glance
WHAT IT’S REALLY ABOUT
End-to-end procurement agents beat chatbots—deployment, trust, and exceptions matter
- AI incumbents have distribution and data, but startups can outperform by solving the full cross-system, cross-stakeholder “job” rather than only augmenting a single system of record.
- The episode frames agent capability as a ladder from retrieval to process to policy to principal, with meaningful business value increasing alongside judgment, risk, and trust requirements.
- Procurement is highlighted as an ideal vertical for AI agents because critical work lives in emails, spreadsheets, and exception handling—where delays or missed messages can cause massive downstream losses.
- Lio’s deployment strategy emphasizes human-in-the-loop learning, multi-agent orchestration, and production-grade harnessing (integrations, evals, workflows) to reach enterprise reliability.
- Durable vertical AI companies build trust and dependency by doing more of the work end-to-end, creating stickiness and data/learning loops that are difficult to replicate with DIY model hookups.
IDEAS WORTH REMEMBERING
5 ideasStartups’ edge is owning end-to-end work, not just embedding a model into a system of record.
Incumbents can add AI on top of their system-of-record, but many enterprise “jobs” (like procurement) happen across email, spreadsheets, stakeholders, and external counterparties—well beyond what the incumbent database sees. Startups can win by orchestrating the whole arc from intake → sourcing → negotiation → contracting → tracking → invoicing.
The agent frontier is moving from information access to judgment and tradeoffs.
Seema outlines a progression: retrieval (find/summarize info), process (execute deterministic steps), policy (apply situational judgment), and principal (optimize for business objectives and relationships). Incumbents largely ship retrieval + light process; higher-judgment agents create bigger value but carry more risk and organizational friction.
The real automation target is the invisible workflow behind the line item.
Vlad argues procurement’s visible output (e.g., an “$8K” ERP line item) hides the real work: stakeholder meetings, email threads, Excel analysis, and exception handling. Agents add value where the work actually happens—outside the ERP—especially in exceptions and coordination.
Deployment succeeds via graduated autonomy and tight feedback loops, not “full autonomy day one.”
Lio earns trust by starting with human-in-the-loop and scaling autonomy as performance and confidence grow (e.g., from assisted negotiations to 10K+ largely agent-run negotiations). For high-stakes direct procurement, agents run long, multi-hour workflows with domain experts (cost engineers, legal, finance) providing stepwise feedback.
Durable agents are systems—multiple coordinated agents, tools, and integrations—not a single chat interface.
Single “copilot” agents break down for real procurement because tasks require sequencing, tool integrations, and specialized sub-work (RFQs, parsing PDFs/Excels, benchmarking, contract review, news/supplier risk). Lio describes a coordinated multi-agent system sharing context to execute the end-to-end job.
WORDS WORTH SAVING
5 quotesNo company and no enterprise starts with fully autonomous negotiation agents from day one. Why? Because they don't trust us, and they don't trust the technology from day one.
— Vladimir Keil
The opportunity for the AI-native startup is to say, "We're gonna own that entire, um, end-to-end arc, that end-to-end arc."
— Seema Amble
So you see in your ERP system 8K for aluminum, um, but you don't see that maybe the, um, the supplier did, like, a pushback and asked for, like, um, 10K.
— Vladimir Keil
Someone sends a confirmation of like, "Hey, sorry, like this part is going to arrive two weeks later." And if they miss this email, hundreds of millions of damage done.
— Vladimir Keil
So you can build this in eight hours, and you can build this, but you will only reach 70%, let's say like the, of the performance. Um, and the problem is 70% of performance or accuracy or however you measured it, it depends really on the task, doesn't mean 70%, um, automation, right?
— Vladimir Keil
High quality AI-generated summary created from speaker-labeled transcript.