Skip to content
a16za16z

Decagon’s Playbook for Building Enterprise AI Applications

Sarah Wang and Kimberly Tan are joined by Jesse Zhang and Ashwin Sreenivas, co-founders of Decagon, to discuss the evolution of enterprise AI agents, why the company increasingly relies on open-source models, and how it is helping some of the world’s largest companies deploy AI in production. Decagon has become one of the fastest-growing AI companies by building agents that automate customer support, sales, and operational workflows. Jesse, Decagon’s CEO, and Ashwin, its president, explain how the company is building enterprise AI at scale. They unpack why Decagon moved most of its inference to open-source models, how latency, evaluation, and fine-tuning shape production AI systems, and why enterprise AI requires far more than simply plugging into frontier models. The conversation also explores forward-deployed engineering, enterprise sales, AI’s impact on jobs, and why application companies will continue to thrive alongside the foundation model labs. Timestamps: 00:00 - Intro 01:07 - Decagon's Journey from Frontier APIs to 90% Open-Source 05:00 - The False Trade-off: Why Fine-Tuned Small Models Win 09:26 - Decagon Labs as a Model Factory 15:07 - Are Frontier AI Labs the Last Startups? 21:21 - The Forward Deployed Trap: Product vs Consulting Truck 28:36 - Duet Autopilot: The Agent That Builds the Agent 37:02 - Winning Enterprise: Glass Box vs Black Box (and Beating Sierra) 47:55 - From Customer Support to AI Concierge 01:14:45 - Will AI Kill Jobs? Jevons Paradox in Customer Support Resources: Follow Jesse Zhang on X: https://x.com/thejessezhang Follow Ashwin Sreenivas on X: https://x.com/AshwinSreenivas Follow Sarah Wang on X: https://x.com/sarahdingwang Follow Kimberly Tan on X: https://x.com/kimberlywtan Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends! Find a16z on X: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711 Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.

Jesse ZhangguestSarah WanghostAshwin SreenivasguestKimberly Tanhost
Jul 31, 20261h 20mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Decagon’s enterprise AI playbook: open-source models, productized deployment, iterating fast

  1. Decagon migrated from frontier-model APIs to a stack that is ~90% open-source to achieve lower latency (especially for voice), better controllability, and cheaper inference at scale.
  2. They argue the usual “small model = worse performance” framing is a false trade-off: task-specific fine-tuning can make smaller models outperform frontier models on narrowly defined enterprise workflows while also being faster and cheaper.
  3. Decagon Labs functions as a continuous model factory, rapidly turning newly released base models into fine-tuned, production-ready components using highly customized evals tied to end customer outcomes.
  4. They differentiate in enterprise go-to-market by productizing deployment (testing, governance, integrations, iteration) and offering a “glass box” experience that lets customers build/iterate themselves rather than relying on a black-box FDE team.
  5. The company’s scope is expanding from customer support automation to an “AI concierge” that becomes the front door for all customer interactions (support, sales qualification, proactive operations), with job impacts framed via Jevons paradox—more demand emerges as service becomes cheaper.

IDEAS WORTH REMEMBERING

5 ideas

Latency and controllability drive open-source adoption more than cost.

Decagon moved heavily to open source primarily to make voice and high-volume interactions fast and controllable; cost savings came as a secondary benefit once they decomposed tasks into smaller model calls.

“Smaller model = worse” is often wrong for enterprise workflows.

By fine-tuning and narrowing scope, Decagon reports smaller models can beat frontier models on specific sub-tasks (topic classification, policy checks, process steps) while also being faster and cheaper.

Evals are the real barrier to post-training in enterprises.

They emphasize fine-tuning is non-trivial because you need bespoke benchmarks tied to customer outcomes and system-level behavior, not generic public evals or loss curves.

Treat model training as a continuous factory, not a one-time project.

Because model capabilities change quickly, Decagon Labs is designed to compress the time from “new model released” to “useful fine-tuned production model,” with frequent retraining and deprecation.

Use frontier models where exploration and synthesis matter most.

They still rely on frontier models for broader, open-ended work—e.g., Duet Autopilot reviewing massive conversation corpora, generating variants, and running exploratory improvements—while small tuned models run the core paths.

WORDS WORTH SAVING

5 quotes

An AI agent should just be the front door of your business, and every interaction, whether it's like reactive or proactive with a customer, should be handled by AI.

Jesse Zhang

I actually think that is a false trade-off.

Ashwin Sreenivas

When we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific tasks we want them to do, they actually outperform the large, smart state-of-the-art models.

Ashwin Sreenivas

And if you can't do that, then you're just building a glorified consulting firm.

Ashwin Sreenivas

Most jobs are made up.

Jesse Zhang

Open source vs closed source model strategyLatency, cost, intelligence trade-space for production agentsFine-tuning small models for task dominanceDecagon Labs as continuous training/evals pipelineDuet and Duet Autopilot (agent-building agent)Forward deployed engineering vs scalable productizationEnterprise selling: governance, rollout, and “glass box” controlConcierge vision: support → sales → ops workflowsMoat in an AGI world: enterprise deployment infrastructureJobs narrative: Jevons paradox in customer support demand

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.