a16zDecagon’s Playbook for Building Enterprise AI Applications
At a glance
WHAT IT’S REALLY ABOUT
Decagon’s enterprise AI playbook: open-source models, productized deployment, iterating fast
- Decagon migrated from frontier-model APIs to a stack that is ~90% open-source to achieve lower latency (especially for voice), better controllability, and cheaper inference at scale.
- They argue the usual “small model = worse performance” framing is a false trade-off: task-specific fine-tuning can make smaller models outperform frontier models on narrowly defined enterprise workflows while also being faster and cheaper.
- Decagon Labs functions as a continuous model factory, rapidly turning newly released base models into fine-tuned, production-ready components using highly customized evals tied to end customer outcomes.
- They differentiate in enterprise go-to-market by productizing deployment (testing, governance, integrations, iteration) and offering a “glass box” experience that lets customers build/iterate themselves rather than relying on a black-box FDE team.
- The company’s scope is expanding from customer support automation to an “AI concierge” that becomes the front door for all customer interactions (support, sales qualification, proactive operations), with job impacts framed via Jevons paradox—more demand emerges as service becomes cheaper.
IDEAS WORTH REMEMBERING
5 ideasLatency and controllability drive open-source adoption more than cost.
Decagon moved heavily to open source primarily to make voice and high-volume interactions fast and controllable; cost savings came as a secondary benefit once they decomposed tasks into smaller model calls.
“Smaller model = worse” is often wrong for enterprise workflows.
By fine-tuning and narrowing scope, Decagon reports smaller models can beat frontier models on specific sub-tasks (topic classification, policy checks, process steps) while also being faster and cheaper.
Evals are the real barrier to post-training in enterprises.
They emphasize fine-tuning is non-trivial because you need bespoke benchmarks tied to customer outcomes and system-level behavior, not generic public evals or loss curves.
Treat model training as a continuous factory, not a one-time project.
Because model capabilities change quickly, Decagon Labs is designed to compress the time from “new model released” to “useful fine-tuned production model,” with frequent retraining and deprecation.
Use frontier models where exploration and synthesis matter most.
They still rely on frontier models for broader, open-ended work—e.g., Duet Autopilot reviewing massive conversation corpora, generating variants, and running exploratory improvements—while small tuned models run the core paths.
WORDS WORTH SAVING
5 quotesAn AI agent should just be the front door of your business, and every interaction, whether it's like reactive or proactive with a customer, should be handled by AI.
— Jesse Zhang
I actually think that is a false trade-off.
— Ashwin Sreenivas
When we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific tasks we want them to do, they actually outperform the large, smart state-of-the-art models.
— Ashwin Sreenivas
And if you can't do that, then you're just building a glorified consulting firm.
— Ashwin Sreenivas
Most jobs are made up.
— Jesse Zhang
High quality AI-generated summary created from speaker-labeled transcript.