Skip to content
EO StudioEO Studio

The Problem with AI Agents No One is Talking About | Yutori, Abhishek Das

Most AI agents don't actually work. Abhishek Das, co-founder & co-CEO of Yutori, argues that the agent industry has quietly normalized unreliability. Even at 90% accuracy per step, errors compound fast across a 10, 20, or 50-step workflow, and the whole thing breaks. Backed by Fei-Fei Li and Jeff Dean, Abhishek is pushing back on that normalization. In this conversation, he lays out what it takes to ship agents that actually work on the first try, why taste matters more than code in the LLM era, and what Grad-CAM taught him about building "proof of work" into AI. 00:00 Find the Dopamine of Building 01:37 IIT Roorkee and SDS Labs 03:26 The Last Generation to Use a Browser 05:22 Stop Normalizing Broken Agents - What Separates Real Agents From Demos 07:30 The 80/20 rule 08:42 How to build taste - the weekly dogfooding ritual 09:27 Why Reliability Matters More Than Raw Performance EO stands for Entrepreneur& Opportunities. As we're looking to feature more inspiring stories of entrepreneurs all over the world, don't hesitate to contact us at partner@eoeoeo.net LinkedIn | @EO STUDIO X | @eostudi0

EO Studiohost
May 4, 202612mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Why unreliable AI web agents fail—and how to build better

  1. Most “do anything on the web” agents fail because small per-step error rates compound into very low end-to-end success on long, multi-step workflows.
  2. Das argues the industry is dangerously normalizing non-determinism and “slop,” and that agents should work on the first try to be considered shippable.
  3. Yutori focuses on robustness via comprehensive production evals, guardrails, and the ability for agents to recognize mistakes, backtrack, and recover.
  4. The product differentiator is shifting from quick prototypes to taste and craft, developed through rigorous internal dogfooding and an 80/20 focus on high-leverage features.
  5. Trust is built through reliability and “proof of work” transparency—showing users what the agent did (sites visited, steps taken)—a philosophy connected to Das’s interpretability background (Grad-CAM).

IDEAS WORTH REMEMBERING

5 ideas

Long workflows magnify small agent errors into frequent failures.

If an agent is 90% correct per step, a 10–50 step workflow quickly collapses in end-to-end success, which is why many impressive demos don’t hold up in real use.

“Usually works” is a bad standard for agentic products.

Das argues reliability shouldn’t be negotiated downward; if an agent can’t succeed on the first attempt most of the time, shipping it as a general solution erodes user trust and product credibility.

Recovery behavior matters as much as raw task competence.

Because agents will encounter unfamiliar websites and layouts, the critical capability is noticing mistakes and backtracking to try a different path, not pretending errors won’t happen.

Continuous evals in production are essential for improving agents.

Yutori runs each production query through comprehensive evaluations to pinpoint where agents succeed or fail and to prioritize which domains and behaviors need more work.

Web-agent generalization will always be incomplete—design for it.

No team can train on “every website,” so systems must expect novelty and variance and rely on guardrails, error detection, and correction loops to stay dependable.

WORDS WORTH SAVING

5 quotes

In this day and age, there's basically 100 different agent products out there that are saying that, like, this can do anything on the web, and you try it once, and it doesn't really work.

Abhishek Das

If we think of a 10-step, 20-step, or 50-step workflow, even if the accuracy at each step is, like, 90%, the 10% error rate compounds very quickly, and so the overall success rate of a task, of a workflow is, like, quite low, right?

Abhishek Das

Yeah, if it's not good enough to work on the first try, it's not good enough.

Abhishek Das

In a world where it's very easy to come up with first prototypes using these coding LLMs, the true differentiator is in taste and craft in how intuitive and well-designed the product is.

Abhishek Das

Whenever we ship something, we have to get it right. It has to really work. It has to be reliable. Users have to trust that it works well.

Abhishek Das

Compounding errors in long-horizon agent workflowsNormalization of unreliable agent demosBacktracking and self-correction as core agent capabilityProduction evals and guardrails for web agents80/20 prioritization and intuitive feature designDogfooding rituals to build product tasteTransparency/proof-of-work to build user trust

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.