Skip to content
YC Root AccessYC Root Access

Tavus: The AI Human Platform

Tavus is building real-time AI humans — systems that can see you, hear you, and respond with natural expression, emotion, and context. What began as personalized video has grown into a full platform used by companies from startups to the Fortune 10. The team recently raised a $40M Series B to advance this vision, introducing PALs: agentic AI humans that can perceive, reason, and act on their own. In this conversation with YC’s Diana Hu, founders Hassaan Raza and Quinn Favret share how they made the leap from generative video to real-time AI humans, the foundational models behind rendering and perception, and why they believe AI humans will become the next major interface for work and communication. Learn more about Tavus at https://www.tavus.io. Chapters: 00:24 – From Personalized Video to AI Humans 01:18 – Why Real-Time Matters 02:36 – How AI Humans See, Hear, and Respond 04:05 – Introducing PALs: Agentic AI Humans 05:42 – The Foundational Models Behind Tavus 07:28 – Building Emotion, Expression, and Context 09:10 – Use Cases From Startups to the Fortune 10 11:00 – Raising the $40M Series B 12:52 – The Future: AI Humans as the Next Interface

Diana HuhostHassaan RazaguestQuinn Favretguest
Nov 14, 202515mWatch on YouTube ↗

CHAPTERS

  1. 0:05 – 0:50

    Tavus in one line: building “AI humans” + a real-time demo

    Diana Hu opens by introducing Tavus and their new $40M Series B. Hassaan defines Tavus as an AI research lab focused on teaching machines to behave like humans, then they immediately jump into a quick demo conversation with an AI assistant.

    • Tavus positions itself as an AI human platform/research lab
    • Goal: machines that can see, hear, respond, act, and look human
    • Live demo setup: an always-on, agentic “PAL” you can call or text
    • Emphasis on naturalness—feels like a coworker/friend
  2. 0:50 – 0:58

    Why Tavus obsesses over real-time interaction

    The conversation highlights that the “magic” comes from being interactive in the moment, not pre-rendered or delayed. They frame low latency as a requirement for believable human conversation dynamics.

    • Real-time responsiveness is central to the experience
    • Believability depends on conversational timing, not just visuals
    • AI humans react to user expressions and emotions during the interaction
  3. 0:58 – 1:40

    Seeing the assistant work: schedule look-up + proactive email drafting

    A short clip shows an AI human taking a request, retrieving a schedule, and drafting an email automatically. The demo showcases agentic behavior (taking action) rather than only answering questions.

    • AI assistant reports today’s meetings and obligations
    • User asks to email a meeting delay; assistant drafts it immediately
    • Demonstrates task execution, not just conversation
    • Reinforces the “AI employee” framing
  4. 1:40 – 2:36

    Who’s using Tavus: from startups to Fortune 10 + the core use-case buckets

    They describe a wide customer base, including large enterprises, using Tavus models and interfaces to build AI employees. Quinn outlines three primary application buckets spanning training, healthcare, and go-to-market roles.

    • Customers range from startups to Fortune 10 (examples mentioned: Amazon, Alibaba, Better.com)
    • Three buckets: learning & development, healthcare, and go-to-market
    • Examples: patient intake, nutrition coach, elderly companionship, AI SDR/support
    • Positioning: customers build “AI employees” on Tavus
  5. 2:36 – 3:17

    Product form factor: multimodal AI humans delivered via SDK

    They clarify that Tavus is not just a video avatar: it can interact over video, audio, and text. Diana frames Tavus’s current go-to-market as an SDK that lets customers build custom AI employees.

    • AI human can video call, audio call, or text—multichannel presence
    • Must look/sound/behave human to feel like a coworker
    • Current model: SDK/API enabling customers to build tailored experiences
    • High customization across different customer needs
  6. 3:17 – 4:19

    Origins pre-ChatGPT: personalized video via lip-sync ‘infill’

    Hassaan recounts Tavus’s early days (2020–2021) when model capabilities were limited. The initial product focused on scaling personalized videos by generating lip-synced variants from a single recording.

    • Built in an era with far less capable generative models
    • Early technique: lip-sync infill to personalize a recorded message at scale
    • Use case: mass outbound personalization (e.g., changing names/attributes)
    • Over time, model advances enabled far more interactive capabilities
  7. 4:19 – 6:18

    The pivotal pivot: from AI sales tooling to research lab + SDK platform

    They describe a “crucible moment” after Series A: decide between being an AI sales company or committing to foundational human-computing models. They chose to churn existing traction and refocus on core technical DNA and platform direction.

    • Strategic fork: AI sales company vs. human-computing research platform
    • Decision driven by team DNA and belief in a larger vision
    • Churned customers/short-term traction to rebuild around SDK/API
    • Reflection on how external opinions can misdirect hiring and priorities
  8. 6:18 – 7:40

    Foundational models: rendering + perception as the ‘yin and yang’

    Diana frames Tavus’s stack as not only rendering human faces, but also perception—understanding expressions, gestures, and context. Hassaan argues the face alone is insufficient without teaching machines how humans perceive and react in conversation.

    • Two pillars: rendering (how the AI looks) and perception (how it understands you)
    • Human communication includes facial cues, gestures, and what’s unsaid
    • Collect and interpret nuanced signals (expression, timing, context)
    • Model the relationship between words, delivery, and reactions
  9. 7:40 – 8:04

    Latency as a hard requirement for human-like conversation

    They quantify the timing constraints of natural dialogue and explain why this is newly feasible. Fast back-and-forth is positioned as essential to achieving a “waltz” instead of robotic interaction.

    • Human conversational back-and-forth requires very low latency
    • They cite strong interaction happening in under ~200ms
    • Real-time performance unlocks more natural emotional responsiveness
  10. 8:04 – 9:20

    Introducing Tavus PALs: agentic, proactive AI humans for consumers

    Hassaan announces PALs as a product leap from SDK-only to a consumer/prosumer experience. They compare the moment to the transition from command-line computing to graphical interfaces—making AI accessible through a natural human interface.

    • PALs aim to bring AI humans to regular consumers and prosumers
    • Proactive + agentic: can act, reach out, and complete tasks
    • Multimodal: video, voice, and text interactions
    • Framed as a new interface layer for computing (like CLI → GUI)
  11. 9:20 – 11:26

    The future interface—and the concerns: jobs, access, and alignment

    They paint a future where everyone has an AI doctor/therapist/sidekick, then address worries about displacement. Hassaan argues the initial target is often replacing bad automated systems or filling gaps where no human service exists, such as inaccessible therapy.

    • Vision: AI humans as ubiquitous companions/assistants (Jarvis/Cortana analogies)
    • Acknowledges some job replacement risk rather than dismissing it
    • Focus on improving degraded experiences caused by existing automation
    • Access argument: AI can provide services where the alternative is nothing
  12. 11:26 – 13:14

    How Tavus builds empathy: perception signals, human simulation, and demos that teach users

    They describe empathy as a data and modeling problem: collecting subtle cues (micro-expressions) and linking them to conversational context. Quinn adds that the product experience and demos matter because users must relearn how to communicate naturally with machines.

    • Human conversation is nuanced—an ‘art’ and a ‘dance’
    • Perception models capture subtle signals (e.g., eyebrow twitch, slight smile)
    • Core work: model relationships between what was said and how it was expressed
    • Go-to-market lesson: great demos are required to convey the new interaction paradigm
  13. 13:14 – 15:19

    Founder lessons + closing: conviction, momentum, and ‘faster, faster’

    They end with advice to founders: develop conviction in the vision and drive momentum daily. The segment closes with their speed-focused culture and congratulations on the Series B.

    • Advice: build deep conviction and don’t over-index on outside opinions
    • Momentum is framed as the essential driver—something must move every day
    • Startup moat = speed; they aim to stay months ahead
    • Company motto emphasizes relentless pace; closing Series B congratulations

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.