CHAPTERS
- 0:05 – 0:50
Tavus in one line: building “AI humans” + a real-time demo
Diana Hu opens by introducing Tavus and their new $40M Series B. Hassaan defines Tavus as an AI research lab focused on teaching machines to behave like humans, then they immediately jump into a quick demo conversation with an AI assistant.
- •Tavus positions itself as an AI human platform/research lab
- •Goal: machines that can see, hear, respond, act, and look human
- •Live demo setup: an always-on, agentic “PAL” you can call or text
- •Emphasis on naturalness—feels like a coworker/friend
- 0:50 – 0:58
Why Tavus obsesses over real-time interaction
The conversation highlights that the “magic” comes from being interactive in the moment, not pre-rendered or delayed. They frame low latency as a requirement for believable human conversation dynamics.
- •Real-time responsiveness is central to the experience
- •Believability depends on conversational timing, not just visuals
- •AI humans react to user expressions and emotions during the interaction
- 0:58 – 1:40
Seeing the assistant work: schedule look-up + proactive email drafting
A short clip shows an AI human taking a request, retrieving a schedule, and drafting an email automatically. The demo showcases agentic behavior (taking action) rather than only answering questions.
- •AI assistant reports today’s meetings and obligations
- •User asks to email a meeting delay; assistant drafts it immediately
- •Demonstrates task execution, not just conversation
- •Reinforces the “AI employee” framing
- 1:40 – 2:36
Who’s using Tavus: from startups to Fortune 10 + the core use-case buckets
They describe a wide customer base, including large enterprises, using Tavus models and interfaces to build AI employees. Quinn outlines three primary application buckets spanning training, healthcare, and go-to-market roles.
- •Customers range from startups to Fortune 10 (examples mentioned: Amazon, Alibaba, Better.com)
- •Three buckets: learning & development, healthcare, and go-to-market
- •Examples: patient intake, nutrition coach, elderly companionship, AI SDR/support
- •Positioning: customers build “AI employees” on Tavus
- 2:36 – 3:17
Product form factor: multimodal AI humans delivered via SDK
They clarify that Tavus is not just a video avatar: it can interact over video, audio, and text. Diana frames Tavus’s current go-to-market as an SDK that lets customers build custom AI employees.
- •AI human can video call, audio call, or text—multichannel presence
- •Must look/sound/behave human to feel like a coworker
- •Current model: SDK/API enabling customers to build tailored experiences
- •High customization across different customer needs
- 3:17 – 4:19
Origins pre-ChatGPT: personalized video via lip-sync ‘infill’
Hassaan recounts Tavus’s early days (2020–2021) when model capabilities were limited. The initial product focused on scaling personalized videos by generating lip-synced variants from a single recording.
- •Built in an era with far less capable generative models
- •Early technique: lip-sync infill to personalize a recorded message at scale
- •Use case: mass outbound personalization (e.g., changing names/attributes)
- •Over time, model advances enabled far more interactive capabilities
- 4:19 – 6:18
The pivotal pivot: from AI sales tooling to research lab + SDK platform
They describe a “crucible moment” after Series A: decide between being an AI sales company or committing to foundational human-computing models. They chose to churn existing traction and refocus on core technical DNA and platform direction.
- •Strategic fork: AI sales company vs. human-computing research platform
- •Decision driven by team DNA and belief in a larger vision
- •Churned customers/short-term traction to rebuild around SDK/API
- •Reflection on how external opinions can misdirect hiring and priorities
- 6:18 – 7:40
Foundational models: rendering + perception as the ‘yin and yang’
Diana frames Tavus’s stack as not only rendering human faces, but also perception—understanding expressions, gestures, and context. Hassaan argues the face alone is insufficient without teaching machines how humans perceive and react in conversation.
- •Two pillars: rendering (how the AI looks) and perception (how it understands you)
- •Human communication includes facial cues, gestures, and what’s unsaid
- •Collect and interpret nuanced signals (expression, timing, context)
- •Model the relationship between words, delivery, and reactions
- 7:40 – 8:04
Latency as a hard requirement for human-like conversation
They quantify the timing constraints of natural dialogue and explain why this is newly feasible. Fast back-and-forth is positioned as essential to achieving a “waltz” instead of robotic interaction.
- •Human conversational back-and-forth requires very low latency
- •They cite strong interaction happening in under ~200ms
- •Real-time performance unlocks more natural emotional responsiveness
- 8:04 – 9:20
Introducing Tavus PALs: agentic, proactive AI humans for consumers
Hassaan announces PALs as a product leap from SDK-only to a consumer/prosumer experience. They compare the moment to the transition from command-line computing to graphical interfaces—making AI accessible through a natural human interface.
- •PALs aim to bring AI humans to regular consumers and prosumers
- •Proactive + agentic: can act, reach out, and complete tasks
- •Multimodal: video, voice, and text interactions
- •Framed as a new interface layer for computing (like CLI → GUI)
- 9:20 – 11:26
The future interface—and the concerns: jobs, access, and alignment
They paint a future where everyone has an AI doctor/therapist/sidekick, then address worries about displacement. Hassaan argues the initial target is often replacing bad automated systems or filling gaps where no human service exists, such as inaccessible therapy.
- •Vision: AI humans as ubiquitous companions/assistants (Jarvis/Cortana analogies)
- •Acknowledges some job replacement risk rather than dismissing it
- •Focus on improving degraded experiences caused by existing automation
- •Access argument: AI can provide services where the alternative is nothing
- 11:26 – 13:14
How Tavus builds empathy: perception signals, human simulation, and demos that teach users
They describe empathy as a data and modeling problem: collecting subtle cues (micro-expressions) and linking them to conversational context. Quinn adds that the product experience and demos matter because users must relearn how to communicate naturally with machines.
- •Human conversation is nuanced—an ‘art’ and a ‘dance’
- •Perception models capture subtle signals (e.g., eyebrow twitch, slight smile)
- •Core work: model relationships between what was said and how it was expressed
- •Go-to-market lesson: great demos are required to convey the new interaction paradigm
- 13:14 – 15:19
Founder lessons + closing: conviction, momentum, and ‘faster, faster’
They end with advice to founders: develop conviction in the vision and drive momentum daily. The segment closes with their speed-focused culture and congratulations on the Series B.
- •Advice: build deep conviction and don’t over-index on outside opinions
- •Momentum is framed as the essential driver—something must move every day
- •Startup moat = speed; they aim to stay months ahead
- •Company motto emphasizes relentless pace; closing Series B congratulations
