YC Root AccessThe Data Layer for the Robot Economy
CHAPTERS
- 0:05 – 0:30
Encord’s mission: an AI-native data layer for physical AI and robotics
Nicolas introduces the Encord founders and the context of their recent Series C. Ulrik frames Encord as data infrastructure for top AI teams, especially those building robotics and other physical-world AI systems.
- •Encord positions itself as AI-native data infrastructure
- •Primary focus on physical AI and robotics teams
- •Series C announcement sets the stage for growth and momentum
- 0:30 – 0:51
What “data infrastructure” means: curate, annotate, evaluate—and keep bad data out
Ulrik explains what a data infrastructure platform does in practice. The core is managing the full lifecycle of training and evaluation data so models receive high-quality inputs and avoid harmful noise.
- •Ensuring the “right data in” and “wrong data out”
- •Universal platform to create/manage/annotate/evaluate data
- •Small dataset errors can have outsized real-world impact
- •Data demands grow as model size and complexity scale
- 0:51 – 1:22
Why the problem gets harder after deployment: continuous data improvements
The conversation clarifies that data work doesn’t stop once a model ships. Teams must keep feeding models new, correct data to maintain and extend performance in production settings.
- •Models require ongoing data refresh to push performance
- •Production introduces new edge cases and drift
- •Scaling AI increases dataset size and operational difficulty
- 1:22 – 2:19
Founding story: spotting data as the defensible bottleneck (pre-ChatGPT)
Ulrik and Eric describe how they met and why they chose the data layer as their wedge. They saw data wrangling and outsourced labeling as slow, painful, and ripe for modernization.
- •Late-2010s deep learning context at Imperial and HFT experience
- •Data identified as the slowest, most time-consuming ingredient
- •Early AI development relied heavily on offshore manual labeling
- •Decision to build a better system for the AI data problem
- 2:19 – 3:13
Early skepticism: raising money when AI data tooling looked “unsexy”
Eric recounts how the market and investors underestimated AI’s trajectory during their YC days. Encord even struggled to raise seed capital as other categories dominated attention.
- •AI tooling wasn’t a hot category in their YC batch
- •Fintech/crypto/remote work drew more investor focus
- •Seed fundraising challenges due to perceived market size limits
- •Anecdote: investors misjudged AI’s market potential
- 3:13 – 3:42
Encord today: customers, scale, and traction across physical-world AI
Ulrik provides concrete business scale—customer count, team size, and funding totals. He highlights the range of use cases from autonomous driving to major robotics programs.
- •300+ AI teams as customers
- •Robotics and autonomous driving companies (e.g., Toyota)
- •150-person team across London and San Francisco
- •$110M raised total including $60M Series C
- 3:42 – 4:43
The first product: automating computer-vision annotation workflows
The founders describe Encord’s initial wedge: making annotation faster and less dependent on slow, outsourced processes. The focus was primarily on computer vision labeling and segmentation.
- •Started in Winter ’21 with annotation automation focus
- •Targeted computer vision datasets and labeling workflows
- •Aimed to replace slow send-out-and-return labeling loops
- •Built tooling to reduce tedious manual effort
- 4:43 – 6:14
ChatGPT changes the market: trust in AI-enabled automation emerges
Eric explains that early users didn’t trust AI to operate on their data—even within AI companies. ChatGPT served as a proof point that AI could be trusted, unlocking broader adoption of AI-assisted data workflows.
- •Encord built “micro models” to automate labeling from few examples
- •Customers initially resisted AI touching their data
- •ChatGPT created widespread confidence in AI outputs
- •Market education accelerated demand for automation
- 6:14 – 7:53
Expanding to multimodal and physical AI: why embodied systems need a data layer
The conversation moves from vision-only tooling to multimodal systems and physical AI. Eric argues that scaling laws worked for digital AI because data was abundant, while physical AI’s bottleneck is real-world data collection and management.
- •Shift from images/videos to multimodal (text+image+audio+sensor)
- •Physical AI requires embodied, real-world training data
- •Compute is available; data is the limiter to scaling in physical AI
- •Encord’s opportunity: enable the “ChatGPT moment” for robots
- 7:53 – 8:37
New offering: Bay Area R&D facility for data collection and robot training environments
Ulrik describes Encord’s move into supporting pre-training and data collection for physical AI—an area they previously avoided. They’re building environments where robotics companies can bring robots to capture high-quality training data at scale.
- •Pre-training/data collection is easy for LLMs (scrape the internet) but hard for robots
- •Encord opens an R&D facility to support embodied data capture
- •They don’t build robots; they provide environments and infrastructure
- •Goal: make scalable, real-world data collection practical
- 8:37 – 9:53
From training to operations: post-deployment observability and exception handling
The founders emphasize that production robotics requires operational tooling, not just training datasets. They discuss exception handling, observability, and connecting real-world behavior back to digital feedback loops.
- •Post-deployment needs: exception handling and observability
- •Robots nearing market readiness increase urgency of ops tooling
- •Infrastructure must couple physical-world events to model iteration
- •End-to-end view enables a tighter data flywheel
- 9:53 – 11:03
Why Encord can be a platform: the data flywheel and consolidated pipeline view
They explain how indexing, curation, annotation, and model-in-the-loop pre-labeling reinforce each other. A unified system can increasingly automate the stack and accelerate iteration from pre-training through deployment.
- •Customers index/curate/annotate within a single platform
- •Embedding customer models enables pre-labeling and automation
- •Unified visibility improves speed-to-production and revenue
- •Multi-modality adds complexity vs text-only workflows
- 11:03 – 12:20
Humans in the loop: frontier labeling, supervision, and higher safety stakes
The discussion addresses why human oversight remains critical, particularly in early-stage physical AI. Compared to digital AI, errors can cause real-world harm, raising data quality and governance requirements.
- •Humans needed at the frontier for new tasks (e.g., folding laundry)
- •Desire for humans as supervisors/managers of AI systems
- •Exception handling is essential when models fail
- •Physical AI has lower error tolerance and higher safety stakes
- 12:20 – 13:47
Customers and value proposition: speed to market and better models (with example)
Eric and Ulrik define the ideal customer as any embodied/real-world AI team and explain what they’re really buying: faster iteration and improved performance without building internal data infrastructure. They share Weave Robotics as a concrete example.
- •Target customers: robotics, self-driving, autonomous systems
- •Primary ROI: faster time-to-market and better model quality
- •Encord offloads data infrastructure so teams focus on robots/products
- •Example: Weave Robotics using Encord for a laundry-folding robot
- 13:47 – 18:43
Series C rationale and the long-term vision: powering the robot economy + founder lessons
They justify the timing of the Series C as physical AI investment accelerates and the market expands. Eric outlines an ambition akin to Stripe’s in payments—be the default pipe for physical AI data—then closes with hiring plans and founder advice on decision-making and navigating toward long-term goals.
- •Funding used to accelerate capturing the physical AI opportunity
- •Physical AI’s potential is enormous given the share of the global economy tied to physical work
- •Timeline expectations: hype, consolidation, then broad adoption of general-purpose robots
- •Hiring humans and internal “agents” across functions
- •Founder lessons: avoid indecision; be flexible on path while fixed on destination