Y CombinatorWhy Old Startup Ideas Like Recruiting Now Work With LLMs
Through LLM code evals and gross margin thinking about unit economics; Triplebyte took years to build what AI does in weeks, making recruiting startups viable.
CHAPTERS
- 0:00 – 1:07
AI tooling and agent infrastructure: why the opportunity surface is exploding
The hosts open by framing the moment: model capabilities are jumping quickly, which creates a wave of “old ideas that only work now.” They set the theme that big startup opportunities exist not just in apps, but in the infrastructure required to deploy AI and agents reliably.
- •Rapid model improvements (e.g., longer context) are unlocking new product categories
- •Many startup ideas are newly feasible despite being “old” concepts
- •Large greenfield space remains in agent deployment, tooling, and infrastructure
- •Taste, evals, and good prompts/datasets can create “magical output”
- 1:07 – 4:19
Recruiting marketplaces reborn: AI makes evaluation feasible on day one
Harj explains how LLMs change recruiting startups by collapsing the hard, slow part: candidate evaluation. He contrasts Triplebyte’s multi-year effort to build labeled datasets and human-in-the-loop interviewing with modern AI-first approaches that can expand beyond engineers immediately.
- •Triplebyte required years of interviews to build a labeled dataset pre-LLMs
- •LLMs enable code evaluation and more general knowledge-work assessment quickly
- •AI reduces complexity of multi-sided marketplaces by removing human intermediaries
- •Founders can expand from one job category to many much faster with LLMs
- 4:19 – 6:06
Psychological hurdle: entering “dead categories” after past failures (Webvan → Instacart pattern)
They discuss why founders must push through investor skepticism in categories that previously burned capital. The Instacart/Webvan analogy illustrates how new enabling tech can make an old model work—if teams are willing to re-test the space when the maze shifts.
- •Capital-heavy failures create cynicism even when tech shifts the equation
- •LLMs are an enabling tech analogous to smartphones for Instacart
- •Category stigma can be a moat for determined founders (less competition)
- •You only find the new openings by actively experimenting “in the maze”
- 6:06 – 7:34
Technical screening agents: from basic filters to nuanced senior-level assessment
Diana highlights AI agents that automate technical interview screening, reducing a task engineers dislike and that consumes massive time. The group argues LLMs expand the addressable market because evaluations can be more sophisticated and applied to senior candidates, not just junior filtering.
- •AI agents can run technical screening interviews end-to-end (Apriora example)
- •Pre-LLM tools mostly weeded out non-engineers or very junior candidates
- •LLMs allow nuanced assessment suitable for senior applicants
- •Solving a single painful workflow slice can unlock significant adoption
- 7:34 – 9:48
Hyper-personalized education: tutors, exam prep, and teacher grading
The conversation shifts to education, emphasizing personalization as the long-standing “internet dream” now becoming practical. They cite AI-powered exam prep tools and teacher grading assistants as examples where agents take on disliked work and can improve learning outcomes.
- •LLMs enable “personal tutor in your pocket” experiences for the first time
- •RevisionDojo: tailored exam prep with strong DAU/power-user engagement
- •Edexia: agent support for grading—reducing a major driver of teacher churn
- •Early adoption may skew toward nimble private schools; public policy lags
- 9:48 – 13:21
Distribution vs product quality: when does ‘better’ translate into growth?
Harj questions whether superior AI products automatically win distribution, especially in consumer. Garry argues economics are key: as intelligence cost falls, freemium and mass adoption become more viable—unlocking consumer AI at scale.
- •Great AI UX doesn’t guarantee distribution; pricing and channels still matter
- •Falling inference costs and distillation may make intelligence “nearly free”
- •Freemium could return: free base with paid power features (subscriptions)
- •Examples of freemium-ish AI: OpenAI, Perplexity, education apps
- 13:21 – 14:39
New consumer business models: charging like a tutor, not like an app
They explore how AI can shift willingness-to-pay by upgrading outcomes to human-equivalent services. For education, this means parents may pay more if the product matches a high-quality tutor—changing unit economics and reducing the need for massive user counts.
- •AI can move products from “cheap app” to “service replacement” pricing
- •Parents pay for tutors; AI tutors that work can capture that spend
- •Better outcomes and retention can justify higher ARPU with fewer users
- •Same dynamic seen in enterprise: budgets rise when replacing teams, not software
- 14:39 – 16:07
Moats in the AI era: brand, switching costs, and ‘good enough’ leadership
They discuss defensibility beyond “adding AI,” focusing on brand, integrations, and switching costs. ChatGPT’s dominance over technically strong alternatives illustrates that being first, being good, and owning mindshare can outweigh objective model parity.
- •Moats come from brand, workflow lock-in, integrations (e.g., school auth/logins)
- •OpenAI’s lead shows mindshare and product execution matter as much as model quality
- •A product can win even if not objectively best—if it’s good enough and trusted
- •OpenAI increasingly focuses on apps, raising questions for startups
- 16:07 – 17:40
Platform neutrality: why assistants should be selectable like browsers
Garry argues that big platform gatekeepers suppress innovation by bundling weak default assistants (Siri/Google Assistant). He draws parallels to net neutrality and Windows browser-choice interventions, advocating for user choice in voice/assistant layers to enable a free market of AI agents.
- •Bundled assistants can block better third-party agents from reaching users
- •Historical parallels: net neutrality and Windows browser/search choice
- •Siri as an example of stagnant default UX despite available AI capability
- •Policy and platform rules could dramatically reshape distribution for AI apps
- 17:40 – 23:24
Big Tech’s AI execution gap: ‘shipping the org’ and innovator’s dilemma
They critique Google and Meta for confusing product strategy, fragmented APIs, and poor integrations despite strong underlying models/hardware. The discussion ties execution issues to company structure and the innovator’s dilemma—why incumbents hesitate to cannibalize revenue.
- •Google: strong models and TPUs, but weak integrations and duplicated APIs/orgs
- •‘Ship the org’ dynamics can replace coherent product decisions
- •Innovator’s dilemma: replacing core products (e.g., search) risks revenue collapse
- •Meta AI’s presence in apps feels intrusive and often lacks useful permissions
- 23:24 – 25:14
AI horseless carriages: system prompts, user control, and ‘vibe-coded’ UX
A critique of Google’s Gmail Gemini integration leads to a broader product lesson: empower users by exposing control (like editable system prompts) and building for real workflows. They riff on new product ideas enabled by AI-first creation, including interactive/vibe-coded publishing.
- •Good AI UX may require giving users deeper control (system prompt customization)
- •Many integrations fail due to rigid tone/behavior and poor workflow fit
- •Interactive, prompt-driven content experiences can be native to the web
- •New “AI-first” platforms (e.g., blogging) may be ripe for reinvention
- 25:14 – 30:03
Full-stack startups and gross margins: why the 2010s wave struggled
Jared revisits the tech-enabled services/full-stack startup wave and why many underperformed relative to SaaS. They emphasize gross margin as the core constraint: heavy ops complexity slows scaling and distracts from product and distribution.
- •Full-stack captured more value in theory but often lacked SaaS-like margins
- •Scaling required hiring more people, creating operational drag
- •Zenefits as a cautionary tale: too much human-heavy process vs software leverage
- •Low gross margins also increase management complexity and reduce focus
- 30:03 – 31:54
The comeback case: AI makes full-stack companies look like software
They make a bullish argument that agents can remove the ops burden that doomed earlier full-stack models. Examples like Atrium’s early timing and newer legal AI tools suggest a path where software agents do the work, potentially evolving into service-delivery giants.
- •Agents can replace operational headcount, improving margins and scalability
- •Atrium’s failure framed partly as ‘AI not good enough yet’
- •Legal tools (e.g., fast-growing YC examples) hint at eventual AI-run service firms
- •Virtual assistants could coordinate both digital work and real humans
- 31:54 – 37:13
MLOps timing and infrastructure bets: being early vs being wrong
They return to infrastructure: in 2019–2020, ML tooling often lacked buyers because ML wasn’t broadly useful yet. Companies like Replicate and Ollama illustrate how persistence through a ‘not-ready’ market can pay off when a capability breakthrough (diffusion, LLaMA) arrives.
- •Earlier ML tooling waves suffered because end-user ML apps weren’t working/ubiquitous
- •Replicate and Ollama succeeded when new model breakthroughs created demand
- •Lesson: timing matters; persistence can win if you’re positioned before the inflection
- •Still-large opportunity set in evals, deployment, and agent infrastructure
- 37:13 – 40:46
Updated startup advice for the AI age: explore the frontier, then build
Jared argues that classic lean-startup “sell before you build” guidance is less dominant in an era where the idea space has expanded dramatically. The new playbook is to follow curiosity, experiment with cutting-edge capabilities, and let new ideas emerge from hands-on exploration—while noting many incumbents still aren’t adapting.
- •Lean-startup canon formed when good ideas were scarce; AI expands the frontier
- •Follow curiosity and work at the edge of the future to ‘bump into’ ideas
- •Magic comes from prompts + data + evals + taste—still underutilized
- •Many established startups are slow to adopt AI internally, leaving openings