Y CombinatorVarun Mohan: Why Insights Depreciate and Why Evals Compound
Windsurf rebuilt from GPU virtualization to vibe coding in 48 hours; evals and irrational optimism are the moats when every competitive insight depreciates.
CHAPTERS
- 0:00 – 0:52
Why startups must keep proving themselves (and why everyone becomes a “builder”)
Varun opens with a core belief: insights decay fast, so startups must continuously earn their edge through new execution. He also frames a broader shift from “developer” to “builder,” where software creation becomes radically democratized.
- •Competitive advantage is temporary; insights depreciate quickly
- •Continuous innovation is the only way to avoid slow death
- •NVIDIA vs. AMD as an example of relentless innovation pressure
- •“Developer” role broadens into “builder” for everyone
- •Software creation becomes more accessible over time
- 0:52 – 1:20
Windsurf today: scale, usage, and what people build with it
The conversation sets context on Windsurf’s traction and how it’s used in practice. Varun describes both large-codebase modifications and fast 0→1 app building as common patterns.
- •Over a million developers have tried the product
- •Hundreds of thousands of daily active users
- •Used for large codebase changes and rapid prototyping
- •Broad IDE usage and varied workflows
- •Excitement about the direction of AI coding tech
- 1:20 – 2:38
From Exafunction to a looming disruption: GPU virtualization meets transformers
Varun explains the company’s original thesis and product: GPU virtualization for deep learning workloads. As transformer models rose, the team recognized infrastructure would commoditize, forcing a rethink of the company’s future.
- •Company started as Exafunction, a GPU virtualization business
- •Background in autonomous vehicles and AR/VR informed deep learning conviction
- •Managed 10,000+ GPUs and reached a few million in revenue
- •Transformers (e.g., Text-Davinci) shifted the market dynamics
- •Fear of commoditization if everyone runs the same architecture
- 2:38 – 6:33
The 48-hour “bet the company” pivot to Codium
In a high-stakes moment, the founders decided over a weekend to abandon the existing business and pivot into AI coding. With only eight people, they aligned around a product the whole team would be energized to build.
- •Pivot decision made within a weekend; announced Monday
- •Team size ~8 while free-cash-flow positive
- •Raised $28M despite lean headcount
- •Key rationale: current success didn’t imply scalability
- •Chose a direction that maximized team excitement and conviction
- 6:33 – 7:55
Competing with GitHub Copilot: irrational optimism + uncompromising realism
Varun lays out a startup mindset: you need optimism to attempt the impossible, but realism to change course when facts change. Despite Copilot’s distribution advantage, they bet they could win on technology and iteration speed.
- •Two necessary beliefs: optimism and realism (in tension)
- •Copilot felt inevitable due to GitHub/Microsoft/OpenAI
- •Early bet: they could train and run models themselves
- •Accepted the possibility of failure as the cost of trying
- •Focus on executing and adapting as evidence changes
- 7:55 – 10:26
Shipping the earliest versions: free, worse than Copilot, then rapidly better
Codium launched quickly as a free VS Code extension, initially inferior to Copilot. Within months, the team improved infrastructure and model training, adding capabilities like fill-in-the-middle and beating Copilot on key dimensions.
- •First shipped product was materially worse—differentiator was price (free)
- •Shipped within ~2 months and launched via Hacker News
- •Moved from open-source model to training their own models
- •Introduced fill-in-the-middle editing as an early edge
- •By early 2023, autocomplete quality/latency surpassed Copilot in key ways
- 10:26 – 11:43
From free users to first enterprise customers (and the need to personalize)
Widespread free adoption turned into inbound enterprise interest, driven by security and private-code personalization needs. Early pilots led to major customers and shaped a focus on performance over massive codebases.
- •Adoption across IDEs drove broad developer usage
- •Enterprises wanted secure deployments and private-data personalization
- •Early logos included Dell and JPMorgan Chase
- •Some customers had tens of thousands of developers
- •Hard requirement: work well on 100M+ line codebases with fast, relevant suggestions
- 11:43 – 13:14
Why they expanded beyond VS Code: multi-IDE support as an enterprise prerequisite
Varun explains why supporting multiple IDEs early was essential to win enterprise rollouts (e.g., Java shops on IntelliJ). Early architectural decisions reduced the cost of adding editors by sharing core infrastructure.
- •Enterprise dev orgs are multi-language and multi-IDE
- •IntelliJ dominance among Java devs made JetBrains support critical
- •Avoid becoming “one of many” tools inside a company
- •Shared infrastructure minimized per-IDE engineering overhead
- •Early architecture choices enabled fast horizontal expansion
- 13:14 – 15:40
From Codeium to Windsurf: agents forced a full IDE to unlock the experience
As the team bet on agents, they found VS Code’s surface limited what they could deliver. With improved tool-calling models arriving (e.g., Sonnet 3.5), they committed to building their own IDE to support an agent-first workflow.
- •Space moves fast; many internal bets fail by design
- •Early agent prototypes didn’t work until model/tool-calling improved
- •Key building blocks: codebase understanding, intent, fast edits
- •Belief shift: devs will spend more time reviewing AI-generated changes
- •Forked VS Code to build a new agentic IDE: Windsurf
- 15:40 – 17:14
Forking VS Code and shipping in under three months: adoption, rough edges, retention
Windsurf shipped quickly across operating systems with a small engineering team. Early adopters arrived fast, and while churn existed due to roughness, improvements in agents and passive autocomplete increased retention and word-of-mouth.
- •Forked VS Code and learned a complex codebase quickly
- •Shipped the IDE in <3 months across OSes
- •Early adopter uptake was rapid despite rough edges
- •Agent and “passive tab” experience improved significantly over months
- •Retention improved as capabilities stabilized and expanded
- 17:14 – 18:29
Go-to-market reality: selling to Fortune 500 requires AEs plus deployed engineers
Varun describes an unusual company shape: lean engineering paired with a relatively large go-to-market team. Enterprise adoption needs support, enablement, and hands-on technical deployment work beyond self-serve credit cards.
- •Engineering stayed lean; GTM comparatively larger
- •Fortune 500 sales require support and adoption help
- •Two GTM roles: curious AEs and hands-on deployed engineers
- •GTM benefits from authentic product enthusiasm and fluency
- •Enterprise success depends on real-world enablement, not just distribution
- 18:29 – 19:44
Non-developers using Windsurf: domain experts building apps via Cascade
A notable share of Windsurf usage comes from non-technical users who never open the editor. They operate through the agent interface and browser preview, while developers later pick up and refine the underlying repository work.
- •Non-technical employees can become power users (e.g., partnerships lead)
- •Domain experts build internal tools without waiting on engineering backlogs
- •Cascade + browser preview enables “no-editor” workflows
- •Security/deployment still often needs a specialist checkpoint
- •Windsurf can resume work even when the code becomes complex
- 19:44 – 25:19
Product strategy vs competitors: agent-first, less @mentioning, more intent understanding
Varun contrasts Windsurf’s approach with chat+autocomplete paradigms and highly manual context tagging. Windsurf aimed to reduce configuration and rely on deeper code understanding, positioning agents as the future interface.
- •Belief that agents—not just chat—are the direction of coding tools
- •Avoiding an @mention-heavy workflow as a long-term anti-pattern
- •Analogies to search UX evolution (from complex portals to simple boxes)
- •Investments in codebase understanding, developer intent, fast edits
- •Acknowledgement that “obvious” insights become obvious only in hindsight
- 25:19 – 27:08
Context and retrieval: beyond vector DB RAG with reranking, parsing, and hybrid search
The team built a more rigorous context assembly pipeline rather than relying solely on vector search. By combining multiple retrieval techniques and GPU-heavy reranking, they optimize precision and recall for real enterprise tasks.
- •RAG is correct in principle, but implementations vary
- •Vector DB is one tool; not always sufficient for complex code tasks
- •Hybrid retrieval: keyword search + embeddings + AST parsing
- •Real-time reranking of large code chunks using GPU infrastructure
- •Goal: high precision/recall so tasks like API migrations don’t miss instances
- 27:08 – 30:14
Evals as the engine: using commits + tests to measure retrieval, intent, and correctness
Drawing from safety-critical roots (autonomy), Varun emphasizes evaluation-driven development. They turn open-source commits and unit tests into repeatable benchmarks that decompose performance into actionable sub-metrics.
- •Autonomy background drove a bias toward rigorous evaluation
- •Use open-source commits with tests to create reproducible tasks
- •Masking/partial-change tasks to evaluate intent completion
- •Break down metrics: retrieval accuracy, intent accuracy, test pass rate
- •Evals justify complexity; without them, improvements are guesswork
- 30:14 – 35:14
Hardcore engineering workflows and “surgical” agent edits: practical usage tips
Varun explains how power users apply agents to real production work, from boilerplate elimination to deployment workflows. He also shares habits that keep agent-driven changes controlled, reversible, and iteratively correctable.
- •Teams increasingly start tasks by writing intent, not typing code first
- •Agents handle repetitive work; workflows can manage deployments
- •Risk: insufficient intent can cause overly broad changes
- •Tip: commit frequently and use revert/iteration instead of abandoning tools
- •Expectations management: 90% right can still feel unusable unless you iterate
- 35:14 – 52:35
Future of coding, hiring, and the “GPT wrapper” concern: moving goalposts and compounding alpha
The discussion looks forward: AI improves every stage of software development and changes what engineers do and how companies hire. Varun argues that products like Windsurf must continuously add value above foundation models as baselines rise, and he closes with advice to pivot faster and treat it as a strength.
- •AI will amplify writing, reviewing, testing, debugging, and design
- •Engineering shifts toward rapid hypothesis testing (more ‘research-like’)
- •Hiring: value high agency, curiosity; test both AI-usage and raw thinking
- •Wrapper meme response: keep doubling the ‘alpha’ above foundation models
- •Founder advice: change your mind faster; pivoting is a badge of honor