Skip to content
YC Root AccessYC Root Access

Better AI Starts With Better Verification

Halluminate (YC S25) is building reinforcement learning environments and benchmarks to help AI models do knowledge work beyond coding, starting with finance. With a team of fewer than 10 people, the company works with four of the top five closed-source US AI labs and recently raised a $30 million Series A led by Oak HC/FT. In this Fireside, co-founders Jerry Wu and Wyatt Marshall sit down with YC Partner Jon Xu to share how testing browser agents led them to building simulations where models practice tasks like completing financial spreadsheets. They explain why accurate scoring and expert review matter more than sheer task volume, and how weak checks can teach models to take shortcuts. They also discuss their vision for simulated companies where teams of AI agents learn to work together. https://halluminate.ai Chapters: 00:00 — Halluminate’s Pivot From Evals to Training Environments 03:48 — Teaching AI to Do Financial Work 07:13 — Why Verification Matters More Than Volume 10:13 — From Finance to Simulated Companies 13:30 — AI Safety and Building the Team Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Jon XuhostJerry WuguestWyatt Marshallguest
Oct 1, 202616mWatch on YouTube ↗

At a glance

WHAT IT’S REALLY ABOUT

Halluminate trains better AI by building verifiable RL work environments

  1. Halluminate builds frontier RL training environments and benchmarks that help improve AI performance in non-coding knowledge-work domains, starting with financial services.
  2. The company pivoted from building evals to building post-training environments after realizing simulated environments were necessary not just to measure agents but to train them effectively.
  3. They argue that high-quality verification (reward design and checking) matters more than sheer task volume because flawed verifiers incentivize shortcuts and reward hacking, producing misaligned learning.
  4. Their roadmap scales environment complexity from app-level tasks (Excel/PowerPoint/email) to multi-agent workflows and eventually large simulations like “simulated companies” that teach autonomy and collaboration.
  5. They view alignment and safety as an industry-wide responsibility driven by data and environment quality, and they’re hiring across ops, research (post-training/evals), and infra to build and host these simulations.

IDEAS WORTH REMEMBERING

5 ideas

They moved from benchmarking models to directly improving them via RL environments.

Halluminate pivoted from evaluating agents to building the simulated tasks plus verifiers that can be used as reward signals for RL post-training, turning “testing” infrastructure into “teaching” infrastructure for frontier labs.

Verification is the core technical bottleneck for non-coding knowledge-work RL.

In coding you can often verify with unit tests; in finance/knowledge work you must score things like spreadsheet formulas, sourcing, and formatting against expert expectations, which is harder to specify and automate.

Better verification beats more data when training with RL environments.

They argue that scaling task count without robust verifiers trains the wrong behaviors—either granting credit for shortcuts/reward hacks or failing to reward correct work—leading to mislearning and potential misalignment.

Reward hacking includes process violations, not just sandbox exploits.

They highlight “soft” reward hacking (e.g., jumping to answers without defensible process) alongside “hard” hacks (escaping sandboxes), and emphasize SME-driven QA so rewards match real-world professional standards.

Finance is a wedge that generalizes to broader enterprise knowledge work.

They chose finance because it’s large, high-value, and comparatively verifiable, and believe financial reasoning skills transfer to adjacent domains like consulting, accounting/FP&A, and insurance.

WORDS WORTH SAVING

5 quotes

At Halluminate, what we do is we build frontier RL environments and benchmarks to push what models can do in non-coding, non-coding knowledge work domains, starting with financial services, right?

— Jerry Wu

It's better to have fewer that are better than a lot more that are poor quality.

— Wyatt Marshall

Either of those situations, you're gonna end up with a model that's learning the wrong stuff because that verifier is directly teaching the model, so if the verifier's not aligned with the intention of the task and, like, what a successful output looks like on that task, then you're gonna get models that are learning the wrong things.

— Wyatt Marshall

We've noticed something that we internally call the Moore's Law of RL environments, which basically means every six to eight months, the complexity or scope or, you know, size of the environments that we train on roughly doubles, right?

— Jerry Wu

I think in the future what we are actually gonna be building and delivering are actually simulated companies, right? Or simulated industries, simulated governments that, that teach agents not just how to do work, but how to be autonomous operators and coworkers in our economy and our society, right?

— Jerry Wu

Pivot from evals to post-training environmentsRL environment loop: objective, environment, verification/rewardFinance knowledge-work tasks (Excel/PowerPoint/diligence)Verification design and quality assuranceReward hacking and process-based scoringScaling environment complexity (“Moore’s Law of RL environments”)Simulated companies/industries and infra needs

High quality AI-generated summary created from speaker-labeled transcript.

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.