At a glance
WHAT IT’S REALLY ABOUT
Halluminate trains better AI by building verifiable RL work environments
- Halluminate builds frontier RL training environments and benchmarks that help improve AI performance in non-coding knowledge-work domains, starting with financial services.
- The company pivoted from building evals to building post-training environments after realizing simulated environments were necessary not just to measure agents but to train them effectively.
- They argue that high-quality verification (reward design and checking) matters more than sheer task volume because flawed verifiers incentivize shortcuts and reward hacking, producing misaligned learning.
- Their roadmap scales environment complexity from app-level tasks (Excel/PowerPoint/email) to multi-agent workflows and eventually large simulations like “simulated companies” that teach autonomy and collaboration.
- They view alignment and safety as an industry-wide responsibility driven by data and environment quality, and they’re hiring across ops, research (post-training/evals), and infra to build and host these simulations.
IDEAS WORTH REMEMBERING
5 ideasThey moved from benchmarking models to directly improving them via RL environments.
Halluminate pivoted from evaluating agents to building the simulated tasks plus verifiers that can be used as reward signals for RL post-training, turning “testing” infrastructure into “teaching” infrastructure for frontier labs.
Verification is the core technical bottleneck for non-coding knowledge-work RL.
In coding you can often verify with unit tests; in finance/knowledge work you must score things like spreadsheet formulas, sourcing, and formatting against expert expectations, which is harder to specify and automate.
Better verification beats more data when training with RL environments.
They argue that scaling task count without robust verifiers trains the wrong behaviors—either granting credit for shortcuts/reward hacks or failing to reward correct work—leading to mislearning and potential misalignment.
Reward hacking includes process violations, not just sandbox exploits.
They highlight “soft” reward hacking (e.g., jumping to answers without defensible process) alongside “hard” hacks (escaping sandboxes), and emphasize SME-driven QA so rewards match real-world professional standards.
Finance is a wedge that generalizes to broader enterprise knowledge work.
They chose finance because it’s large, high-value, and comparatively verifiable, and believe financial reasoning skills transfer to adjacent domains like consulting, accounting/FP&A, and insurance.
WORDS WORTH SAVING
5 quotesAt Halluminate, what we do is we build frontier RL environments and benchmarks to push what models can do in non-coding, non-coding knowledge work domains, starting with financial services, right?
— Jerry Wu
It's better to have fewer that are better than a lot more that are poor quality.
— Wyatt Marshall
Either of those situations, you're gonna end up with a model that's learning the wrong stuff because that verifier is directly teaching the model, so if the verifier's not aligned with the intention of the task and, like, what a successful output looks like on that task, then you're gonna get models that are learning the wrong things.
— Wyatt Marshall
We've noticed something that we internally call the Moore's Law of RL environments, which basically means every six to eight months, the complexity or scope or, you know, size of the environments that we train on roughly doubles, right?
— Jerry Wu
I think in the future what we are actually gonna be building and delivering are actually simulated companies, right? Or simulated industries, simulated governments that, that teach agents not just how to do work, but how to be autonomous operators and coworkers in our economy and our society, right?
— Jerry Wu
High quality AI-generated summary created from speaker-labeled transcript.
