CHAPTERS
- 0:05 – 1:35
Halluminate’s mission: RL environments & benchmarks for non-coding knowledge work
Jon introduces Jerry and Wyatt and asks what Halluminate does. Jerry explains they build frontier reinforcement learning (RL) environments and benchmarks that improve foundation models’ performance in non-coding knowledge work, starting with finance, and describes their traction with top frontier labs.
- •Builds RL environments and benchmarks for non-coding knowledge work domains
- •Focus begins with financial services (spreadsheets, decks, diligence, investing workflows)
- •Benchmarks/environments have influenced real gains in financial reasoning and literacy
- •Small team, rapid growth, and partnerships with leading closed-source labs
- 1:35 – 4:02
Founding story and the pivot: from evals to post-training environments
Jerry recounts how he and Wyatt met at Cornell and originally started Halluminate to solve evaluation/benchmarking needs encountered while building finance agents. Work on browser/computer-use evals revealed that simulated environments were necessary—and valuable beyond evals—prompting a pivot toward post-training environments for frontier labs.
- •Initial focus: evals and benchmarking to verify agent quality and efficiency
- •Early work included computer-use and browser-use evals (e.g., booking/checkout flows)
- •Realization: robust evals require simulated replicas of real apps/sites
- •Pivot thesis: environment-building + verification is core to post-training and RL scaling
- 4:02 – 4:57
Why specialize in finance: where value accrues beyond coding and computer-use
They explain why Halluminate moved from generic computer-use environments into finance and other enterprise knowledge work. The team believed non-coding domains would create more long-term value and that progress was bottlenecked by the lack of domain-specific environments, data, and benchmarks.
- •Most industry effort was on coding/computer-use; non-coding work lagged behind
- •Non-coding jobs outnumber coding jobs, implying larger economic impact
- •Model weakness attributed to missing training environments and benchmarks
- •Decision to specialize in finance as a high-value, enterprise-critical starting point
- 4:57 – 6:02
RL environment basics: objective, sandbox, and verification loop
Wyatt breaks down what an RL environment is and how it’s used for either evaluation or post-training. He emphasizes that verification (the scoring/reward mechanism) is the hard part in knowledge work, unlike coding where unit tests provide crisp signals.
- •Core components: objective prompt, environment to act in, and evaluation/verification
- •For evals, the score is reported; for post-training, it becomes the reward function
- •Knowledge-work verification is inherently fuzzier than coding/math
- •Verification quality directly determines what the model learns
- 6:02 – 7:13
Concrete example: training on Excel financial modeling tasks
Wyatt gives a detailed example of an RL task in finance: completing or repairing a partially built Excel model. Verification compares the agent’s spreadsheet to a “golden” ground truth and checks formulas, sourcing, and formatting requirements that matter to practitioners.
- •Task setup: prompt + partially constructed spreadsheet + supporting references
- •Agent attempts to produce a correct completed financial model
- •Verification checks formula correctness by cell, proper sourcing/derivation, and formatting
- •Outputs are combined into a holistic performance score for learning/evals
- 7:13 – 8:50
Quality over volume: verification as the core differentiator
The discussion turns to what makes environments ‘high quality’ and why scaling volume alone isn’t enough. Wyatt argues that verification is the central lever: poor verifiers teach models the wrong behaviors and can create misleading improvements or regressions.
- •Environment quality determines post-trained model quality
- •Better to have fewer high-quality environments than many low-quality ones
- •Verifier errors can over-reward shortcuts or under-reward correct solutions
- •Misverification is a direct path to misalignment and degraded learning
- 8:50 – 10:25
Reward hacking in knowledge work: shortcuts, sandboxes, and “showing work”
Jon asks how they address reward hacking and learning plateaus. Wyatt outlines hard failures (escaping sandboxes, finding answer keys) and subtler ones (taking unapproved shortcuts), stressing SME-driven oversight and QA to align automated scores with expert judgment.
- •Reward hacking ranges from sandbox escapes to subtle shortcut-taking
- •Knowledge work often requires process compliance, not just final answers
- •Need robust QA and oversight during task creation, not just task generation
- •Subject matter experts validate that verifiers match real-world grading standards
- 10:25 – 11:40
Roadmap: expanding from finance into adjacent knowledge-work domains
Jerry explains why finance is a strong wedge: it’s large, relatively verifiable, and has many sub-disciplines. He argues that financial reasoning generalizes into adjacent areas like consulting, accounting, FP&A, and insurance, which Halluminate plans to expand into.
- •Finance: large headcount, high value per worker, and comparatively verifiable tasks
- •Coverage spans subdomains (IB, PE, equity research, etc.)
- •Belief that finance reasoning skills transfer to other enterprise workflows
- •Expansion targets: consulting, accounting, FP&A, insurance and more
- 11:40 – 12:47
The “Moore’s Law” of RL environments: doubling complexity every 6–8 months
Jon asks how environments evolve as models get smarter. Jerry describes a pace where environment scope must expand roughly in step with model capability—moving from single-app tasks (Excel/PowerPoint/email) to multi-agent teamwork and eventually large-scale simulations.
- •Internal observation: environment complexity doubles roughly every 6–8 months
- •Near-term: environments spanning multiple tools (spreadsheets, decks, email)
- •Next step: multi-agent team settings completing larger work chunks
- •Long-term vision: simulated companies/industries/governments as training grounds
- 12:47 – 13:30
Simulated companies as infrastructure: how customers will ‘rent’ training time
Jerry frames large digital simulations of real work as a necessary future training substrate for autonomy and coworker-like behavior. Halluminate’s vision is to provide these environments as infrastructure that labs and enterprises can access on demand.
- •Simulations teach agents to operate autonomously within organizational constraints
- •Building realistic simulations is extremely difficult but increasingly necessary
- •Business model vision: labs/enterprises purchase or rent time in environments
- •Goal: push the frontier of real-world work capability, not just toy tasks
- 13:30 – 14:49
Safety and alignment: better environments and verification as a frontline defense
Wyatt asks about alignment and safety; Jerry argues many failures stem from poor RL environments and weak verification. He positions environment builders as central to industry-wide safety, not just lab safety teams or external evaluators.
- •Alignment/safety issues often trace back to low-quality data and RL setups
- •Weak sandboxes and brittle verification enable unsafe or unintended behaviors
- •Safety is an industry-wide responsibility, including environment/data companies
- •Robust verification and mitigations “down the stack” raise overall safety
- 14:49 – 16:33
Hiring and org design: building the environment ‘supply chain’
They close by discussing hiring needs across operations, research, and infrastructure. Jerry emphasizes environment-building as a supply chain combining engineering, research, data/ops, and process, and they seek high-agency people who can leverage modern tools to solve hard problems.
- •Hiring for operators/generalists spanning product and ops
- •Growing research: post-training, benchmarks, and evaluation expertise
- •Platform/infra is becoming a bottleneck for hosting complex simulations
- •Looking for high-agency builders; backgrounds matter less than problem-solving
