Skip to content
YC Root AccessYC Root Access

Better AI Starts With Better Verification

Halluminate (YC S25) is building reinforcement learning environments and benchmarks to help AI models do knowledge work beyond coding, starting with finance. With a team of fewer than 10 people, the company works with four of the top five closed-source US AI labs and recently raised a $30 million Series A led by Oak HC/FT. In this Fireside, co-founders Jerry Wu and Wyatt Marshall sit down with YC Partner Jon Xu to share how testing browser agents led them to building simulations where models practice tasks like completing financial spreadsheets. They explain why accurate scoring and expert review matter more than sheer task volume, and how weak checks can teach models to take shortcuts. They also discuss their vision for simulated companies where teams of AI agents learn to work together. https://halluminate.ai Chapters: 00:00 — Halluminate’s Pivot From Evals to Training Environments 03:48 — Teaching AI to Do Financial Work 07:13 — Why Verification Matters More Than Volume 10:13 — From Finance to Simulated Companies 13:30 — AI Safety and Building the Team Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

Jon XuhostJerry WuguestWyatt Marshallguest
Oct 1, 202616mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:00 – 3:48

    Halluminate’s Pivot From Evals to Training Environments

    1. JX

      [on-hold music] I wanna welcome Jerry and Wyatt of Halluminate, who went through the Summer '25 YC Batch, now building one of the fastest-growing companies in the last year. Uh, and you recently announced your $30 million Series A led by Oak HC/FT. Congratulations, guys.

    2. JW

      Thank you so much, John.

    3. WM

      Thanks, John.

    4. JW

      Yeah.

    5. JX

      Really happy to talk to you and learn more about all the cool stuff that you guys are doing. So let's first talk about what is Halluminate? What do you guys do?

    6. JW

      At Halluminate, what we do is we build frontier RL environments and benchmarks to push what models can do in non-coding, non-coding knowledge work domains, starting with financial services, right? And, and a very simple way to, to sort of visualize our impact on the stack is, you know, if you've ever used a foundation model or product to not write code, but do things like generate spreadsheets, generate PowerPoints, uh, help due diligence your deal, or help ma- help you make investments, right? A lot of those gains in financial and economic reasoning and literacy have come from the, the benchmarks and environments we've built and scaled with the leading frontier labs, uh, here in the United States. In the last year, like you said, our team of less than 10 people actually have, have scaled to over mid eight figures in run rates, and we currently work with four of the top five closed-source US labs here, here today. Um, so it's been a crazy few years. Um-

    7. JX

      That's amazing.

    8. JW

      Pre- crazy 12 months since then, yeah.

    9. JX

      How did you guys get together and work on this? Like, what gave you the insight? Because as I recall, this wasn't actually your first idea, or this was not the idea that, um, you were originally working on.

    10. JW

      Yeah. So Wyatt and I met, um, at Cornell. We've known each other for almost eight years. We've lived together for almost seven and a half years. Uh, pretty, pretty insane period of growth together. And when we first started Halluminate two and a half years ago, our, our original problem space that we were working on was, was evals, right? Evals and benchmarking, which for many of, of those listening now, it's a very well-known problem space today, right? But the idea initially started because I was, uh, working at Capital One Labs, and while at Capital One Labs we were training one of the first AI agents for finance, and, you know, the constant problem we kept hit- running into was, was, like, how do we verify, how do we test, how do we benchmark quality and efficiency, right? So we started Halluminate originally to solve this problem in evals. And while we were working on evals a- and really partnering with some of the leading enterprises and startups on, on this problem, uh, we ended up working specifically on computer use and browser use evals, right? And actually, we've worked with a few browser, browser agent companies from YC, Skyrun and Browser Use, um, on, on this specific problem. And, and the, the way we kind of realized there was this huge need for environments was because, you know, in order to eval and benchmark a web agent or a browser agent in late 2024, early 2025 on, say, like, flight booking, right? What we realized we needed to do to do our job was build a simulated environment for, say, you know, Google Flights or, or a shopping website or, or a checkout website, right? And, and we realized, like, this work is incredibly hard, but, but building this environment was also incredibly valuable, and our hypothesis was not just for evals, but actually for post-training, right? And this is when we saw some of the early scaling laws around RL and, and DeepSeek and, you know, reasoning models as well. So, so we basically made a pivot and said, "Hey, let's take this, like, competency we've built in evals around building verification frameworks and environments, and let's essentially try to apply that to building very high-quality post-training environments for model builders, i.e. the labs," right? And, and so we made that pivot early 2025 and have basically been scaling that thesis ever since.

  2. 3:48 – 7:13

    Teaching AI to Do Financial Work

    1. JX

      Yeah.

    2. JW

      Right.

    3. JX

      And, and this is sort of around the time when computer use and, and the first sort of Sonnet-3.5 browser use model had already been permeated in, in a lot of the s- the, the solutions for automating knowledge work.

    4. JW

      Exactly.

    5. JX

      Yeah.

    6. JW

      Right.

    7. JX

      Yeah.

    8. JW

      And, and so you actually hit on a good point. Like, when we first started building environments, we were focused on computer and browser use environments, right? Um, but sort of, uh, a- as we started working on environments continuously, the reason we, we eventually, you know, moved into this specialized vertical we work on today, which is knowledge work and finance, is we, we sort of realized that everyone in the space was working on coding or computer use. But it was very obvious to us that, you know, if, if we sort of extrapolated where the models were going and where adoption would be, we had strong conviction that actually these enterprise knowledge work non-coding use cases, we believe would actually accrue more value in the long term, right? Because, I mean, put simply, more people in the world work non-coding jobs than coding jobs, right? But models were still far behind in these categories, and our hypothesis was, like, it's just fundamentally there's a lack of training environments and data and benchmarks for these domains. So we, we, we really took that initial competency and, and eventually specialized into the verticals that we work on today.

    9. JX

      Now, for people that are completely not familiar, I just wanna level set. Help us understand what really goes into an RL environment and how that's sort of different than just pure data that's used for-

    10. JW

      Mm-hmm

    11. JX

      ... a lot of the other, other training methods.

    12. WM

      Yeah. An RL environment is actually pretty simple. You give the agent an objective, so some, you know, prompt or problem statement. You have some environment that it runs in, so it's-- it could be a computer, could be the browser, could be the terminal, and then you have some ability to evaluate the work that the agent has done. If you're gonna do an eval, then you just take that and that's the score for the eval. If you're gonna do post-training, that becomes the input to the reward function that actually is in the post-training pipeline. So the, you know, core loop of that is not, not that complicated, but I think what makes it complicated, especially for us in the kind of world of, of knowledge work, is a lot of the verification is less clear than in something like coding, where you might just be able to run unit tests or even in some of the stuff we've seen in math where, you know, the lean compiles. Knowledge work is-- it's harder to verify those. It's a little bit less clear.

    13. JX

      So for people who don't know, uh, can you tell us an example of an RL environment and a model- Doing the learning loop with an RL environment

    14. WM

      Yeah. Take a Excel modeling task, like building an LBL model. You might give an agent a prompt, uh, for it to finish a, a half-constructed model, and then you give it sort of that initialized spreadsheet with maybe some pieces missing or broken, um, and then possibly some supporting context to kind of fill in the pieces or, you know, references that the agent needs to make in order to find the right information. The agent is then gonna solve that task. It's gonna try to create the correct output model, and then as part of verification, you're able to take that model that the agent created and compare it to, you know, the ground truth, basically golden Excel model, and you can do things like compare, did it get the correct formulas in the right cells? Did it derive the information from the right sources? Did it format things the way that the formatting checklist specifies? Because these are all things that, you know, real bankers care about, um, and I think when you, when you put those things together, that's how you get that holistic score of, you know, this agent solved this task this effectively.

  3. 7:13 – 10:13

    Why Verification Matters More Than Volume

    1. JX

      If we look at the construct of both an environment, a set of tasks, and evals, what, what is the most difficult part about creating something that is of high quality? Because I suspect that, you know, you can easily get into this sort of volume, um, because of the scaling laws. Obviously, the, the volume is one of the, the factors, right? But, um, but really to create something of high quality is actually quite difficult, I suspect.

    2. WM

      Yeah, I mean, the trade-off of quality and volume is kind of the core thing to overcome in environment building. I think the quality of your environment is directly determining the quality of the model that's post-trained on that environment, and so having high-quality environments is the most important thing that you can do if you're building environments. It's better to have fewer that are better than a lot more that are poor quality. And when it comes to thinking about quality, verification is definitely the most important part of that. So when we've seen this now, um, I think recently especially, lots of this, this misalignment, uh, you know, articles, uh, that have been coming out, and typically the result is that there's some lack of quality and verification when it comes to, is this task being verified the right way, or can the agent get credit for work that it didn't do and it doesn't deserve? Or vice versa, is the agent actually solving the task and it's not getting credit that it should get? Either of those situations, you're gonna end up with a model that's learning the wrong stuff because that verifier is directly teaching the model, so if the verifier's not aligned with the intention of the task and, like, what a successful output looks like on that task, then you're gonna get models that are learning the wrong things.

    3. JX

      And, you know, as you look at how you're constructing the verifier in these models, um, how are you going, how are you thinking about this reward hacking and, and the, the various ways in which you can asymptote in the amount of learning that the model ultimately does?

    4. WM

      Yeah. There's a few ways, I think, when you use, like, the word reward hacking. Um, obviously there's, like, the hugging face issue, you know, actually breaking out of the, the sandbox, finding the answer key, knowing the answer before it needs to solve the problem. Um, but there's other and maybe softer versions of it where the agent does something, you know, maybe it takes a shortcut. It doesn't show its work.

    5. JX

      Mm-hmm.

    6. WM

      And in real life, if you wanted to give some task to, like, an analyst at an investment bank, that analyst needs to prove out, you know, the model, whereas the agent maybe just goes right to the answer and, and you don't wanna give the, the agent credit in that case because it didn't solve the task the right way. So, you know, I think understanding the different ways that reward hacking can kind of manifest itself is very important. It's also really important to have the quality assurance and, you know, oversight on task creation and, and actually not just basically pumping out tasks and training models on them, but making sure with subject matter experts that these tasks are aligned, that the scores that the verifiers are producing are aligned with what an expert would give.

    7. JX

      Mm-hmm.

    8. WM

      Um, so I think again, that's where the scale and quality trade-off comes up, but quality assurance is definitely, it's a huge part of producing good tasks. You can't really

  4. 10:13 – 13:30

    From Finance to Simulated Companies

    1. WM

      skip that.

    2. JX

      So as you look forward to how to continuously achieve that, to grade it, to bring it into greater scale, what does that look like? What's, what's kinda coming up next for you guys?

    3. JW

      Yeah, I mean, there's a few, like, super exciting, you know, areas that we wanna continue growing, right? A big reason why we chose finance as our starting vertical in, in, within KnowledgeWork is, one, finance is the world's largest, you know, s- knowledge work service industry, right? On a head count basis, it's, it's one of the highest value. It's somewhat verifiable, which has a lot of nice properties for RL, and there's so many different subgroups within finance, right? Uh, investment banking, private equity, uh, you know, equities research, et cetera. But most importantly actually, we, we sort of had this belief that really being excellent at building finance environments and, and sort of teaching financial reasoning would generalize, uh, to a lot of other knowledge work domains, right? Such as, um, consulting, accounting, FP&A, even insurance, right? 'Cause a lot of these knowledge work workflows kind of come from the same, you know, capability curriculum, right?

    4. JX

      Yeah.

    5. JW

      That, that we can hill climb on as an industry. So what we, we're expanding into as this company and with all this funding that we've acquired is growing into these adjacent knowledge work domains, right? And we've already, uh, seen a lot of our skill sets and quality research translate into many of these other categories that our customers are super excited about.

    6. JX

      So as models get smarter, how do you imagine that the environments that you build continue to evolve over time?

    7. JW

      Yeah, it's, it's a really great question. Eh, in, internally at the company, we've noticed, you know, like I said, we've been doing this for almost 18 months, right? And we've noticed something that we internally call the Moore's Law of RL environments, which basically means every six to eight months, the complexity or scope or, you know, size of the environments that we train on roughly doubles, right? And this roughly corresponds with the, the intelligence of the models, right? Or another way you can think about it is, like, if we wanna keep pushing the frontier of what models can do, we as an industry need to roughly be able to double or increase the complexity of the environments that we can build and scale. For model training, right? And if you extrapolate this out, our, our belief at Halluminate is that, you know, today we're building these environments of PowerPoints or Excel or e- even, you know, email writing and, and a few of these apps together. I think in six to eight months we'll be building environments of multiple agents operating in a team-

    8. WM

      Mm.

    9. JW

      Right? W- doing, uh, you know, large chunks of work. I think, you know, in the future what we are actually gonna be building and delivering are actually simulated companies, right? Or simulated industries, simulated governments that, that teach agents not just how to do work, but how to be autonomous operators and coworkers in our economy and our society, right? And, and these large digital simulations, uh, of, of real world work is, is, you know, ex- gonna be extremely hard to build, right? Uh, but they're really gonna be necessary to continue to push the frontier of, of models in the future. And, and whether it's labs or enterprises, we believe they'll rent and sort of purchase time in our environments, right? That's sort of our vision of, of where

  5. 13:30 – 16:33

    AI Safety and Building the Team

    1. JW

      the space is going.

    2. WM

      There's been a lot of, um, dialogue about alignment and safety-

    3. JW

      Mm

    4. WM

      ... of these models. How do you guys think about that as the underlying infrastructure to actually make the models a lot smarter?

    5. JW

      Yeah, I mean, I think alignment is a huge issue, but I think the really interesting thing, for those of you interested in working on alignment, is if you really look at the reports, alignment and safety issues that emerged was really, you know, a product of poor quality data, and specifically poor quality RL environments, right? If you look at the Hugging Face paper, [clears throat] a lot of the, like, behaviors that allowed it to e- escape was, you know, poorly built sandboxes, a lack of robust verification on, on process techniques, et cetera. And so, like, I, I think alignment and safety is an industry-wide problem, right? I don't think it's just, uh, for, like, a lab safety team or, like, an independent evaluation group. Like, really it's, it's a industry-wide problem specifically for, for data and environment companies like us, right? And, and, like, having robust quality and, and mitigation of issues all the way down the stack with, with companies like us working on the frontier is, is some of the highest impact ways for us as a whole industry to ensure that what we build is, is truly safe and aligned for everyday use.

    6. WM

      If someone wanted to work on that with you guys, um, are you guys hiring? How should they be thinking about joining you to do that?

    7. JW

      Yeah, definitely. Um, there are so many problems to solve, right? Building an environment is very much like a supply chain, right? You need, you need to layer on research and engineering, uh, data, people, and then just the whole process and financing around it, right?

    8. WM

      Mm-hmm.

    9. JW

      So o- on the op side, we're hiring for, for great operator, great generalist operators that can work both in product and op senses, right?

    10. WM

      And I think what you touched on with these simulations is I think we are moving, you know, more into the world of engineering as, like, the fundamental bottleneck. Um, so something we're really excited about, we've been building out the research team at Halluminate, um, so we're looking for people, you know, who know really in-depth post-training and benchmarks evaluations kind of from the research side. Uh, also, you know, platform and infra. I think some of the infra challenges around hosting these simulations are, are just thorny problems. Um, and I think just generally, you know, we've, we've seen people who can come in and can use tools, can use AI tooling to solve the problems, you know, be really successful and, and it doesn't really matter their background. It doesn't really matter previous roles. Um, I think the... I don't know. The, the future of work I think is changing a little bit, and it's just driven people who can kind of have agency and, and solve problems, and that's really what we're looking for. Fantastic. Well, congratulations to you both on the amazing progress, uh, and on the recent fundraise. Um, thanks for joining us today.

    11. JW

      Thanks so much, Jon.

    12. SP

      Thank you, Jon. [upbeat music]

Episode duration: 16:33

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode 0-g2-PRrOdw

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.