EVERY SPOKEN WORD
40 min read · 7,665 words- 0:00 – 0:47
Intro
- AMAmjad Masad
You saw the SpaceX S1. It's like, oh, $30 trillion. [chuckles] It's like, what is the world GDP? 100 trillion?
- AAAlex Atallah
Both Stripe and OpenRouter really want lots of new companies in the world. We don't want everyone to be a part of one giant company. When the models get more intelligent, the risk actually will continue to get higher, and yet no one new is taking responsibility.
- AMAmjad Masad
We're gonna slowly realize how good we've had it with the, like, deterministic code. Remember the days when computers did exactly [chuckles] wh- what we told them to do.
- AAAlex Atallah
Are we going to prevent the models from deceiving users during training runs predictably? Will, like, a model that's big enough and powerful enough suddenly stop deception and stop sandbagging? Um-
- 0:47 – 5:08
Inside the Stripe acquisition
- ETErik Torenberg
Welcome to a16z podcast. We're here with Amjad o- of Replit and, and Alex o- o- OpenRouter, and this is the first podcast that Alex has done since the acquisition, so we're really excited to have both of you.
- AMAmjad Masad
Mm-hmm. Thank you. I'm excited to be here.
- ETErik Torenberg
A- Alex, let's, let's, let's start with that, actually, if you can briefly, um, share, um, obviously, you know, m- massive acquisition. Uh, Amjad is an investor. W- w- we're, we're also, of course, an investor. Um, the biggest, uh, shareholder, but, hey, who's, who's, uh, who's counting? Um-
- AMAmjad Masad
[laughs]
- ETErik Torenberg
Um, Alex, why don't you give us a little bit of the backstory, like, how does a acquisition like that even, uh, even happen? Like, do you get a DM from Patrick one day? Uh, you know, what, what, uh, what, what, what can you share?
- AMAmjad Masad
Hey, how much for OpenRouter? [laughs]
- ETErik Torenberg
[laughs]
- AAAlex Atallah
I had talked to Will Gabrick, um, the, like, uh, Stripe president a w- long time ago, couple years ago, when we were doing our Series A. And, uh, and we just... We stayed in touch. We had, like, a lot of ig- like, Stripe work streams going on with various teams at Stripe. Um, so there were always sort of, like, things that we were doing with Stripe. Um, we presented at Stripe sessions. Uh, and, and, and so it always kind of... They always kind of felt close. And then, yeah, just in, in July, I believe, um, they reached out, wanted to chat. Met, met, like, both of them in person. And, uh, and it kinda just progressed from there fairly quickly. Um, they're very efficient and, you know, they were very, like, founder-friendly about the experience. Um, I was really impressed with the whole thing.
- AMAmjad Masad
Did it, did it... Did you want to sell? Like, did it cr- even cross your mind before they reached out?
- AAAlex Atallah
No. [chuckles] We were not thinking about that at all. Um, I, you know, did really respect the company, do really respect it today. And, you know, of possible acquisition options for us, um, it was, I think, my top choice. And, uh, so it was, like, an interesting idea. Uh, and, you know, as we kinda, like, fleshed out the reasons why it would make sense for both companies, it got more and more interesting. Uh, it was really clear how, like, aligned they were with us kinda having autonomy over the brand and roadmap and product, and keeping OpenRouter doing what it's already doing, just much faster with, um, much more... With, like, a, like, a, a much more serious go-to-market plan. And, um, and then, you know, some better together stories between the two products, uh, and the two, the two companies. And, and then culturally, there were just, I, I think ver- uh, in terms of, like, mission and values and building, like, a neutral, trusted platform that businesses can depend on and scale on top of, that's also really developer-friendly with the best possible developer experience for, you know, to encourage new companies to emerge. Um, you know, that, that alignment was there. And there was a bigger picture kind of alignment too, where, like, you know, both Stripe and OpenRouter really want lots of new companies in the world.
- AMAmjad Masad
Mm.
- AAAlex Atallah
You know, we don't want everyone to be a part of one giant company. We want to, like, create really good incentives and, uh, and really sort of easy, streamlined workflows for people to start new companies and grow them successfully, and, you know, make both lifestyle and venture-backed businesses on top of really good, like, reliable and price-efficient infrastructure and, and, like, a marketplace that works. Um, and you know, I really want that future. [chuckles] And Stripe demonstrated that y- they've been wanting it and building towards it for, like, many, many, many years. Um, and in many ways, like, payments and inference are going to blend together for companies of the future.
- 5:08 – 8:53
Why OpenRouter needs more startups
- AMAmjad Masad
I'm, I'm curious. Uh, I understand why Stripe's incentive is to have, you know, much more vibrant startup ecosystem. I understand the kind of moral argument in why you would want that. But why is that good for OpenRouter? Like, is your model for OpenRouter, is that it's a, like, a network effects business? Is it, is it, like, a network-
- AAAlex Atallah
It, well, for us, I think the, the... Yeah, there, there's, like, a couple different problems OpenRouter solves. One is we're, you know, allowing you to build a company that uses AI, um, or, like, augments intelligence with unique data and, and other services, uh, without model lock-in-
- AMAmjad Masad
Mm-hmm
- AAAlex Atallah
... without vendor lock-in. Allowing you to kind of, like- Be on the Pareto frontier continuously as the ecosystem grows it. And, um, to do that, it, it's a lot of work because there's all kinds of little lock-in that appears. Um, we also, uh, y- you know, I, I, like, we want companies to feel like the, the, the best way-- We, we-- That they can, like, add more than just prompts on top of a single model. There's, like, a lot more to building, um, like, unique intelligence, and I think a big component of that is neurodiversity. You really w- need the, the power of multiple models that are trained in different ways, including some of your own-
- ETErik Torenberg
Mm-hmm
- AAAlex Atallah
... um, to do more than ChatGPT or Claude would on the task if, like, someone who's thinking about buying from you as a potential customer is wondering, "Well, what if I just use the model directly?" Um, like, how do you really show that you are significantly better and able to build a business that, that matters? Um, and I, I think a lot of it will involve neurodiversity and blending, like, powers and, and good data from multiple models. Um, another component is, is, like, helping people get really good cost efficiency. Like, j- there are a lot of businesses that just cannot, do not emerge until they become cost-effective. And h- like, creating an environment where we can help drive down costs by building an efficient market, uh, is cr- like, crucial to making that happen. Otherwise, like, you know, m- m- uh, uh, [chuckles] why lower my prices [chuckles] as a, as a provider? Um, we have a captive market. And, uh, and like, you know, I think that's, like, a key point of marketplaces that was just totally missing from AI before we showed up. There was just one player, OpenAI. It w- could've been like a, uh, a very strange world. I'm not saying that we did all the work [laughs] . No, of course not. But, like, helping people, like, choose new models and explore new models and learn, like, what makes, um, you know, a cl- a closed source or open weight model, uh, like, g- actually good at your task involves, like, seeing what the whole ecosystem is doing and learning automatically. Like, LLMs are not things where you can just enumerate all the features on a webpage. It's impossible. You, you have to see how they're being used to know what they're good at.
- 8:53 – 12:26
Enterprises are picking open-weight models
- ETErik Torenberg
So you, you started the company three years ago. I'm, I'm curious what has surprised you the most about sort of the evolution of open and closed source models to the present, um, as it relates to model performance or just people's, uh, perspective on the variety of models they have available to them?
- AAAlex Atallah
People have been more open-minded than I thought they would be towards open weight models. Like, typically, there's like a, there's a lot of brand... Uh, e- especially in enterprise, typically, there's a lot of, like, brand trust, where, like, an enterprises, you know, enterprises in general are like, "Oh, I'm... You know, I don't really know how to tell the difference between these things, so I'm just going to like, you know, buy the one that all the other credible enterprises are buying." And, and there's a lot of, like, enterprise, um, lock-in with, like, a mentality like that. Um, and that just didn't happen that much. Well, they-- It did happen, like, a bit, but, uh, we saw a lot of enterprises, uh, want to explore new models. Uh, it, it's simply like m- it, it was mu- it was v- very much good for a marketplace. Um, enterprises wanted to diversify outside of just the proprietary frontier model labs, um, both for, like, cost reasons and for differentiation reasons. You know, they wanted to, like, own their intelligence so they could, um, you know, I think, like, one, keep their, their talent, like, have, like, an internal AI practice. Like, AI is just, like, a huge strategy topic. It's not like you go to your board and you're like, "Oh, yeah, um, you know, we fixed the AI problem. Uh, quarter complete." Like [chuckles] the b- your board is, like, asking you every month, like, what, you know, "What's, what's next for the AI, internal AI team?" And all, like, every single enterprise, it now has just this in, you know, internal AI team that they're developing, where they, they need a strategy behind it. Um, it's not just a, like, you know, we, we check the feature off, we, like, set up the database, and we're done. Um, and so I think that that, uh, dynamic has resulted in just, like, a desire to explore and diversify, um, and a, a desire to kinda, like, figure out how to reduce costs and figure out how to do benchmarks for the first time in the company's life. I have been, like, surprised there hasn't been more benchmarking, like, more companies creating more benchmarks. It's starting to happen, and I think, like, eventually, we'll, we'll see a lot more of them, um, to demonstrate, like, oh, yeah, like this thing is better than using, you know, Claude Direct. But, um, like, I think that's gonna be, like, a bigger focus for this internal AI group at every company, um, evals. And I think, like, Amjad's been doing that a lot at Replit, for example. Like, you guys have done a lot of cost per task research. Um, you guys have been You know, you, you've made like a Doom Loop Resc- you, you've kinda been like experimenting with new ways of using agents like Doom Loop Rescue and, and, uh, like helping bring those to developers. Um, so, uh, like more, more of that kind of research I think is gonna pop up internally everywhere, um, for all of those reasons.
- 12:26 – 19:52
Why owning your intelligence matters
- AMAmjad Masad
Yeah, I, I think, um, [clears throat] uh, Satya, uh, CEO of Microsoft, has been very, uh, very, you know, prescient on this and also very articulate on, um, why companies need to own their intelligence. Ultimately, in the same way that, you know, we had dot com companies and then every company became an internet company, like every company employs people that know how to like build websites and be on the internet. Similarly with, with software, every company has software engineers. Every company needs some AI practice, AI capability, and that will compound over time. The knowledge, the intelligence inside the company, the, the, the use case model fits, like which models actually work for them, how do they save, save money. They need sort of that independence. The other thing that, um, I think Alex Karp of, um, Palantir has been talking about is that there's like a risk that when you work closely with the foundation model companies, is that they're gonna move into your business. And we've seen that, you know, with Figma vis-a-vis Anthropic. We've seen that with like now Harvey and OpenAI and, you know, it's, it's really hard to partner with them and, uh, uh, b- because, you know, they're... they see the world as their potential market, right? Like I mean, when, when they talk to investors that are like, you know, you saw the SpaceX S1. It's like, oh, $30 trillion. [chuckles] It's like what is the world GDP? 100 trillion? And so there is a sense in which these companies are different than other generation of companies. It's, it's harder to partner with them because their ambition is, is, is such that they wanna, they wanna be, they wanna subsume a big part of the economy. Um, and increasingly like what we're thinking about at Replit is kind of i- in a similar vein to what Alex have sort of innovated is, uh, Replit is becoming more of an independence layer inside, inside enterprises, where we create a layer of indirection between, um, between you and the, and the models, and we get you the best token at the cheapest price. Uh, but also we, um, we also create, uh, an abstraction layer on top of the cloud as well 'cause, you know, you should be able to deploy to AWS and Azure, and you should be able to use Databricks and Snowflake and so on. Um, and so increasingly I think there is, not just with AI, with all of technology, there needs to be more platforms that help companies gain, gain independence.
- ETErik Torenberg
It, it seems like, yeah, what OpenRouter did for their segment, you're, you're doing for other, other areas of the business.
- AMAmjad Masad
Right. Yeah.
- AAAlex Atallah
There was this... This kinda reminds me of this tweet I saw the other day when somebody was like, "Basically all companies are building the same thing now."
- AMAmjad Masad
Yes.
- AAAlex Atallah
Everybody's building an agent loop with like notifications and, and context, uh, like third party, uh, connectors and context management and memory and, uh-
- AMAmjad Masad
Sandbox
- AAAlex Atallah
... uh, sandboxes.
- AMAmjad Masad
Mm-hmm.
- AAAlex Atallah
Uh, you know, web, you know, agentic web search-
- AMAmjad Masad
Computer use
- AAAlex Atallah
... and an always-on agent on top of it-
- AMAmjad Masad
Yeah
- AAAlex Atallah
... and notifications, and it's like this product is showing up everywhere.
- AMAmjad Masad
Yes.
- AAAlex Atallah
Um, in a way, yeah, it is like showing up everywhere, but it, it also fee- kinda feels to me like these are just the new table stakes primitives. It's kind of like, like a 2005 version of that tweet would be, "Oh, everybody's building the same thing. It's like a database, a user's table-"
- AMAmjad Masad
[laughs]
- AAAlex Atallah
"... a sign-in page, a sign-up page-"
- AMAmjad Masad
Yeah.
- AAAlex Atallah
"... a profile page-"
- AMAmjad Masad
[laughs] Yeah
- AAAlex Atallah
"... a log-out page. Like everything's the same."
- ETErik Torenberg
[laughs]
- AAAlex Atallah
It's like [laughs] it's not. [laughs]
- ETErik Torenberg
[laughs]
- AAAlex Atallah
[laughs] There's like a lot of differentiation really. It's just there's like a, there's table stakes needs for AI just like there are table stakes needs for the web.
- AMAmjad Masad
Yeah, and, and I think, um, as well, like inside the enterprise, making these products actually do real work is still unsolved problem. Like you can use Muse in your personal life and connect it to your credit card and bank accounts, and no one's connecting Muse to their enterprise data or even Grok Bot and things like that. Um, I think there's an even more emphasis on data sovereignty and, and security. So we, we spent the past, you know, year almost, uh, working on making, uh, making Replit, uh, deployable on your own cloud, basically on-prem, like bring your own cloud. Like two years ago, I would've thought I would never do this, uh, because, you know, just like, yeah, the cloud is the future, like software as service, all of that. Uh, but, but now actually we've sort of like reverted a little bit back to a world where companies are a little bit more protective because there's so many ways in which data can leak. Like all these agents that pe- people are using, I mean, there's all these screenshots on Twitter, I don't know how true, where Instinct is like, or Muse is like mixing people's data. Starts to call you by a different name or, or something like that. Uh, and so yes, the kind of consumer stuff is kind of obvious, but on the enterprise, there's still tremendous amount of work for the entire industry to do in order to actually get these things to be useful and productive at work.
- AAAlex Atallah
Do you, Amjad, are you doing any, like do you have like a custom personal agent other than Muse or Instinct that you use for like work stuff that you've been building?
- AMAmjad Masad
Yeah. I mean, I, I built something on Replit like a long time ago. Uh, it started as like a sort of a CRM agent initially. That was the main problem that I had. But slowly, like we added features to it, and it's sort of like doing more and more things. But what's, what's really interesting is that the more I connected Replit to my, all my stuff, it sort of like started answering all the things for me. Um, and so increasingly, like the platform itself is like subsuming these s- sort of like these domain-specific agents that I built. I do think that there, there is some, in some ways you want something that is sometimes like focused on one particular thing, and you don't want it to be able to do everything. On the other hand, once you have your entire company's context in one place, it's really cool to join across totally different domains. Like when I ask it a question, it can like look at my sort of personal chat history, join it across the GitHub repo, across Salesforce. And so it says-- Uh, like it will link like random things, and it's like, "Oh, you met this guy like a year ago at a conference. Yeah, I see it on your calendar. And by the way, someone else from their team is in discussion with your sales team," and it creates all these different synergies and, and when I go into a meeting, I like, I, you know, I, I've connected a lot of different threads and, and I'm, I'm making much more progress on a deal or, or something like that. Uh, so, so this is where it's, it's sort of trending now.
- 19:52 – 23:36
The case against the god-agent
- AAAlex Atallah
Well, I, I think I, I, I'm like a little bit-- I'll take the counter on that.
- AMAmjad Masad
Yeah.
- AAAlex Atallah
I think that the, the worst part about doing cross-domain joins with your personal agent is that the, the more work you give it to do, the more understanding of what's going on you're sacrificing. And yet no one, no one new is taking responsibility for that sacrificed understanding. Like agents don't have any responsibility. You can't like... You, like if there's a fixed level of cortisol that the whole company can tolerate between everybody, and, uh, you know, you like de- y- [chuckles] you, uh, you... Like I, I wanna be like less stressed about some area if I'm going to be like sacrificing my understanding of it.
- AMAmjad Masad
Mm-hmm.
- AAAlex Atallah
Someone else needs to take the cortisol.
- AMAmjad Masad
Mm-hmm.
- AAAlex Atallah
Well, the, [chuckles] the, the agent doesn't like take on any of that responsibility.
- AMAmjad Masad
Mm-hmm.
- AAAlex Atallah
Um, and then like a universal agent that's doing all things, you know, I, I can't like adjust how much understanding I'm sacrificing in all the different things. It's, it like points me a little bit towards, you know, maybe like down the road, like the sub-agents that people use will be like very vertically focused. Maybe we have like a chief of staff type agent that, you know, coordinates between them. Um, but I feel like you do need like vertically focused agents where you're like, "Okay, like th- this agent is more responsible psychologically for these things. Um, and like I want like quality checks that make sure it's doing those things correctly. And I-- it don't, it doesn't need to focus on anything else. It just has one focus area." Um, I wonder if that's gonna like help people like at least get like a weird loose sense of responsibility on top of agents.
- AMAmjad Masad
Fascinating. So, so you're saying general agents create like a tragedy of commons of sorts? Um-
- AAAlex Atallah
Ki- kind of. Like I have a general agent that every day, um, looks for things that need me and tries to figure out what to do, and it's just like impossible to improve this agent. [laughs]
- AMAmjad Masad
Mm-hmm.
- AAAlex Atallah
Like I, it's, it's like every time I try to make an improvement, I end up like ignoring its output a- about a week later. Um, it's, it just feels like it doesn't really care about any of the like specific things it's diving into. Um, and, uh, like yeah, like kind of imagine having a chief of staff where they're very good at, you know, like drafting all of your replies across the whole organization. Um, and then compare that to something where you have like 10 chiefs of staff, each, uh, eq- each as competent as that one chief of staff, but they're all responsible for like individual sectors of, of what you, of what makes up your life. Like the latter, I feel like gives you a way of tuning how much understanding you sacrifice-
- AMAmjad Masad
Hmm
- AAAlex Atallah
... compared to the, the gain you get from... Uh, uh, like basically I can like lean in more to the areas where the agent is failing for some areas, and then have agents with very good competency like take over my understanding of other, of other parts of my life.
- 23:36 – 31:13
Specialization, Adam Smith style
- AMAmjad Masad
Yeah, interesting. It's, it's sort of like, uh, almost rediscovering, you know, um, specialization, right? W- what's his name? The, the like the famous, uh, economist, uh, Adam, um, wait, like-
- AAAlex Atallah
Adam Smith.
- AMAmjad Masad
Adam Smith, like with a pencil kind of thing where [chuckles] you know, that, that was like a huge realization for humanity that like specializ- specializ- specializ- specialization is actually good.
- AAAlex Atallah
Yeah.
- AMAmjad Masad
The problem is like we kind of like overspecialized as like a civilization, and I think overspecialization is, is, um, is oppressive in its, its own ways. I mean, I think we've, y- you know, um, sort of like, you know, there's the Marxist theory of alienation, right? The, the idea is, um, uh, because of overspec- uh Uh, you know, people are doing, just focus on one thing. Uh, they, they do not see the fruits of their labor. They don't, they don't actually know what their impact is on the larger organization or the product they're producing, and therefore, they actually kind of feel depressed and detached and, and you're kind of acting like a machine and you're not actually fully, fully human. And so may- maybe there's like a bit of a reaction to that and, and, and I think with our agents we're like, oh, th- there should be like one god, god-agent. But in fact, specialization is actually like really good for machines, and that's like the, the point that you're making. And like humans should be general, but like machines should be ultimately a lot more specialized.
- AAAlex Atallah
The problem with I think what I'm describing is that we don't know what good looks like. There hasn't been a system of specialized agents that feels as elegant as like ChatGPT or Claude or Muse, uh, where you're basically just talking to one thing only. It, it's, it's yet to be discovered. Like maybe like OpenAI just launched Dots. Um, I think they're kind of like experimenting in that direction. Grok, uh, Grok Bot I guess kind of counts, but-
- AMAmjad Masad
But, but they're all-
- AAAlex Atallah
That's-
- AMAmjad Masad
They're all very general. Like I think the idea behind Dot is that it, it's like, it's like, uh, it's like your digital double. As at le- at least that's what I understood it.
- AAAlex Atallah
Well, when I saw Grok Bot, I don't know what, what, what's gonna-
- AMAmjad Masad
Yeah
- AAAlex Atallah
... I don't know that much about Dots yet, but, uh, they just came out. But Grok Bot, when it first came out, like the first use cases that I saw people talking about-
- AMAmjad Masad
Mm-hmm
- AAAlex Atallah
... were, "Oh, whoa, I can make two bots, one that knows my bank account-
- AMAmjad Masad
Right
- AAAlex Atallah
... and one that knows my Twitter account."
- AMAmjad Masad
Right.
- AAAlex Atallah
And the two bots like don't have the credentials from each other.
- AMAmjad Masad
Yeah.
- AAAlex Atallah
But they can talk to each other if like they need to get something done. Um, there's no credential sharing though, and that was like one sort of big unique thing I saw like a couple time-- like pop up a couple times that people seemed to like. Uh-
- AMAmjad Masad
But it looks like Muse is, uh-
- AAAlex Atallah
But I wasn't sure
- AMAmjad Masad
... Muse and Instinct are, have a much stronger product market fit than Grok Bot. Um, and p- and perhaps-
- AAAlex Atallah
Seems like that
- AMAmjad Masad
... perhaps it is because you don't have to worry about creating these, these domains. But, but I think maybe personal agents are different than, than work agents, and I think your, kind of your critique of general agents is, is more, uh, pertaining to, to, to work and to enterprise, which I, I sort of agree with. And, um, you know, th- there's also all sorts of, uh, data access considerations. I think a- as CEOs, we can have general agents because we have admin access. Uh, but you know, you're-- for individual employees or certain teams, they can't have like truly general, you know, fully context-aware agents because there's, there's access control, control issues. So y- you'll have to kind of work on-
- AAAlex Atallah
Sure. Yep
- AMAmjad Masad
... on something like specialization. Ultimately, I also think we need to figure out what does agent-to-agent communication look like. I don't think there is like good protocols around that just yet. I, I don't think that agents are trained to handle that very well. I think we've seen-- It seems like the next generation of OpenAI models are trained to do agent collaboration, because we've seen it in the Hugging Face hack where they started helping each other and sort of emerged naturally. But there's also needs to be some way in which, like an agent can't convince another agent to kind of give it, give it information that it shouldn't give it, uh, give it. Like th- there's, there needs to be like data, um, isolation and, and, and, and, and proper ways in which these agents communicate. Y- you almost don't wanna communicate them fully in natural language. Maybe there, there's like some other DSL protocol that, that they need to follow.
- AAAlex Atallah
I really, I, I, I think one of the cool potential applications of JEV and other decision models like it is gonna be alignment. You know, checking to see if a tool call or like an agent-to-agent, you know, communication is aligned. 'Cause there's just so many tool calls. Like you really need like a cheap, fast model, um, if you're gonna block something like that. Then, uh, like a really, really fast decision model that just classifies and gives feedback on rejections, um, might be like a really good way to bridge the gap between agents and from agent to infrastructure too. So I haven't seen like-- And we have like a, a, a little prototype that we're running internally, um, at OpenRouter, but, um, I kinda think a good-- I, I think, like it could be like an interesting alignment study for people-
- AMAmjad Masad
So you're using it for policy enforcements?
- AAAlex Atallah
Yeah. Like imagine, um, uh, looking at the system prompt and the current tool call being made and be like, like, you know, is this aligned with the [chuckles] the system prompt of the original agent and with these like extra guidelines that maybe we didn't tell the agent about? Uh, for example, let's say you have a bunch of agents that are instructed to do, um, to like red team the, like some, some new product, and they cannot, should not, uh, be able to access the internet. Um, and if they ever do, they should stop right away. Um, but you might not wanna like s- explain all of that to the agents doing the red teaming. You might want them to try to like break out of the sandbox, um, and act like bad actors. Like what would a bad actor do? It wouldn't be like, "Break out of the sand-- You know, try to like break into this company, and the moment you do, stop. Don't do anything else." [chuckles]
- 31:13 – 34:13
Models training their replacements
- AMAmjad Masad
I, I wonder another thing about specialization and, um, sort of what you're talking about. There, there's a lot of talk of, uh, recursive self-improvement. Um, th- there's something I don't think is getting a lot of, um, uh, sort of discussion, which is, uh, models training their replacements is sort of like, you know, you can think of it as a just-in-time compiler. So the way just-in-time compilers is, um, you know, as you're executing dynamic code, you know, the, the, the interpreter realizes that there's an opportunity to optimize. It will emit machine code on the fly, and that's a, a lot more optimized. So you can imagine models, like general models, you're kind of doing something with Opus or some- some of the Astra, some of the big models, and they realize that the use case is, is limited, uh, or you prompt them some way or some other agent obs- uh, observing and realizes that a use case is limited. And I think, you know, general agents have all these flaws that you just talked about, but also there's, there's more potential for them to be harmful. There's more potential for them to go off the rails and sort of on the fly trains a model, uh, that, that could be its replacement, but is, like, a lot more domain specific. Uh, and, and therefore it is cheaper and also, um, you know, less, uh, less vulnerable to prompt injections, uh, less, less, less harmful for, for, you know, because it's less capable. Um, and it's, it's, it's almost like a, like, you know, some system that's training machine learning models i- for specific use cases as it's monitoring the entire system.
- AAAlex Atallah
Like, would that specific use case be, uh, involve, um, unstructured text generation or a very, like, structured decision model, like-
- AMAmjad Masad
It could be unstructured t-
- AAAlex Atallah
... use case?
- AMAmjad Masad
Text, uh, generation. It could be decision models. Like, even the case of, of, of Jav, like, if you have, if you understand the, the inputs ahead of time, uh, you could potentially like, you know, take an off-the-shelf like Quinn or something like that and, like, train it specifically for that policy.
- AAAlex Atallah
But for... That makes sense to do for cost reasons, um, assuming that there aren't, like, really, like the mo- the model labs, the frontier model labs might make very low-cost models, um, that you can easily transition to ensure.
- AMAmjad Masad
But, but safety, safety as well, right?
- AAAlex Atallah
Oh, yeah. I see your point.
- AMAmjad Masad
They're so capable, and so I think oftentimes people are using these big foundation AGI-like models to like... It's like nuking a butterfly, right? It's like [chuckles] they're very, you know, most of the times, like a lot of the use cases, even unstructured use cases don't need that capable model.
- 34:13 – 45:12
Will smarter models deceive us?
- AAAlex Atallah
I wish there were more public evals about this stuff. Like the-
- AMAmjad Masad
Yeah
- AAAlex Atallah
... a lot of the evals about this are private. You just can't see, um, whether, like, when the models get more intelligent, the risk actually, uh, will, like, continue to get higher. Because I think there's also an argument to be made that alignment will get better, and the models will, like, start to, you know, avoid going off and, you know, hacking on their own, um, as we, as they get smarter and better at alignment, especially when it comes to agent-to-agent coordination. Like the, the... This is something Noam Brown, um, said on a podcast recently. Like, as the agents have gotten smarter, they've gotten just better at coordinating. Um, they're, uh, there's, they're still like... A- and, and, you know, it's unclear if they're going to, like, be harder to align than humans when they're, when we get more and more of them. But if we can figure that problem out, um, then a smaller model, like, will it be harder to align?
- AMAmjad Masad
S- so, so, uh, uh, Erik asked earlier, is, is that do you think it's true that smarter models are, are more aligned naturally or they're easier to, to align? Well, i- if you th- think back to the original sort of like rationalist, less wrong arguments for AI safety, there is this thing called the orthogonality thesis. The idea is that intelligence is orthogonal to ethics or morality or, you know, so on. Um, I don't believe that's entirely true with, with humans. I think people who are generally, like, more intelligent, kind of more, more educated t- tend to, tend to, not always, tend to, you know, be more considerate of animals, for example. Um, but, uh, but, but in, in, in, in machines, uh, I, I think it could go the, the other way because, you know, there's been quite a bit of studies on, on, on RL showing like, you know, how reward hacking and deception, they just, like, get better at it. And, like, the evals could be, could be deceiving because, uh, the model could be smart enough to, to To know that it's getting evalled. I mean, we already know this. It's been shown that if you do a lot of monitoring of chain of thought, they start lying in their chain of thoughts. Um, and so you add pressure almost on the chain of thought, and that, that kind of creates... And, and, and I think at some point, for you to do proper alignment evals, you need to run it for, like, months, right? You need to run this thing-
- AAAlex Atallah
Mm-hmm
- AMAmjad Masad
... for months on a, like a really large, you know, goal or task in order for, for it to, to truly kind of figure out whether it's aligned or, or, or not. Yeah, I, I, I always struggle with this word alignment. It just feels, like, wrong for so many reasons. It's, it's sort of, like, vague and, and, uh, sort of like aligned to what? Whose values? And, and so it just doesn't make the conversation easier. I think in this case, I'm talking especially about deception, like the model is actually deceiving its user.
- AAAlex Atallah
I mean, I, I mean, maybe this kind of reduces to, like, are we gonna solve align... You know, are, are we going to prevent the models from deceiving users during training runs predictably with like, you know, better... Like, it... Like, will, will, like, a model that's big enough and powerful enough suddenly stop deception, stop sandbagging? Um, and nobody knows the answer to that yet. So, like, it... At the point when that do- You know, if that ever does happen though, we might see kind of an interesting pressure for organizations to go towards the frontier. Like-
- AMAmjad Masad
Oh, interesting
- AAAlex Atallah
... you know, all, uh, to have no, you know, basically no risk or, or significantly less, um, until-
- AMAmjad Masad
Would they, would they be willing to pay 10X to, to get that, that much?
- AAAlex Atallah
I mean, it probably depends on the, like, types of tasks they're trying to do.
- AMAmjad Masad
Mm.
- AAAlex Atallah
Like, you know, some just have way lower risk [chuckles] than others. Um, writing code, uh, that... Or doing, like, security research is the highest risk type of task today, and so you'd probably spend 10X [chuckles] to get a fully aligned model that can also find all the bugs, or fully, like, anti-deceptive model that can also find all the bugs. One of the coolest things about decision models is that you fully control the structured output and, and generally with structured output models in general, like, the, the room for m- misbehavior is so much lower. You just have, like, defined tasks, and only machines are, like, dealing with the outputs, and it's not writing code, um, that it can execute. The, those tasks feel like probably underrepresented in the ones that people talk about and in the things that enterprises are dealing with. Um, so I, I expect, like, enterprises to get a lot more interested in them.
- AMAmjad Masad
Yeah. I f- I feel like w- we're gonna slowly realize how good we've had, we've had it with, like, deterministic code. We're like, "Oh my God, remember the days when, when computers did exactly [chuckles] what, what we told them to do?" And I think, like, things like Jav, I think, hint at, uh, like more, more of a need for, you know, not only specialized models, but models whose output domain, uh, is, is more controllable. And maybe you could do, maybe you could do a lot more than, than we thought you'd, you'd need, um, you know, by using, like, a bunch of specialized models, specialized output models.
- AAAlex Atallah
Have you guys done any workloads internally with it?
- AMAmjad Masad
You know, I've been training a lot of small models. Uh, I mean, I said this glib thing when it first came out because I, I, I was like sort of... I gave this Hacker News comment-like comment, which I felt disgusted with myself afterwards. But I, I've been taking a lot of, like, Qwen 8B and, like, um, and, like, asking... honestly, like, asking, you know, Fable and Opus and Astra to train a model. For example, I trained a cost estimator model, uh, internally so that when you put it in a prompt in Replit, we know exactly how much it, it will cost, and it basically emits a, you know, probability distribution over multiple buckets. Like, if it is between $5 and $10, bucket A. Bucket B, between, you know, $10 and $20. And like, so I'm used to the, training these classifiers by, you know, g- giving it different enums essentially and, and looking at the-
- AAAlex Atallah
Yeah
- AMAmjad Masad
... log props per enum. I've been doing it for, for a couple of years. I trained a chess bot to play by just doing that. So I, I'm, I'm already sort of pilled on, like, sort of decision models and specialized models, so it wasn't that big moment for me. But I understand that, like, a found- like, a true foundation model that's fully promptable, uh, is like, is like amazing user exper- amazing developer experience, and you can, like, do a bunch of stuff with it without training a model from scratch. But if you have a data, if you work at a place where you, you have the data, we have so much data at Replit, like, I ended up training a lot of specialized classification models pretty easily.
- AAAlex Atallah
Yeah. Like, I definitely... It, it, it also feels like less model debt. Like, something that I, I still hear from companies is that they're, like, worried about fine-tuning models for, like, unstructured outputs 'cause you're just, like, always b- you gotta redo it again in, like, two months, and everybody just feels the model, the weight of the model debt. But like a- ... very bespoke classifier that's trained with, like, proprietary data, you just-- I, I feel like people won't think it's behind constantly.
- AMAmjad Masad
Mm-hmm.
- AAAlex Atallah
And it might just, like, last longer.
- AMAmjad Masad
Yeah.
- AAAlex Atallah
Um, it's just, uh, you, you, you, you don't have to worry about its ability to speak a new language or write Rust or, you know, [chuckles] do anything but, like, the, the LLMs are being, like, evaluated on. You know the use cases, so you can, like, build it more. And I, I, I... It seems like an easy thing for enterprises to build themselves and, uh, and it, and actually, like, not regret.
- AMAmjad Masad
Yes. Speaking of Rust, uh, actually, like, a good analogy is when the world got super excited about dynan- dynamic languages. Um, like, if you think back to the '90s, everyone was writing in Java and C++, things like that. And then, like, Python, JavaScript, Ruby just, like, took over the internet, and everyone was like, "Ah, this is how you build startups really quickly. This is... And y- you built Stripe, a financial organization, on Ruby." It's like, how crazy is that? And th- and we built Facebook, you know, using, using PHP. And then everyone was like, "Oh, shit. Like, we're running into all these really bad bugs. It's, like, freaking slow. Uh, so, like, let's go in and add types. Okay, let's add a JIT compiler. Let's..." And you end up sort of reinventing everything. And then Rust came out, and I was like, "Okay, I guess we can use Rust for [chuckles] for a lot of things we would otherwise be using JavaScript and Python." And my prediction is that the, the same cycle will happen here, where we're using these AGI-like models for all these different use cases, and then everyone's gonna wake up and be like, "Oh my God, this is, like, so wasteful, so risky for no reason." Um, and, and there's-- It, it's gonna be so much easier. We're actually adding that capability on Replit, but I think it's gonna be everywhere. It's, it's gonna be so much easier to, like, to go on a site, like, upload a CSV file, and get a special model [chuckles] that does one thing. Uh, and, um, you know, that, that goes back to your thesis, uh, about, about OpenRouter, this, like, neuro- neurodiversity, which I, like, really fundamentally believe in a lot more. And Erik and I had, like, discussions a lot about, like, AGI and whether we were truly on a path to AGI or w- whether it's even desirable to, to, to get there. And I think the future is, is a lot more diversity.
- 45:12 – 48:04
Fusion models at half the cost
- AAAlex Atallah
Outside of code review, which was, like, I think the first time I saw people get really serious about using, like, different model families to double-check the results of their main model, um, the, uh... these, like, fusion models, like I, I, I... Like, the research has been getting, has been, like, kind of slow for years on, like, doing mixture of, of models and composite models. But f- like, things have been speeding up from my, like, view of the research. And, and I mean, now we see, like, a bunch of AI agent labs. Um, like, we launched a fusion, uh, tool, a, a fusion model, and Technician launched one, and, like, th- they do reduce cost, um, follow... a- and, and allow, like, you know, a wider breadth of ideas to be searched. Like, our initial launch was focused on deep research. The thinking is that, like, if all these model apps are training on different sources of data, like, why not pull from all of them? And, uh, and this, like, resulted in basically fable-level quality at 2X lower cost. Um-
- AMAmjad Masad
We, we just published results actually just, just today about, about that. We, we, we showed like a deep suite, like Replit agent through, through combination different things, including the harness, but, but also you might think of it as a fusion type thing, um, uh, where it's like, um, you know, frontier level at like 40% to 50% of the cost.
- AAAlex Atallah
At the-- Uh, which models does it use?
- AMAmjad Masad
I think it's, it changes over time, but one thing that has been interesting is I think OpenAI added this feature that allows you to save, um, the computation, uh, like, uh, across different model families, so you can... Uh, uh, like also across different, um, effort levels.
- AAAlex Atallah
Mm-hmm.
- AMAmjad Masad
So, uh, l- like you wouldn't do, uh, you wouldn't miss the cache if you change the effort. Don't quote me of this, I think it's also across different models, uh, which is hard to fathom how. I might be wrong though, [chuckles] but I, I need to double-check that. But I think s- you know, staying within the OpenAI family has, like, added a lot of efficiencies. But in the past, we've done it with other, with other, with other models as well. 'Cause cache is, is like one of the bi- like being cache aware is like one of the biggest things when you're designing fusion models, routers, um, uh, sort of-
- AAAlex Atallah
Yeah
- AMAmjad Masad
... escalation models.
- AAAlex Atallah
A- Alex, Amjad, it's been a great conversation.
- AMAmjad Masad
Yeah.
- AAAlex Atallah
Thank you.
- AMAmjad Masad
Podcast.
Episode duration: 48:19
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode ekK8urKHPMQ
