Y CombinatorOpen Models Change The Economics of AI
EVERY SPOKEN WORD
60 min read · 12,304 words- 0:00 – 1:16
Intro
- JMJeffrey Morgan
Cost is by far the largest pain point that open models can jump in and solve. But, you know, every business has a vision of getting better control over AI and customizing it for their business, and that's really their North Star. You know, cost is something they can solve in the short term, but it then enables them to then go and, and customize these models for their unique use case.
- SPSpeaker
Early 2024, there was lots of interest in fine-tuning your own custom models, then it sort of went away and all of that will just be wasted effort. It'll get stomped by the next model release. Seems like it's coming back now. You have a front seat to all of it. Do you think we're going through, like, another cycle, or is it here to stay this time?
- GTGarry Tan
[upbeat music] Welcome back to another episode of The Lightcone. Today, we're talking to Jeffrey Morgan, co-founder and CEO of Ollama, the easiest way to run open source AI models locally and in the cloud. Ollama is used by 9 million developers, has 178,000 GitHub stars, and is used by 85% of the Fortune 500, which means Jeff knows a lot about the state of the art of AI, what models score highest on benchmarks, and what developers actually download and keep using. Jeff, welcome to The Lightcone.
- 1:16 – 2:13
Who's Actually Using Open Models?
- JMJeffrey Morgan
Thank you for having me.
- GTGarry Tan
We're down to hear, like, what is the state of AI? What are you seeing out there?
- JMJeffrey Morgan
Well, I think the biggest thing we're seeing is a shift to open models, especially in enterprise, and that's from a mix of US and, and Chinese origin models, and it's predominantly driven by coding agents and also AI assistants, more co-work use cases like OpenClaw and Hermes.
- SPSpeaker
And because you sit in the token flow of, like, so many tokens, you have really good data on what models people are actually using and how it's changing. What are the trends that you're seeing?
- JMJeffrey Morgan
Yeah, you know, Ollama started as a way to run open models on your MacBook or other hardware, NVIDIA, AMD, Intel, and, um, earlier this year we launched Ollamas Cloud. And what we're seeing there is that it's predominantly Chinese models right now of Chinese origin, um, but they're being r- accessed by businesses all over, all over the world, um, especially US and, and Germany is actually a, a big source of where open model tokens are being
- 2:13 – 3:31
Is It All About Cost?
- JMJeffrey Morgan
accessed.
- SPSpeaker
Uh, is it all about cost? Is it our enterprises coming 'cause they just wanna get the cost down, or is there anything more to it?
- JMJeffrey Morgan
Cost is by far the largest pain point that open models can jump in and solve. But, you know, every business has a vision of getting better control over AI and customizing it for their business, and that's really their North Star. You know, cost is something they can solve in the short term, but it then enables them to then go and, and customize these models for their unique use case.
- JFJared Friedman
Is there a particular, uh, large enterprise that you can name that has done this?
- JMJeffrey Morgan
I think there was a great article in The Information yesterday, uh, from AT&T, and it ends up they've already shifted 40% of their token consumption to open models, and that's right now predominantly through US, uh, and, and Europe models, but they're also evaluating the Chinese models.
- JFJared Friedman
What kind of workflows do they run?
- JMJeffrey Morgan
Predominantly coding agents. I think what we've seen just from the extreme growth and, you know, per developer or per user token usage has predominantly been from coding agents. And then earlier in March and April, we saw OpenClaw take off and subsequently the Hermes project, Hermes Agent project take off, which has then opened up that ability to automate a huge chunk of work over a long span of time to non-developers too, whether it's like finance or support or marketing or
- 3:31 – 5:31
The Token Usage Explosion
- JMJeffrey Morgan
sales.
- JFJared Friedman
You had this actually very cool graph on the takeoff exponential for, for OpenClaw.
- JMJeffrey Morgan
Yeah. So earlier this year... You know, this is a graph of token usage by developer on Ollama's cloud. You know, the average amount of tokens they use per week. And we kind of start-
- SPSpeaker
Just to be clear, this is per developer. So, like, it looks, this graph looks like it should be an aggregate of Ollama's growth, but this is actually the per user growth.
- JMJeffrey Morgan
Exactly.
- SPSpeaker
Right [laughs] .
- JMJeffrey Morgan
This is on an individual user ba- user basis, how many tokens are they using a week? And so there's kind of like two big inflection points. One is that initial run-up at the start of the year, which was driven by coding agents. So we saw Kimi, the GLM models, Minimax launch. Finally, we had open models that could power coding agents. And then in April, we saw this incredible [laughs] growth from OpenClaw, really, which was then not just developers, but the rest of the world could take a hard problem, give it to an open model, and let it go complete the task, which obviously consumes a ton of tokens as it's figuring out what tools to use, what data to go fetch. We went from a context window of 128k to a million-plus with open models, and so all that enabled this explosive growth.
- JFJared Friedman
So it went from, uh, roughly 5X, so under somewhere around, I don't know, 15 million tokens before, before all these co-work type of use cases.
- JMJeffrey Morgan
I think that's about right for that OpenClaw jump we saw in April. Um, obviously in aggregate it's, you know, in the 10 to 20x if not more. As a whole through Ollama's cloud, we saw 150x since the start of the year.
- SPSpeaker
Wow.
- JMJeffrey Morgan
And so it's just... The, the big, you know, interesting thing here is this huge surge in demand for open models, right? Whereas open models, I'd say in 2024, 2025 from the large models being served were mostly being served as custom models. So you'd take an off-the-shelf model like DeepSeek or Kimi and you'd fine-tune it for your use case, um, like for example, you know, Cursor had famously done, and from there, you know, you could serve that at scale. But seeing out of the box open models being served, that really only took off at
- 5:31 – 7:27
Fine Tuning: Hype Cycle or Here to Stay?
- JMJeffrey Morgan
the start of this year.
- SPSpeaker
Even since we started this podcast, these things sort of come in cycles maybe. Like, I feel like early 2024 there was lots of interest in fine-tuning your own custom models, then it sort of went away and it's like all of that will just be wasted effort. It'll get stomped by the next model release. Seems like it's coming back now. You have a front seat to all of it. Like, what's... Do you think we're going through like another cycle or is it here to stay this time?
- JMJeffrey Morgan
I think the release of these models, the cadence is only speeding up, which makes it ever more harder to- You know, stay on top of that and have a, a post-training-
- HTHarj Taggar
You mean the, the, the, the latest closed frontier models are releasing faster than ever-
- JMJeffrey Morgan
I think on the open source side-
- HTHarj Taggar
... ever does
- JMJeffrey Morgan
... it's getting faster and faster. You know, this summer we've already seen three iterations of the DeepSeek Flash model-
- HTHarj Taggar
Yeah, I see
- JMJeffrey Morgan
... as an example, um, what used to be pot- more of a six-month cycle. And so-
- HTHarj Taggar
Yeah, so the gap is, like, closing essentially
- JMJeffrey Morgan
... it's closing, and I think that makes it even harder to custom train models. On the flip side, I think the tooling's getting better, and so it allows teams that wanna fine-tune their models to stay on top of it.
- SPSpeaker
I mean, we're also entering this moment where AI safety is becoming more and more of an issue at, uh, the frontier labs. So the frontier, uh, may well slow down to figure out its alignment and containment issues. [chuckles] And then meanwhile, the open source models and open weight models are continuing to grow and get better.
- JMJeffrey Morgan
Yeah, and you know, we saw the announcement and the release of the GLM 53 model and its capabilities from a cybersecurity standpoint, uh, you know, being extremely impressive. It creates a big opportunity for whether it's startups or existing businesses in the security and governance space to really jump in and help because I think if you look at, you know, that AT&T article I was speaking about, it largely... The, the blocker to adopting open models is largely around security and safety.
- SPSpeaker
Hmm.
- JMJeffrey Morgan
Um, but from our experience talking to customers, whether it's in Europe, whether it's here in the US, if you can solve the safety problems, by and large, adopting the Chinese origin model labs is completely on the table, and it's really exciting for these
- 7:27 – 8:26
Where Open Models Beat Claude
- JMJeffrey Morgan
businesses.
- SPSpeaker
There's sort of this interesting moment right now where a Hugging Face had to use open weight models to actually even detect the hack from the frontier. [chuckles]
- JMJeffrey Morgan
A common question we get is like, "Well, where, where can I use open models that are leaps and bounds of an advantage over, uh, using a frontier m- closed model?" And one of the key use cases is security testing and making sure that, you know, your software's secure.
- HTHarj Taggar
Yeah, 'cause, like, if you try to get Claude to, like, pen test your product, it will just refuse to do that.
- JMJeffrey Morgan
Correct, yeah, by and large. [laughs]
- SPSpeaker
Whereas there are, uh, literally obliterated, uh, security researcher-
- JMJeffrey Morgan
[laughs]
- SPSpeaker
... models that you can find on Hugging Face that allow you to do it.
- JMJeffrey Morgan
That is true. There are ones that are, are, you know, custom trained to be, you know, even more, you know, liberal to, to go and attack these problems. But even the out-of-the-box models, um, they, they do come with safety training, but they're a little better understanding if you're doing this for, you know, a, a, a good, you know, use case versus, versus one that's more of a negative or, you know, malicious
- 8:26 – 11:31
Launching Models at Scale
- JMJeffrey Morgan
use case.
- HTHarj Taggar
A cool thing about Ollama is that because you guys are such a key distribution channel for these models, my understanding is that typically the model developers are contacting you before the general release to, like, coordinate launches and stuff like that, and you often get sort of like previews of what's about to, about to happen, and you get these incredible growth spikes when a new model drops. Like, I wonder if you could, like, tell us a bit about, like, what it's like to operate this thing at scale.
- JMJeffrey Morgan
Yeah, absolutely, and like I said before, the models are coming out faster and faster-
- HTHarj Taggar
Yeah
- JMJeffrey Morgan
... and faster, and so we've developed a playbook to what does a su- successful day zero model launch look like? And there are a lot of things to get right. There's making sure that it's supported in your favorite inference engine, which is generally sometimes a multi-week process to make sure it's fast, to make sure it's accurate-
- SPSpeaker
Mm-hmm
- JMJeffrey Morgan
... to make sure it's up to spec with the reference. There's also finding the right use cases and harnesses for developers and users to, to make use of this model. Generally, these models have new capabilities. This morning, t- you know, there's announced, uh, DeepSeek launched their first multimodal model, um, from a large model LLM standpoint. They had previously had some smaller OCR models which unlocks a whole bunch of use cases, but what they also changed was the DeepSeek harness which launched recently so that it could support this capability. So step two is then to find the harnesses, make sure they're prepared to actually run this model and to do that effectively. But, you know, every model's different. They all have different challenges and architecture changes and tool calling mechanics, and getting all this right is really hard. I think the key thing to do at the end of the day is to run benchmarks against the final product ahead of release and make sure that, you know, it's, it's running as, as it, as the research team at the lab specified.
- SPSpeaker
So I think one very important part that you play in the whole ecosystem is you kind of create a very legible standard to be able to make sure you have the best way to use the hardness for each new model, each new tool call, and all of it is consistent across all the different models, which is pretty hard to do.
- JMJeffrey Morgan
Yeah, I think there are three things really that, you know, we try to package together. One is harness, for example. You know, whether it's an off-the-shelf harness or an SDK to help use the model, some existing harness that is designed for this. And, and what's great now, there's so many great open source harnesses. Um, the Codex Harness is open source. Uh, OpenCode's a great one and one of the most popular harnesses from Ollama users. Step one's getting that right, but then you gotta package it with the model, making sure that it's, you know, available, it's reliable, if it's in the cloud, that there's enough capacity for it because day zero tends to be the largest growth day, obviously. And then i- importantly, there's the hardware and the providers, and that's actually where there's a lot of collaboration to be had, whether it's like an inference provider optimizing the model or, you know, with some of our partners, whether it's NVIDIA or, or, you know, for example, working with the Apple Silicon stack, making sure that it actually runs the model really fast. 'Cause if the model's capable but it's really slow, that's not a great experience. So getting those three things packaged into a box, and generally you get the model, you're lucky if there's a, you know, uh, the, the ability to access the model a few weeks in advance. A lot of this k- stuff comes together in the last 24 hours before the model gets released, and so it's generally a fire
- 11:31 – 13:58
Olama as the OS Layer
- JMJeffrey Morgan
drill.
- SPSpeaker
So you kind of almost, uh, like an operating system in the old world where you needed to really integrate very tightly with all the drivers, all the hardware, and then at the application level to make sure that all the apps were f- really tuned up well, and you are the glue for all of it, right?
- JMJeffrey Morgan
I do think the OS, which is generally a cliche analogy-
- SPSpeaker
Yeah
- JMJeffrey Morgan
... to use, is a good one because you've got the drivers for the hardware and the providers and the inference layer, but you also have the application runtime and making sure that the harness works. Gluing that together, it's a very combinatorically large problem to solve, and so doing that well is really difficult. And so, but, you know, over time you develop pieces that, you know- Allow you to quickly develop that, test it, release it, um, and kind of have this common runtime that can match any harness to any model, and that's kind of the role we're really playing for developers.
- JFJared Friedman
You guys are very hardcore engineers, and you fine-tune things all the way to from the origins with Apple Silicon all the way now to DGX. How do you build such a deep technical bench with that?
- JMJeffrey Morgan
I think a lot of our team, you know, we aren't AI researchers by background. We're from VMware and Docker, um, and from, you know, other networking companies. And so the classic compute problems are kind of reinventing themselves in the inference land, whether... And, and so largely, you know, that's where we like to focus our time. But I think at the end of the day, you know, it all comes down to the developer experience. What is it like when the developer makes a call to the API and gets tokens back? What happens? And it- there's more and more happening in that layer right now, and getting that right needs all the layers of the stack to work well together. And I think, you know, one of the challenges with open models has been that hasn't been happening at the rate of what a frontier model lab puts out, where they have, you know, the classic five-layer cake, right, that Jensen mentioned, which is like the apps. You have, you know, the model. You have the infrastructure and inference. You've got the chips, and you've got the energy, and they've got all that ready to go for developers on day zero, and that's really the thing that we're trying to reproduce for open models. Of course, we're not gonna do every layer of the stack, but we can help orchestrate that.
- JFJared Friedman
You layer the lay- layers. [laughs]
- JMJeffrey Morgan
Yeah, and maybe in open models there's more than five layers. Like, that model layer actually has a lot to it, right? There's the model weights, but there's also a lot of the orchestration components that each model's uniquely good at. There's this developer API layer. There are so many opportunities to build between the model and the application layer that are kind of hidden in today's, you know, five-layer cake
- 13:58 – 17:28
Hidden Layers Between Model and App
- JMJeffrey Morgan
stack.
- GTGarry Tan
That's interesting. Do you, do you think any of those, like, hidden layers might get unbundled and become their own companies or providers?
- JMJeffrey Morgan
Absolutely. There was a really good talk from some of the Anthropic team, the platform team, and they talked about three big things. One was knowledge, how do you connect your company's data and context to the model? One is coordination, so as, as you know, when you make a request to in your Claude app or Codex, it goes off and spins out a bunch of sub-agents, some of those in the cloud, some locally. There's a coordination problem. Um, and then lastly, there's an execution problem, which, you know, we talk a lot about as sandboxes, but these agents, more and more of them are moving to the cloud. There's a huge compute problem to be solved there. And if you look at what's cl- happened classically in op- in, in the cloud business, open source has meant that there are best-of-breed companies for each of those problems. Whereas you might have had kind of the... If you look back to the original generation cloud products, you have, like, the Herokus of the world. You have Google App Engine, where all those things were bundled together, but what developers ended up preferring is best-of-breed products for each one.
- GTGarry Tan
It's the Navalism, which is, uh, all things are just bundling or unbundling-
- JMJeffrey Morgan
Exactly
- GTGarry Tan
... like the frontier model.
- JMJeffrey Morgan
Exactly.
- GTGarry Tan
Frontier model labs want you to be totally in their walled garden of managed agents and their context, their memory layer, and then meanwhile, Little Tech and all the founders out there and all the open source developers, uh, don't wanna be caged in, so we're gonna make all this other stuff. And it'll, it'll be an interesting moment to figure out, like, what ends up winning. I mean, it'll probably be some mix of both.
- JMJeffrey Morgan
I think so. And you know, you've got this abundance of open model tokens that's being created. There are just dozens of open model providers that are able to serve these tokens. The new scarcity, the problems now are what's above the tokens, right? You know, how do you orchestrate an agent from, you know, A to B? These are problems that are of tons of new, you know, systems and, and engineering problems that are just really hard to solve for an individual dev. Like, there's no way they're gonna build all those layers.
- GTGarry Tan
Well, the interesting thing now is, um, because the coding agents themselves are getting a lot better, and ostensibly this is the worst the models will ever be, the classic reason why there was a moat here was it was just too hard to have really well-maintained software that was properly tested that actually satisfied user need, and what if that goes away? [laughs] Like, we're r- literally at this moment where actually maybe the a- you know, you'll just have a cron, and it runs a markdown file in some s- typescript, and it'll just... You know, there's no lock-in anymore, right? Like, you could be using OpenAI's memory system one day, and then actually, you know, you could have an agent be s- uh, constantly syncing that against your own memory system, and actually it just works. It's fine. Like, it works over MCP. Like, there's all, you know, there's plenty of, like, runtime testing, and then there's no lock-in.
- JMJeffrey Morgan
Yeah, I think a lot of the, you know, what we know of as harnesses today, a lot of those pieces will go down into the model. But, you know, as you kind of push a lot of the core loop of the model and hooks down to the model itself, there are these pieces that kind of come out. Like, memory's a great one, and the, the general guidance we're using is if it's a stateful problem, like there's storage involved, that's something that, you know, in the end can't go into the model because the model's trained, and it's, you know, as we know, training runs are now happening on, like, a monthly basis, but it's still not up to date with the latest data. So generally storing data is, like, a huge problem space that I don't think will ever make its way down into the model layer.
- 17:28 – 20:41
Open vs. Closed: The Steady State
- SPSpeaker
Like managing credentials and that kind of stuff is never gonna be-
- JMJeffrey Morgan
Security credentials, safety. Um, open models don't have all the safety tooling that closed model providers give you out of the box, but that's super important, especially for businesses to adopt them.
- SPSpeaker
I'm curious where you see sort of the, the end or future state for enterprise on the balance between sort of frontier closed models and open source model, um, s- especially on spend. Like, my... It feels like, you know, initially it was just The everyone's just like allocating all of their budget to Anthropic or, um, OpenAI. Uh, my sense now is, yes, open A- open source is clearly growing, but, like, so are, so is, like, Anthropic's spend. So, like, the things seem to be growing together. Like, does that continue, or do you think there's sort of like a steady state where it's like, I don't know, like, half the budget's gonna be on the closed source frontier model and half's gonna be on open source or, or, or something different?
- JMJeffrey Morgan
The super majority of tokens, and this is our take, it will be open models within a business. Call it 80, 90%.
- SPSpeaker
Mm.
- JMJeffrey Morgan
That doesn't mean 80, 90% of the, the budget will go to open models.
- SPSpeaker
Mm.
- JMJeffrey Morgan
In fact, I think what the open model community's doing incredibly well together is lowering the cost to make it more accessible. And so maybe you'll only pay 10 to 20%, uh, of, of the cost towards open models, but your token usage will be high.
- SPSpeaker
But most of your tokens will be going through the open models.
- JMJeffrey Morgan
Right, which will enable a whole bunch of use cases on top because you have this abundance of tokens. You're not thinking about taking away token access from your team. You're giving more and more access. I think for the hardest tasks, that's reserved for these frontier labs-
- SPSpeaker
Yeah
- JMJeffrey Morgan
... where a lot of the best researchers are. And then from there, there's a whole bunch of problems in the middle, right-
- SPSpeaker
Mm
- JMJeffrey Morgan
... where maybe it's a combination of open and closed models working together. But I think the steady state is that most of the software's op- uh, most of the models are open.
- SPSpeaker
It's kind of an interesting idea 'cause as the models get more powerful, ideally you just wanna delegate to, like, your smartest model to figure out when to go to, like, an open model. But the labs who own the models presumably don't want that.
- JMJeffrey Morgan
And I think, look, I think, uh, that all the labs are aligned in many ways to one thing, which is how do you serve the customer? And I think it will be up to the customer to decide if I have a router where some of the scheduling and harder, you know, orchestration happens through a frontier model, so be it. But a lot of the kind of line item work can happen through open models, and the collaboration of the two together. And I think we've seen a ton of projects, whether it's from Sakana AI or Open Router, that have combined the two and have seen really good results.
- HTHarj Taggar
This is not too dissimilar from a human organization, right? Like, you have, you know, like a law firm. There's, like, a partner, and then there's, like, a bunch of associates, and, like, the partner farms out the work to the-
- JMJeffrey Morgan
Yeah
- HTHarj Taggar
... the associates, right? [laughs] It's like the same. [laughs]
- JMJeffrey Morgan
We saw the same thing with cloud computing, where it was really a blend of proprietary software, some of them provided by the, the cloud providers themselves. For example, you know, uh, AWS had DynamoDB, which was kind of their proprietary scale-out database. But then a lot of customers use that in conjunction with PostgreSQL DB. And in the end, you know, what we see is customers will, will use a combination of the two.
- JFJared Friedman
I think it's a very common pattern. I mean, this is also same design for why the Apple silicon is actually more superior, the special, spec- special s- accelerators for different kinds of workloads for, let's say, image processing as opposed to audio. That's, it's been... Or even go way back in, in, in the PC era, you had, like, your standalone audio card, right? Graphics
- 20:41 – 23:12
Local vs. Cloud Models
- JFJared Friedman
card and all that.
- HTHarj Taggar
And speaking of Apple silicon, should we talk about local models? 'Cause you're, you're in a bit of a unique position in- because you have large businesses both in cloud-hosted models and locally hosted models that'll run on your laptop. What are you seeing in those two worlds, and what, what do you think is gonna happen?
- JMJeffrey Morgan
I think it's incredibly exciting because it's similar to the closed versus open model question. It'll be a mix in our mind, and, and that's what we hear from customers as well, where for easier tasks you could run them locally and with lower latency and, of course, lower cost when it comes to the per token cost. Ultimately you're buying hardware up front, and you'll use that in conjunction with these cloud models. What's exciting about this next generation of hardware, which we've had for a few years now, is just how good they are at running the 20 billion parameter to 40 billion parameter range of models, sometimes up to 120 billion parameters.
- HTHarj Taggar
Yeah, Qwen 3.8 38B is now as good as Opus 4.6 for coding. Is that right?
- JMJeffrey Morgan
That's what the benchmarks show.
- HTHarj Taggar
Yeah, that's wild.
- JMJeffrey Morgan
It's incredibly exciting-
- HTHarj Taggar
Yeah
- JMJeffrey Morgan
... 'cause you can run that on not the l- the s- the lowest memory MacBook, but the second lowest memory MacBook you can buy-
- HTHarj Taggar
It's insane
- JMJeffrey Morgan
... from the store. [laughs]
- HTHarj Taggar
That's insane.
- JMJeffrey Morgan
So it's incredible.
- HTHarj Taggar
And are you seeing your users do that? Like w- what are you seeing people use the local models versus the cloud-hosted models for in practice?
- JMJeffrey Morgan
From the side of which models they're running, we're seeing a really solid mix of US and Chinese-trained models being used for local, and we have the incredible models from, you know, the original Llama models of course, but also the Gemma models from DeepMind. These are great choices for local. But when it comes down to use cases, coding agents by and large are most effective with the large cloud models. You're solving really hard problems. You're writing code, tests. It's really difficult. Uh, versus some of the document processing workflow use cases that run extremely well locally because they don't have as difficult of a, of a task in the end-to-end, you know, problem you're trying to solve. And so that's where we see this hybrid execution model where some of the easier, more straightforward tasks run locally, and then you have a router that can help decide, hey, we need to go to a large cloud model for this. And I think what that means for customers is that you're, you're really dropping the costs even further when you're going from open models, not just because they're cheaper to run in the cloud, but because now you can run them effectively for free on the hardware you're buying for your business anyways. And what we're seeing ultimately from the cloud coding agent models, it is pre- predominantly Chinese models being consumed today. And for the local models, it's a really strong blend of, uh, US, Europe, and, and Chinese origin models.
- HTHarj Taggar
Yeah,
- 23:12 – 24:00
Why Chinese Models Dominate Cloud
- HTHarj Taggar
this, these two graphs are pretty stunning in comparison. Like basically for local models, the US and Chinese models are neck and neck. We're like tied. And for cloud-hosted models, like the US is like recoloring the X-axis. It's like 100% Chinese models. Basically we need more US labs to make large models. [laughs] Is, is, is, is, is that what this graph is showing?
- JMJeffrey Morgan
Effectively. And you know, with the launch of the Nemotron Ultra model, we're seeing kind of the first wave of that, and it, it's really exciting.
- HTHarj Taggar
NVIDIA as a company is so interesting because, um, their moat is not like trying to start new software businesses or sell, you know, tokens. They seem to be quite interested in just Releasing a lot of open source and helping the ecosystem, and then the fact that they do that then helps them stay ahead of the game on the hardware
- 24:00 – 26:03
NVIDIA's Open Source Play
- HTHarj Taggar
side [laughs]
- JMJeffrey Morgan
I think so. And, uh, you know, ultimately NVIDIA, what's so incredible is they're helping power an ecosystem around open models, whether that's the hardware, the models. You know, we've seen the, uh, new DGX station, uh, computers that they're working on-
- HTHarj Taggar
I want one
- JMJeffrey Morgan
... which have a GB300 on your desk-
- HTHarj Taggar
How do we get on that list?
- JFJared Friedman
On your desk
- JMJeffrey Morgan
... that isn't deafening-
- HTHarj Taggar
Yeah
- JMJeffrey Morgan
... loud.
- JFJared Friedman
[laughs]
- HTHarj Taggar
Yeah. Do you know the price point on that thing yet?
- JMJeffrey Morgan
I don't know it off the bat. [laughs]
- HTHarj Taggar
Yeah. It's gotta be like-
- JFJared Friedman
It doesn't matter, Garry's buying it
- JMJeffrey Morgan
Yeah.
- HTHarj Taggar
I know, right? [laughs] Well, I looked it up. It's, I mean, you can probably run a frontier model for, like, I mean, very slowly for, like, two, two, $300,000. Is that, is that right?
- JMJeffrey Morgan
I think it's even more competitive than that.
- HTHarj Taggar
Yeah.
- JMJeffrey Morgan
And you can run more than a frontier model at high speeds at a, a price point that isn't very far off what you can buy from a classic workstation computer.
- HTHarj Taggar
Oh, no way.
- JMJeffrey Morgan
So if you think about, you know, quite a few of the customers we talk to, some of them are banks, for example, or industrial businesses.
- HTHarj Taggar
[laughs]
- JMJeffrey Morgan
They already have these NVIDIA workstation GPUs in every single engineer's desk, some of them tens of thousands of them. And so this is-
- HTHarj Taggar
Oh, I want one
- JMJeffrey Morgan
... entirely-
- JFJared Friedman
They've been selling them for, uh-
- HTHarj Taggar
How do we get them?
- JFJared Friedman
... for all the CAD work for a while
- JMJeffrey Morgan
Well, this is for, this is for the kind of original RTX A-
- HTHarj Taggar
Yeah, yeah, yeah, of course
- 26:03 – 27:15
The Return to Local
- JMJeffrey Morgan
at running them.
- HTHarj Taggar
So you'd say, like, this is, this is the platform to get. Like, you could make Stu- Apple Studios work, but, like, if you want something that just can work, get DGX Spark.
- JMJeffrey Morgan
From our testing, both are very competitive.
- HTHarj Taggar
Got it.
- JMJeffrey Morgan
So I think it, a lot of it will come down to-
- HTHarj Taggar
Whatever you can get, really
- JMJeffrey Morgan
... what you can get-
- HTHarj Taggar
Yeah
- JMJeffrey Morgan
... [laughs] and then also, um, the tech stack you're looking for. I think there's an incredibly, uh, mature tech stack through the MLX project with Apple, where they've done some amazing work to run LLMs on the Mac Studio, but also the smaller Macs, and of course the DGX Spark stack's just incredible. We're super excited, uh, as partners with NVIDIA for that.
- JFJared Friedman
And it's gonna be a cool renaissance for personal desktops.
- JMJeffrey Morgan
I think so. And you know, it's, it's funny with Ollama's journey, we started local. Clearly the coding agent demand is in the cloud, but that's gonna come back locally in our minds because the hardware will catch up. When you have a GB300 on your desk and you want the fastest coding loop that's as fast as running your tests or as fast as making code editor change, we all remember the GitHub Copilot experience of having the auto-complete come up in a few millisecond- at 100 milliseconds. That experience will make its way back to the desk, which is, which has been a journey of starting local, going to the cloud, and then we think that'll come back local, and you'll end up using the two
- 27:15 – 29:16
Getting GPUs Is Hard
- JMJeffrey Morgan
together.
- HTHarj Taggar
Speaking of, like, what you can get, in order to run Ollama Cloud, you need, like, a shit ton of GPUs. What are you seeing in the GPU market?
- JMJeffrey Morgan
I think what we're seeing is ultimately the prices are changing very quickly, and the supply and demand volatility is, is very high there. And so, you know, I think ultimately for if you're a startup, getting access to some of the B200, B300 GPUs you need to run these latest models is very hard. Thankfully, there's a great set of inference providers building on top of that, and so we're seeing this extreme demand unlike, you know... And, and when we think it's-
- HTHarj Taggar
Are you able to get all the GPUs that you need? Are you constantly like, like growth limited by how many GPUs you can get your hands on? What's the, what's the current state?
- JMJeffrey Morgan
We're lucky in that we've partnered with quite a few providers to work together to pool a bunch of GPUs together, which allows us to stay on top of our, our demand. But that's a lot of work, and it's definitely a lot of spending time thinking through, you know, which model will get run where, how fast should it be, which region is it in, um, what will the latency be for the customer. So there's a lot of hard problems to solve in that stack, and I think what's really exciting about products like OpenRouter, Ollama, the OpenCode project is for an end user developer, they can sign up and get access to this without having to go negotiate prices on a B200, B300, you know, think about their 24-month forecast [laughs] in order to get access to some of these, these, uh, you know, GPUs.
- HTHarj Taggar
YC's next batch is now taking applications. Got a startup in you? Apply at ycombinator.com/apply. It's never too early, and filling out the app will level up your idea. Okay, back to the video. Suppose you were, like, a s- a startup founder, and you were just starting out now, and you're building some AI company, and you haven't raised a lot of money, and so you, like, want to, like, use as many tokens as possible, like, inexpensively. Like, what would your advice be to that person about, like, how they can get, like, huge mileage with, like, a limited budget?
- 29:16 – 31:35
The Flash Model Revolution
- JMJeffrey Morgan
There's this new class of models, like DeepSeek Flash is a great example, and I think there'll be quite a few more, where it's ultra low cost per token. It's also low cost per task, which is a really important metric. And- That class of models w- in my mind will be the first ones that come down to this idea of, like, unlimited tokens. We all remember ChatGPT, you didn't really have to think about how many tokens you were using. You would just use it every day, you had unlimited. Ultimately, I think we return to that, but it's gonna take a lot of work in the model, the architecture, to be custom trained for high volume token usage. And if you think about the start of the year, we really want- open models got to the frontier of intelligence. We bridged the gap. We're maybe, like, less than three months behind between the frontier closed models and the open models. But the next problem to solve is extreme efficiency. Seeing, for example, the GPT Luna model become very, very price effective for customers has been a huge boom. We talked to a ton of customers where that kind of pricing enables widespread adoption within a team. I think we're gonna see that with open models. We already are seeing that with open models. I think the DeepSeek Flash model is leading that charge. If we go back to the model breakdown on Ollama's Cloud, the highest growth area is definitely the DeepSeek model, and this is largely powered by the DeepSeek Flash, uh, adoption. And so this new class of Flash models where they're good enough for 80% of the tasks, they're really fast, and they're ultra-cheap, s- this new class of model that I think will enable some of those use cases.
- SPSpeaker
Yeah, those are gonna be like the workhorse models that do, like, all the grunt work.
- JMJeffrey Morgan
Exactly.
- SPSpeaker
Yeah.
- JMJeffrey Morgan
You won't have to be thinking about how many requests am I making, how many tokens. You'll be much more inclined to consume as much as you can because, you know, it's able to solve the hardest pr- the h- not the hardest, but-
- SPSpeaker
Yeah
- JMJeffrey Morgan
... you know, difficult problems. If you think back to the coordination layer we were talking about earlier too, being able to coordinate these Flash models together to do different tasks can also yield great results that a, a bigger model can. So by having these cheaper models, not only give- are they more accessible, they can run faster, and you can access them at higher volume, but you can start to chain them together and build new problems that are solved by orchestration on top, and that's a really exciting area for new startups, for existing inference providers, um, for some of the larger businesses today that solve workflow problems. Ultimately, being able to chain these models together is gonna be super helpful, and you won't have to think about the underlying costs.
- 31:35 – 33:37
God Model vs. Orchestration
- SPSpeaker
Yeah, I guess, you know, when we first started talking about AGI, even on this podcast, um, there was this sort of debate about, you know, and I think a lot of AI researchers would come out and say, like, "There's just gonna be a giant God model, and it's gonna do everything." But, you know, I think so far, like, it hasn't quite worked out that way. Like, obviously you still ha- you know, if you're, if you have to literally hack the NSA, maybe you need Mythos [laughs] or something.
- SPSpeaker
[laughs]
- SPSpeaker
But, um, for the majority of use cases, like, you're talking about orchestration, and you're talking about, um, like, smaller models. You, you, you know, that the task composition actually probably gives you a bunch of ways to make it more repeatable, it's more trustworthy, like, it actually does work at a cost that is, like, possible. So y- you know, if, if it was gonna be, um, God model versus, like, lots of, you know, smaller special purpose or even just, like, simpler models, um, it's turning out to be the latter so far.
- JMJeffrey Morgan
I think for most customer use cases, there's a, you know, level at which a model becomes good enough, and then they can continue using that level of intelligence. Maybe the model will get faster, it'll have better architecture, it'll be, it'll have new capabilities, but they won't have to reach for the God model. But I do think there are use cases where the most powerful models unlock them, and that'll continue to be a thing, and it'll be really exciting, you know, and sometimes scary as well on, on what they can do. But for the run-of-the-mill use cases where open models really shine, I think that's where, you know, we're, we're hitting a point where, you know, you're not solving necessarily the hardest problems within the business, but they're hard enough where it's now unlocked by open models. Will there be an open model that becomes a God-tier model? Um, I think it's possible, and we're seeing, you know, really exciting developments from, you know, Zhu AI and GLM, where some of the f- tasks, they are frontier. And we all saw with the Kimi model how for web development it became the best model. And that sent this new shockwave across the market, which is it's less about a gap, and it's more about a head-to-head competition-
- SPSpeaker
Hmm
- JMJeffrey Morgan
... which I think makes all of this even much more exciting.
- 33:37 – 36:14
The Geopolitics Question
- SPSpeaker
This is a bit of a sensitive question, but what do you think about this and the geopolitics around it?
- JMJeffrey Morgan
I think, you know, a lot of the geopolitical angles around this, um, start with, you know, where the c- where the model's from. And the more we spend time with customers and users, a lot of it's actually how the model's run, where it's run, how it's run. Is it run in a secure environment? Um, and that starts to matter a lot more. Um, but I do think, you know, look, it, it's super important that a customer in the US can use a, a model trained in the US, and we have two kind of s- classes of, of customers we speak to. One is they don't really care where the model's from, they care about where it's run. But for every one of those, there's, you know, a customer that's saying, "I really care about where the model's from because it's data, the way it..." It's not even just a, a security issue as much as how does the model speak? You know, we all go through mo- uh, and, and-
- SPSpeaker
Mm-hmm
- JMJeffrey Morgan
... communicate, and we all go through, you know, uh, a lot of the models go through phases where they sound more robotic, they sound more friendly, and a lot of that matters too. Um, but I think the highest order bit is obviously making sure that you have a model that end to end you understand where the data's from, which is great from the Nemotron models, that you can go and introspect what made this model. 'Cause if you're putting in a mission critical task, which people are absolutely using open models for mission critical tasks, there's a post online about how Ollama powers the analytics of a power plant to detect surges in Finland to make sure that the lights stay on. That's where these models, the, uh, the model origin really matters.
- SPSpeaker
For, for, like, critical tasks like that, how do you ensure that a Chinese model [laughs] even if it's hosted in the US, isn't basically, like, booby-trapped to, like, cause problems?
- SPSpeaker
Yeah, the Manchurian Candidate problem.
- SPSpeaker
Yeah.
- JMJeffrey Morgan
[laughs]
- SPSpeaker
Exactly.
- SPSpeaker
Have there any, been any, like, known cases of, of the Manchurian Candidate yet?
- JMJeffrey Morgan
I think not that I can think of off the top of my head.
- SPSpeaker
Yeah, I can't, I haven't heard... I feel like I would've heard about it, but-
- JMJeffrey Morgan
You know, what you don't see a lot on some of the press articles is how robust some of the The IT and security teams are at the businesses that we know of, the top, you know, Fi- Fortune 500 businesses. They're really used to this already because open source software, if you think the average application has thousands of dependencies, this is like isn't a new problem, and all it takes is one dependency for there to be a major security issue in the entire application.
- HTHarj Taggar
Yeah.
- JMJeffrey Morgan
So there are-
- HTHarj Taggar
Supply chain poisoning is insane.
- JMJeffrey Morgan
It's a thing.
- HTHarj Taggar
It's, it's fan out.
- JMJeffrey Morgan
It's been a thing for decades, and it's not new in that sense. It's a little more opaque 'cause you can't like dig into the model. It's more-
- HTHarj Taggar
It, it isn't deterministic, right?
- JMJeffrey Morgan
But it's deterministic, and if you screen the model properly with safety checks, by and large, at least what we're hearing is from customers, is that
- 36:14 – 37:12
From Docker to Ollama
- JMJeffrey Morgan
can be solved.
- HTHarj Taggar
Do you wanna talk about the origins of Ollama? You, you know, you guys came up through the Docker ecosystem, and, uh, a lot of people watching, you know, would love to be in the position you're in, where you have this sort of enduring brand moat that looks like it will extend for, you know, really [laughs] till the end of time. No, I mean [laughs] like just it's, it's a very powerful situation to be in. Um, you basically found yourself on top of a giant oil well, right? Um, for those out there wildcatting, you know, can you tell us that story? You, you were actually working with Jared in, uh, 2021.
- JMJeffrey Morgan
Yeah. You know, my, my co-founder and I previously built Docker Desktop while at Docker, so we really got a understanding of like what makes a great developer experience.
- HTHarj Taggar
Mm.
- JMJeffrey Morgan
But I have to say, the first few years of Ollama as a company was really in search for what's the right problem to solve with this muscle we've built of trying to design a great experience for developers.
- 37:12 – 40:36
Applied to YC With the Wrong Idea
- JMJeffrey Morgan
And-
- HTHarj Taggar
So you applied to YC with a very different idea, right?
- JMJeffrey Morgan
For sure.
- HTHarj Taggar
Do you remember what the, like, tagline was when you guys applied to YC in Winter '21?
- JMJeffrey Morgan
I think it wasn't well-defined. I think we realized, let's go back to building a really great desktop experience for containers and Kubernetes.
- HTHarj Taggar
I remember what I wrote down on the, on the application. It was a Kitematic for Kubernetes or a Docker Desktop for Kubernetes. [laughs]
- JMJeffrey Morgan
Yeah, which was effectively Docker Desktop. They had a great [laughs] Kubernetes product.
- HTHarj Taggar
[laughs]
- JMJeffrey Morgan
I think, you know, uh, it's one of the challenges as a second-time founder that, you know, Michael and I have told ourselves we, we tried to over-engineer the idea in many ways, and I think even the two to three years after, like Ollama, we did YC in 2021, and Ollama wasn't launched until July of 2023, after we raised our Series A, after obviously after we had done YC. That journey was one of really in search for a customer problem that could delight a developer. And in some ways it was almost a good thing that we tried different ideas and pivoted until 2023 because that's when Ollama, Ollama came out and started the open model wave.
- HTHarj Taggar
Which is why it's called Ollama.
- JMJeffrey Morgan
Not necessarily.
- HTHarj Taggar
Okay.
- JMJeffrey Morgan
No. [laughs]
- HTHarj Taggar
Oh, really?
- JMJeffrey Morgan
Ollama means generally from our experience, whether you think of LocalLlama, the subreddit, Ollama really just stands for Open Models. You know, as we were looking through the name, it wasn't necessarily from an existing model. It was-
- HTHarj Taggar
It's more LLM and-
- JMJeffrey Morgan
Exactly.
- HTHarj Taggar
It's like the animal plus LLM.
- JMJeffrey Morgan
Yeah, and I think having that character was important. So we're like, "What's a good name for a character, a face you can put to the name?" Because open models are scary.
- HTHarj Taggar
You need a good animal mascot sometimes.
- JMJeffrey Morgan
Docker had one.
- HTHarj Taggar
Yeah, yeah.
- JMJeffrey Morgan
GitHub. Yeah.
- HTHarj Taggar
He still hasn't taken my advice to have llamas come to actual Ollama events.
- JMJeffrey Morgan
Oh my God. [laughs]
- HTHarj Taggar
[laughs]
- SPSpeaker
I'm curious what your Series A pitch was, 'cause you raised from Benchmark, like fantastic investor, but all of this, the feature we're in now hadn't quite taken off in 2023. So what, what was like the pitch and the vision back then?
- JMJeffrey Morgan
Yeah, and we partnered with Benchmark in 2022, so it was DALL-E days, but pre-ChatGPT. When it came down to the pitch, I think we weighed so much on like, "Hey, we're trying to build this great developer experience. We're solving this security problem." And we had known, uh, Peter, a partner at Benchmark, from our previous lives building, building at Docker because he was the Series A investor in Docker. And so a lot of it was weighted on, you know, the people and also why we exist. I think the what, I mean, you know, solving s- SSO for Kubernetes, which is a real problem, wasn't really our passion, and I think we were really lucky to find a partner that could, could see us for what we stood for and what we were trying to do versus the point in time, you know, problem we were solving at that point.
- SPSpeaker
Oh, I see. So you raised the A sort of pre-pivot then actually.
- JMJeffrey Morgan
Correct.
- 40:36 – 42:15
Lost in the Wilderness for Two Years
- HTHarj Taggar
was the experience like?
- JMJeffrey Morgan
It was definitely scary, and, and, and for a few reasons. You know, one is, like, when you're, when you're trying to solve a problem for devs or for a customer, and you're just getting on the phone with them over and over again, and it's not totally clicking, that's, you know... It, it's less about the are we in the headlines or is, is the project taking off, the product we're building. It was just, are we truly actually solving a problem for somebody? And I think being lost in the wilderness, like, what's your North Star? That, customers are generally a great North Star, but not seeing the North Star is e- even scarier, right? 'Cause often you know what problem you wanna solve, you just haven't figured out what problem. And I, I think the, you know, Michael, my co-founder, and I, we started this company because we'd built a company in the past, and we ended up being acquired by Docker very early. It was just the founding team. And our North Star was saying, "We wanna go solve a great experience for developers with something they find really hard." But man, in the two years where we're just finding that problem, it's really scary. You know, we had a team of, uh, more than 10 people, which made that really hard.
- HTHarj Taggar
Mm.
- JMJeffrey Morgan
And I'm so thankful to that team for [laughs] staying by our side as we went through different ideas, you know? And, and what, what's not really obvious is we went from this Security for Kubernetes to then, like, security for developers on the desktop, which is, like, the pivot that we've never spoken about. And then we kind of took that form factor when models came out, and we said, "Well," it was a leap, but it was... We knew kind of the, the kind of problem and the feeling a developer wanted to have, but LLMs finally made it realize. Like, it was, it was crystal clear at the point when we tried running the Llama model and it was really hard, and we're like, "Okay, this is a problem," and it's really impressive when you get it working. And it's kind of just a zero to
- 42:15 – 44:49
The Pivot That Changed Everything
- JMJeffrey Morgan
one moment.
- HTHarj Taggar
I'm curious for the story of that pivot, 'cause there, there are actually, like, many pivots in the Ollama story, but probably, like, the most critical one was, like, the pivot to Ollama to doing, like, locally hosted LLMs. Like, how did that come about? Were you just, like, tinkering with ideas on the side, and when you found the idea, was it really obvious to everyone in the company that that was the thing to do? Or was there, like, like, like, a, a, a big debate and it wasn't until it took off that it became clear?
- JMJeffrey Morgan
W- we, you know, sat in a room together. I remember we were in Toronto, 'cause we had a team split across Toronto and Palo Alto, and, and now we're predominantly in Palo Alto. And we were saying, "Throw everything out." Like, if we had to start from scratch, if we were just joined YC right now, what would we do? And, you know, we had seen two big problems, 'cause we had talked to some users in LLMs. We had tried using open source LLMs ourselves or just LLMs in general. One problem was, could you build a gateway to access any model and host that and make that really seamless? Back then we were thinking of it as, like, the segment for-
- SPSpeaker
Hmm
- JMJeffrey Morgan
... you know, LLMs.
- HTHarj Taggar
That's a good way to think about it.
- JMJeffrey Morgan
Which I think has become really this big router idea, which is only at the beginning. It's a massive i- uh, opportunity. And the other problem was we were a bunch of ex-VMware, ex-Docker folks. We're like, "We know how to make things run," and so, like, let's help make things run with open models. And then we kind of-
- HTHarj Taggar
Systems
- JMJeffrey Morgan
... Yeah, systems
- HTHarj Taggar
... Run systems.
- JMJeffrey Morgan
And so be- we kind of tried to really introspect our team, which I wish we had done sooner, because security's a very different team and sale than developer tools. And just by doing that, we gravitated towards saying, "Let's just try this thing. Let's give ourselves two weeks to launch the first version of Ollama." And then Llama 2 came out and we said... That was right at the end of the two weeks, so we said, "Okay, we're launching it," and we just had a bias to action. If you think back, like, in two weeks, all of that happened, going from idea to shipping it, to getting to more users than we had ever had with our previous stuff. And before that was two years of just frankly overthinking the customer, the product, and just not getting something out there.
- SPSpeaker
The first time I actually heard about Ollama was on Reddit. I didn't realize it was you guys.
- SPSpeaker
[laughs]
- SPSpeaker
I was on, like, that, like, I just was interested in, like, running local models and it was on, like, the l- I think the local LLM subreddit or whatever, and everyone was just raving about Ollama and how great it was.
- SPSpeaker
I was like, "Oh, it's a YC company." [laughs]
- SPSpeaker
Yeah, [laughs] I found, I found out later actually, 'cause you, you were called a different company.
- JMJeffrey Morgan
Yeah.
- SPSpeaker
You weren't in our internal system as Ollama. [laughs]
- JMJeffrey Morgan
[laughs]
- HTHarj Taggar
[laughs]
- JMJeffrey Morgan
Yeah, I remember catching up with Jared and saying, "Oh, hey, by the way, [laughs] there's-
- HTHarj Taggar
Yeah, you know, I think-
- JMJeffrey Morgan
... all that security stuff. We, we have this Ollama thing now
- HTHarj Taggar
[laughs]
- SPSpeaker
I think you were catching up with Jared, and then I bumped into you on the stairs-
- JMJeffrey Morgan
Yeah. [laughs]
- SPSpeaker
... and I think you had your T-shirt or some swag or something.
- JMJeffrey Morgan
[laughs]
- SPSpeaker
And I was like, "All right," [laughs] like, uh, "you guys are Ollama?" [laughs] I was like, uh-
- JMJeffrey Morgan
You know, I-
- 44:49 – 47:40
100K GitHub Stars, No Revenue
- HTHarj Taggar
But it was impossible.
- JMJeffrey Morgan
Exactly.
- HTHarj Taggar
LLMs didn't exist.
- JMJeffrey Morgan
Exactly.
- SPSpeaker
Llama, Llama didn't launch yet. I guess you were also one of the first GitHub projects that very quickly got to 100,000 GitHub stars, right? Do you remember how long it... It was, like, very quick.
- JMJeffrey Morgan
Yeah. I can't remember exactly how fast, but it was much faster than Docker and Kubernetes. To your point, it, it, things kind of just started working and started taking off, and you're really as a founder just beside yourself because you can't totally explain why. I think it's the best way to explain product market fit, and there are different levels of product market fit. You know, we only started monetizing e- earlier this year with Ollama's Cloud. But just to see people fall in love with a product, it's such a zero to one moment that, um, I wish we had done it during YC-
- SPSpeaker
Mm-hmm
- JMJeffrey Morgan
... but in some ways it wasn't possible.
- HTHarj Taggar
I also think it's just kind of wild to put into perspective. Like, you sort of went from being in sort of like the cranks on Reddit, like, interested in running their own, like, rigs at home to, like, 85% of the Fortune 500-
- JMJeffrey Morgan
[laughs]
- HTHarj Taggar
... in, like, two years or something like that. [laughs]
- JMJeffrey Morgan
[laughs]
- HTHarj Taggar
That's like, that's like a pretty, uh-
- SPSpeaker
That's the Homebrew Computer Club to-
- JMJeffrey Morgan
Yeah
- SPSpeaker
... uh, broad computer adoption, like speed run that took 10 years for the PC.
- JMJeffrey Morgan
Yeah.
- SPSpeaker
It took, like, 18 months, 12 months. [laughs]
- JMJeffrey Morgan
Yeah, and that's one of the things-
- SPSpeaker
Or 11 months
- JMJeffrey Morgan
... that surprised us the most, 'cause I think, look, I think, uh, uh, Open Models, the original users, very much hobbyists just tinkering, "Oh my God, this is even possible." But very quickly because, you know, two things. One is they were free to get started with, and you could run them anywhere. That is incredibly helpful to a Fortune 500 IT developer team, 'cause they don't have to ask for permission to use it. And so it just happened what was really good for a hobbyist user translated very quickly to a developer within a business. Um, it just happened to be a case where that was, that was it. You know, for example, databases, we saw some of this too, where a database that started for devs, uh, like MongoDB very quickly also moved to enterprise. But because LLMs are stateless, it made for such an easy transition. Now, moving to the cloud, there's a lot more in play. There's an economic question if you're a customer. There's obviously security. Where's the model running? But what's beautiful about open models that both hobbyists and I- IT developers loved is you could just get started. You didn't need permission.
- HTHarj Taggar
Can we talk about the monetization angle? 'Cause this is, this is interesting too. So, like, in 2023, Llama 2 takes off. All of a sudden you've got all these users, 100,000 GitHub stars. Like, you've clearly found something. But it's basically, like, Reddit cranks who are using it. You're making no revenue, and there's no obvious path for how you will ever make any revenue-
- SPSpeaker
[laughs]
- HTHarj Taggar
... from all these, like, cranks on Reddit. It was two years before you actually figured out a business model for it, which funny enough is exactly the position that Docker was in. Like, how did you think about it during those two years? Were you, were you worried about it? Were, was the team asking, like, "What's the business model gonna be?" Uh, how did you think about, like, figuring out how to make money from it?
- 47:40 – 49:43
How Do You Monetize Open Source?
- JMJeffrey Morgan
I think there's always two ways that we saw open models, um, being able to monetize in a way that's great for the company, great for the developer, and great for the, the customer. And one of them was- A privacy-focused AI product, which Ollama started, really started with that in its open source incarnation. But we always felt that there was this moment where, you know, you weren't using Llama with, uh, the, the Llama models, for example, with, with tool calling right away when they came out. So there were use cases where it was still reserved for the frontier models, and again, with- at the risk of overthinking it, w- we kind of saw that there wasn't the product level of product market fit with open models that closed models had. And in some ways, philosophically, we want to align with when that happens, we wanna be there to capture that. I think it happened this year with coding agents running with open models, 'cause you had the largest consumption of AI being matched with finally open models being able to service that. There were a lot of opportunities along the way to do it privately, securely. Again, a lot of the Fortune 500 had already adopted Ollama. But we really asked ourselves, "What would be the most important problem we could solve for a customer?" And the local piece, while an important part of that story, never felt like the whole story, which was, how do you access open models for the hardest problems? And so in some ways, waiting, we knew we had to wait a little bit for the market to mature. At the same time, what are the risks of waiting? Well, you build a culture if, if not careful, and we had learned a lot of this from our Docker days, where you don't think about monetization. It's not a priority. I think from our previous battle scars as a team, we, we kinda had... We knew about that. But I think the other component which is really important is making sure you keep in touch with your customers. Like, one of the biggest risks of having an open source project that takes off is you consider your user base and customer base, like your customer, just a blob on the internet.
- GTGarry Tan
Mm-hmm.
- JMJeffrey Morgan
Which is a really ris- risky way to think about customers, because you wanna meet them, figure out their needs. What are they doing? What do they wanna do in six months? What's their story? And I think that's the thing I wish we had done a little more in the last few years, and we're doing a ton of that now.
- 49:43 – 51:49
Why Do YC as a Second-Time Founder?
- JFJared Friedman
One thing I'm curious, when you went through YC, you guys were, uh, second-time founders. I'm curious what got you to decide to do YC, actually.
- JMJeffrey Morgan
You know, we, we went back and forth on this for a lot, which we shouldn't have. We should've just said, "Of course we're doing YC." Um, but by and large, starting a company is a really lonely experience, even if you have a great co-founder, and Michael, my co-founder, was the co-founder of my first company, he was my college roommate at University of Waterloo. Um, but it's still lonely, and I think just having a set of peers, even though we did it during the pandemic, just talking to Jared and, like, five other groups of founders every week really helped you feel less lonely, and I think that's such an important part of it. And then, of course, when we finally moved down here and there was no more COVID, the network was just incredible, and the fact that we could meet founders building on open models, building on any kind of AI, you know, we kinda knew that was gonna happen, 'cause we had known so many founders from the University of Waterloo who had done YC pre-COVID, and they were like, "It's really about getting together." And, like, that was a big part of it, and we knew that was there, and, you know, I think that's, that's what made it a no-brainer. But also just, I think there are a lot of mistakes you can repeat that you don't have to, and what I love about the YC community is how transparent founders are with each other about those. And, you know, I still keep in touch with the founder of Docker, uh, who's an investor in our company, and we're able to talk about some of these challenges we saw in the previous generation of companies that, you know, we don't necessarily have to repeat, or things that worked and we can bring, you know, into the future.
- GTGarry Tan
Yeah, if you just, uh, don't repeat one of those mistakes, um, that, you know, sometimes is the mistake that would've killed the company.
- JMJeffrey Morgan
Potentially.
- GTGarry Tan
Yeah.
- JMJeffrey Morgan
Yeah, yeah. Our, our famous saying, you know, a bunch of our team is from, uh, companies that ended up working great, and Docker's doing phenomenal now. But whether it's, you know, some of our team was early, early at VMware, and it... There are always ups and downs, and I think just having a group of people around the table who have a collection of those, and also what worked. Actually, what worked is actually even more important, and just being able to, like, have that muscle memory is a big part of it.
- 51:49 – 53:38
What Seeing "Good" Actually Does for You
- GTGarry Tan
Oh, man. I was just thinking about this, 'cause we obviously hang out with and work with a lot of 18-year-olds or 19-year-olds, and then sometimes they're always asking like, "Well, what should I do?" And then I'm starting to realize, like, one of the more important things is if you've never worked on a team that shipped really amazing technology to, like, a lot of people-
- JMJeffrey Morgan
Mm
- GTGarry Tan
... or like, like, just real clear product market fit, like, do that once. Like, even if it's a month, even if it's, like, three months, you would learn more in those three months because then you know what good looks like. And then without that, it's like you... I mean, it's not like it's impossible. Like, people at YC do figure it out because, you know, but it's, uh, that much harder. Like, the difference between having seen something that actually works from, like, beginning to, like, some form of, like, this is what the bug database looks like-
- JMJeffrey Morgan
Yeah
- GTGarry Tan
... and this is how we release, and this is the quality that's necessary, and here's, like, the bar that we hold each other to. Having seen that, it just, like, multiplies the chance that people succeed. So it makes sense that, you know, starting off with a co-founding team that has seen a lot of that, pretty powerful.
- JMJeffrey Morgan
Yeah, I think it, it provides you a set of values you can work around, especially when you have so much power in your hands with AI. There are just parts of it that AI can help you with, but it won't hold you accountable to it. And, you know, how does software work? And look at Ollama, for example. I'm sure there's versions of Ollama running in the wild from two years ago. How will your software work when somebody falls in love with it and continues using it for two years? Is it still gonna be working? [laughs] Well, hopefully they update to the latest software, or it's, you know, a cloud service. But I think you build that muscle memory, and we definitely have that from a lot of our more senior engineers on the team who were at VMware or Nicira, for example. But at the same time, I think there are a lot of lessons we learned in the previous generation of DevOps and infrastructure that aren't valid anymore in the AI world.
- GTGarry Tan
Oh, yeah. Te- tell us about it.
- JFJared Friedman
Mm.
- JMJeffrey Morgan
Yeah.
- GTGarry Tan
What have you found?
- JFJared Friedman
What is not valid anymore?
- 53:38 – 57:14
Old DevOps Rules That Don't Apply Anymore
- JMJeffrey Morgan
I think a good example that I classically used is there was this generation of companies called platform as a service-
- JFJared Friedman
Hmm
- JMJeffrey Morgan
... in, in the early '10s.
- GTGarry Tan
Like Herokus.
- JMJeffrey Morgan
The Herokus of the world was a great example of this.
- GTGarry Tan
I mean, Docker started out as that. Docker-
- JMJeffrey Morgan
Docker started out as a platform as a service. And there's this concept that if you're a layer on top of something else, that you're in kind of a vulnerable position as a startup Which is absolutely not true in the AI world. And in fact, going up the stack can sometimes be even better because you're closer to the customer. In an infrastructure world, that, that's also the case. And that, that was like an analogy that we had to like... So many of these muscles we actually had to break building Ollama. Another one was, you know, these LLMs are never perfect, and like in the systems world, you want everything to be exactly as it's designed to run. It's tested, it's validated, but LLMs by definition are not-
- GTGarry Tan
That's a feature, not a bug
- JMJeffrey Morgan
... Exactly.
- GTGarry Tan
Yeah.
- JMJeffrey Morgan
It's a feature.
- GTGarry Tan
You want it to be a little non-deterministic, I suppose.
- JMJeffrey Morgan
And I think from building a team too, it's that, you know, with AI now, there are just problems that you don't need to staff as heavily, whereas you, you, you did 10 years ago, right? If you think about what does your customer support pipeline look like? Um, what does it look like to deliver a cloud service? Like it's a very different world with AI because how do you build a service where no engineer knows exactly how all the code works, which is obviously the case now. And so there's just new lessons we're learning going from like a, you know, some of our team from infrastructure 1.0 in the 2000s to cloud in the 2010s to now the AI space, there are a lot of rules that break.
- GTGarry Tan
I mean, you're probably actually doing an incredible service to, um, like both sides of the ecosystem in that like the end users get this like very clean thing that just works, especially like the tokens just come out and they're very clean, and the API makes sense, and it's rational and logical. And then on the flip side, like, I mean, if you don't have a layer like Ollama, I've directly experienced this where it's like, oh yeah, the underlying inference provider has a weird error for, you know, if you put this parameter in this way, or it expects JSON and, you know, it's not documented. It's just like this insane minefield. Like, you know, the agents can kind of figure it out, but like you're gonna like bang your head into the wall for like a couple hours before, you know, the agent figures it out, and in the meantime you're like, "This is a terrible experience," you know? And so you're like in there probably helping the inference providers fix all these fundamental bugs too.
- JMJeffrey Morgan
Yeah, and it's part of the, the job we do, and I think one of the big opportunities in the open model landscape is curation, and taking a fragmented universe of models and inference technology and cloud services and harnesses and like making that actually just work is a really valuable problem because the end developer, to your point, they just wanna build their software, right? They just wanna build stuff. They wanna build their next company, their next application, and I, I think that's where we come in, but it's where a ton of great services also come in, and we saw, you know, Open Router obviously is a good example of that from a wide model selection, so the developer doesn't have to sign up for, you know, 100 different providers. They can just go to one. They can pay in one place. I think we've seen with Open Code, you know, you can have one harness that integrates with any model. It's a really powerful experience for developers just looking to try the next model to see if it solves their use case better. So this curation and, you know, when there's a abundance of models and providers, now there's a scarcity in bringing that together into something that works.
- GTGarry Tan
Thank you so much for joining us. That's all we have time for.
- JMJeffrey Morgan
Thank you guys for having me. [outro music]
Episode duration: 57:15
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode rY0wnfFHYbs