EVERY SPOKEN WORD
10 min read · 1,844 words- 0:00 – 0:25
AI’s appetite for data and the Scale AI signal
- RKRam Kumar
I learned the world through internet. I think for my daughter, they will learn the world through AI. A company called Scale AI, which is known for being a data-referring company, they've got so much data from across the globe.
- SPSpeaker
People sorting, labeling, and sifting reams of data to train and improve AI for companies like Meta, OpenAI, Microsoft, and Google. Meta is paying nearly fifteen billion dollars for a Scale AI stake. I've confirmed with a source.
- 0:25 – 0:55
The hidden value of user-generated data (and why people aren’t paid)
- RKRam Kumar
This just shows you how much there is a need for data. We all contribute data to AI. Your tweet, the post that you do on Facebook, the videos that we upload on YouTube are the ones AI models are trained on. But us, as users, we don't get paid for it. The data economy today is valued about eleven trillion dollars. This is the data that's across the globe from enterprises, organizations, and individual people on the internet. And let's assume about five percentage of that is contributed by individuals. That's close to about five hundred billion dollars. Five hundred billion dollars worth of data that's
- 0:55 – 1:25
Flipping the model: data ownership and getting paid for contributions
- RKRam Kumar
taken away from you, and you're not getting paid for it. We wanna change that. What if we can flip the switch here and have people have the ownership to that? We wanna build a system where you can contribute your data and you get paid for that. It's not just for AI researchers or developers, right? The common man who owns a lot of data. We have close to about a million users who are contributing datasets for that. It's like how nations went ahead and fought for oil, now larger organizations are gonna fight for data, right? Data is the new oil, and people have to realize that,
- 1:25 – 1:55
What OpenLedger is: an AI blockchain marketplace for datasets
- RKRam Kumar
that they own that oil. They own that data, but it's time to fight back and earn a piece of that. I'm Ram. I'm one of core contributors at OpenLedger. OpenLedger is an AI blockchain where people have the datasets. AI systems need these datasets, so as an application you can use OpenLedger to go ahead and contribute a dataset that you own. We have a lot of data contributors that initially came on board. We have close to about ten ecosystem projects building AI models on us. We have close to about a million users who are contributing datasets for
- 1:55 – 2:26
From enterprise blockchain/ML work to a global, user-facing product
- RKRam Kumar
that. [upbeat music] I've been in this industry close to a decade right now. The idea was to build an R&D company around blockchain and machine learning. We saw there is a need for enterprises to bring in fairness and, like, transparency within their ecosystem, like within their organization. We had an opportunity to work with enterprises like Walmart, Sony, GSK, and many more, and what we realized is that especially a technology like blockchain brings equality among every
- 2:26 – 2:56
Convenience vs. privacy: the “my data is being used” wake-up call
- RKRam Kumar
user that uses that. OpenLedger is a contribution from that. Idea was to not just service enterprises, uh, build a product that can be used by anyone across the globe and figure out how AI is impacting everyone lives. We all have this epiphany at one point of time where you have conversations with your friend about a product that you wanna buy, and you see that ad on Instagram. You know that your data is being used. I've had multiple epiphanies as that. And we've worked with firms where we-- that is visible, right? People's information was used to make
- 2:56 – 3:56
How contributors can monetize specialized knowledge (not just public web data)
- RKRam Kumar
their product better. Sure, it gave convenience, right? But it also took privacy. That's gonna happen with AI as well. AI is gonna make money. All these large organizations are gonna make money out of it, but you're not gonna be part of that. As we evolve, right? As models evolve, models will become specialized where they need datasets from people. And in that case, we need to make sure that we can own our data and we get paid for it, and that's what OpenLedger is trying to solve. It's a platform where users can come and contribute datasets, which could be, let's say, a knowledge that I have about trading or a knowledge about a particular subject. Let's say I know to cook well. I can go ahead and contribute that. And then models can use this data, and if they use that data and they build an AI out of that, and this AI makes revenue or creates an impact, you should be part of that. You should get a piece of that revenue. We wanna bring in fairness to this ecosystem, the people who contribute data or a model developer or a compute provider or any kind of resource provider gets paid as part of the process, and that's what OpenLedger is all about. A lot of people say
- 3:56 – 4:27
AI and jobs: a new gig economy around data contribution
- RKRam Kumar
that AI is gonna make people lose jobs. I don't think so. M- it might be a temporary thing, but it's gonna create a lot of jobs. Data contribution itself could be a great gig economy. Data is also very relevant with enterprises. Data you find on internet is very generalized, but the data that you would find in a firm, in an enterprise is very specialized, right? And all of this knowledge do not come on the internet, right? People don't write blogs about it. A surgeon doesn't write about how-- what his experience in actually doing the surgery. An artist doesn't write about how he actually painted a picture.
- 4:27 – 4:57
Beyond the ChatGPT moment: specialized models need new data sources
- RKRam Kumar
So it's all comes down to the individual person's knowledge that they own. This knowledge would be needed for AI, right? For AI to truly get into all parts of our lives, more than just a chatbot, it needs to know knowledge about the entire world. It needs to know about very intricate details, right? In that case, people would reach out to, you know, individual users to get their knowledge. So we knew that ChatGPT moment were not gonna be there. It's just a spark. It's gonna become much more bigger. We're gonna build AI systems that are very specialized in various use cases, but there is
- 4:57 – 5:57
Decentralization and community power: resisting re-centralization of AI
- RKRam Kumar
not enough data out there on the internet. We've seen lot of independent developers who have a lot of innovative ideas who are actually building interesting AI models in the Asian region. They contribute datasets by using our nodes, which is basically a node they can download and have it as a plugin. They can contribute datasets for that. If you take a look at, uh, Web3's nature, internet was supposed to be decentralized, but because of convenience, we let larger organizations take that. So making sure that it does not go back to bunch of centralized larger firms is very important. We don't have money to fight for it. All we have is our own power, people coming together and building something against the larger organizations. In order for that to happen, you need to make people to come together, and community building is very important as part of that. Building a very strong culture, building a very strong community, having the same goal, building systems that are open, verifiable, and rewarding is what makes people come together. Let's take an example of Ethereum itself. Ethereum is a very community-driven blockchain, and
- 5:57 – 6:27
Proof of Attribution: on-chain tracking of data ownership, usage, and payouts
- RKRam Kumar
that is why it's so strong today. Even though it has its ups and down, uh, Ethereum as an ecosystem is very strong. That's why I think community is very important in what we're building. So to explain, uh, proof of attribution in a very simple manner, what if there's a tracking mechanism where you can see and who actually contributed for all of this? And you can also see it on chain that everyone who contributed are getting paid for it. So it's a tamper-proof record of your data's ownership. It's a record of how your data was used, and it's also a record of how much you're gonna
- 6:27 – 7:28
How the data flywheel works + a concrete example (global sleep health model)
- RKRam Kumar
get paid if your data was used as well. So the reason why we have the proof of attribution is because we need to have a trustless system where they don't have to believe Ram, right? They can believe the system. They can believe the code. I think that's very important. As a data contributor, you can go ahead and choose the model that needs the data, and you can start contributing datasets to that. It's that cyclic, uh, ecosystem we wanna build. All the data that is contributed is recorded on the blockchain because-- so that we can track this, and we can pay this. As a user, I know that I can prove my ownership by having it on chain. Another interesting side that we have started to see are smaller model developers and innovative, you know, people who wanna build something interesting have started using our product. A good example that's being built on OpenLedger is a bunch of doctors are building, uh, a sleep model. The model is trained on sleep datasets, high-quality sleep datasets across the globe. Uh, they wanna get access to the sleep datasets from various parts of the world so they can cover various races. This particular dataset is very unique because they're gonna correlate that
- 7:28 – 9:03
Why blockchain is necessary—and the future of specialized agents everywhere
- RKRam Kumar
with their health vitals. And then once the model is ready, you can just upload your sleep data, and then it's just gonna tell you what's your body vital looks like. So like this, we have very interesting players who are building models on us, very niche, innovative ideas where you need data which is quite unique. We are quite excited about them. So a lot of people ask me, like, why this has to use blockchain. It could be a traditional AI company, but I don't think so. If you take a look at generalized models, the era of that will slowly fade away. Agents will become much more specialized. There'll be an agent for healthcare. There'll be an agent for legal. There'll be agents for every other sector out there. And you can have a general model power that. You need to have a specialized model that powers it. And the dataset is among actually people. And these datasets can eventually become models, specialized AI models, which then can be consumed by apps and agents that are gonna be built on top of that. So I think every aspect of our life will change. How we take a ride home, our doctor visits, and what we learn from, all of that will change. I learned the world through internet. I think for my daughter, AI will create a huge impact, right? They will learn the world through AI. So their AI has to be responsible. So building a responsible AI system is upon us. If we encourage larger ecosystems to go ahead and consume our data and, and not reward us, then that's what is gonna happen, right? I think now it's time to change it. If we can realize how important our data is, probably that's the best output that we can see out of AI. [upbeat music]
Episode duration: 9:03
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode vrY8x-iGsFs
