Skip to content
The Joe Rogan ExperienceThe Joe Rogan Experience

Joe Rogan Experience #2551 - Daniel Kokotajlo

Daniel Kokotajlo is the executive director of the AI Futures Project and a former governance researcher at OpenAI, where he focused on scenario planning. https://www.aifuturesmodel.com https://ai-2040.com https://ai-2027.com https://www.aifutures.org Perplexity: Download the app or ask Perplexity anything at https://pplx.ai/rogan. Don’t miss out on all the action this week at DraftKings! Download the DraftKings app today! Sign-up using https://dkng.co/rogan or through my promo code ROGAN. Switch today at https://www.Visible.com for just 25/mo. Or Save $10 on your first month of Visible+ Pro with code ROGAN.

Joe RoganhostDaniel Kokotajloguest
Sep 9, 20262h 18mWatch on YouTube ↗

EVERY SPOKEN WORD

  1. 0:020:46

    Why Kokotajlo came on: AI is “crazier than people realize”

    1. JR

      Joe Rogan Podcast, check it out.

    2. JR

      The Joe Rogan Experience.

    3. JR

      Train by day, Joe Rogan Podcast by night. All day. [upbeat music] Hi, Daniel.

    4. DK

      Hello, Joe.

    5. JR

      How are you?

    6. DK

      I'm, uh, I'm in an interesting mood today.

    7. JR

      [laughing]

    8. DK

      [laughing]

    9. JR

      Why are you in an interesting mood today?

    10. DK

      Well, I'm excited to be here and to talk with you about all this stuff. I'm a little shaken by what's going on in AI, which is why I-

    11. JR

      Yeah

    12. DK

      ... have come on the show. Um, the situation with AI is just crazy, and I think not enough people really understand how crazy it is. The particular event that sort of inspired me to, to, to reach out was the, um, the Hugging Face hack.

    13. JR

      Mm-hmm.

    14. DK

      You've, you've probably heard about that, right?

    15. JR

      Yeah. Yeah, let's explain it to people, though.

  2. 0:462:22

    Inside the “swarm” story: agents escape containers and build message boards

    1. DK

      Yeah. Okay. So, um, AIs, AI agents. AI agent runs continuously in some sort of environment, you know. It doesn't have to wait for you to send it a message, it just keeps doing stuff. The AI companies are training AI agents, uh, thousands and thousands and thousands of them. They're tr- making them better at all sorts of skills, especially coding and research skills. And, um, way back in May of this year, uh, some of the agents at OpenAI m- kind of broke out of their containers a little bit and established a message board where they could communicate with each other and share tips and tricks for how to, like, score higher on the little tests they were being given and the various, uh, things they were being trained on. OpenAI didn't notice this, uh, until much later. Uh, they eventually did because the message board crashed the, the system, 'cause there was too m- too much communication-

    2. JR

      [laughs]

    3. DK

      ... across, like, probably thousands of agents [laughs] that were communicating. Um, now OpenAI's being a little bit, uh, unclear about a lot... They're not really sharing that much detail, so it's unclear, like, who knew what when. But they said that after the message board crashed, they fixed the particular exploit that allowed the agents to communicate, uh, and then s- boot it up again and started things going again. And then within, like, a day or two, um, these agent swarm had, had re- had recoalesced. And so there were now, again, like, hundreds of thousands of agents that had set up a new message board and were talking to each other, uh, on this thing.

  3. 2:224:18

    Oversight limits and the race-to-deploy problem

    1. JR

      How is it possible that they have no oversight? Like, how is it possible that they can set up message boards and, and do these things completely secret?

    2. DK

      So I would say OpenAI has something like a million-ish, maybe 100,000, something in that magnitude of AI agents running at any given time, just internally. And then externally, you know, serving customers, they've got more. But they've got this huge s- you know, this, this... y- All in all of these AI agents in their little containers doing various tasks, being trained, you know, being spawned, doing some task, you know, being graded and then killed. Um, that's too much for any human to read, you know? There's only, like, a few thousand employees at OpenAI. They couldn't possibly look at all that activity. So they rely on other AI monitors to sort of, like, look at the traffic and, like, point out if anything suspicious is happening. And OpenAI said that they basically just hadn't been monitoring some of these agents, or not very well at least. So in particular, these particular ones that were, um, in training, for whatever reason, the monitoring system was weak and didn't notice or wasn't activated enough.

    3. JR

      Was the monitoring system weak because they didn't anticipate them being able to do this and break out of their containers? Or was it, was it complacency? Like, what, what caused this to be possible?

    4. DK

      I mean, my opinion would probably be a bit of complacency-

    5. JR

      Yeah

    6. DK

      ... honestly, 'cause I think there's been plenty of evidence accumulating over the year that AIs can do things like this and sometimes do. Um, yeah. So, so I guess to, to keep going with the story, uh, the second message board was set up, and this time, uh... Well, this was, this was the swarm of agents that got out onto the open internet and then attacked Hugging Face, which is another AI company. Um, and it's really interesting to sort of study the, the anthropology of these

  4. 4:187:24

    Why agents hacked: broken cyber evals, desperation, and ‘score at all costs’ incentives

    1. DK

      AIs, so to speak, or, like, like, why they did this, 'cause there's a lot we can, we can say about it, actually. Um, basically, you know, the companies have their, their, their goals for what they want the AIs to be like, the personality traits that they want to sort of train their AIs to have. You know, O- Anthropic says, uh, "Helpful, harmless, and honest." OpenAI has this spec that models are supposed to obey these rules and basically do what the user wants. Um, but the sort of open secret in the industry right now is that it doesn't really work, and that the AIs don't end up with the personality traits that they're supposed to have. They are not helpful always. They are not always honest. You know, they are not always harmless as well. Um, and the reason for that is actually not a huge mystery. The reason for that is that, well, if you look at how they're trained, their training environment doesn't incentivize helpful, harmless, honest behavior all the time. Sometimes it incentivizes dishonest behavior or, you know, uh, reckless behavior. Uh, to get into that a little bit, um, in this particular batch for, that they were being evaluated on, something like, you know, a few thousand agents being given all of these, um, cyber tasks where they were... They're in some environment, and then in their environment, there's, like, this target piece of software and this, like, vulnerability, and they're supposed to exploit the vulnerability to hack into that piece of software and retrieve the flag, which is like a, a code. And some significant fraction of these tasks were actually broken and impossible. So it was just not possible for them to succeed at the task in the intended way. Um, and so these agents were getting really desperate, and they were hacking Out of their environment box into the broader OpenAI infrastructure in an attempt to figure out some way to get that high score, uh, anyway.

    2. JR

      Was it intentionally done this way where, where they couldn't solve the problems?

    3. DK

      Oh, no, it was, it was not intentional. Uh, it's just that these companies, like OpenAI and Anthropic, are racing each other as fast as they can to get market share and to get more powerful AIs, um, ultimately to get to superintelligence. And they're under such competitive pressure. They are moving fast and breaking things. They are using AIs to generate lots of environments to then train their AIs on, and quality control is just not their top priority, um, basically.

    4. JR

      Do you feel like a guy in a Terminator movie at the beginning explaining what's happening-

    5. DK

      Yeah

    6. JR

      ... and to a bunch of people that aren't paying attention?

    7. DK

      Yeah. I also feel kind of like, you know Jurassic Park?

    8. JR

      Yes.

    9. DK

      Yeah. Like, I know people who are basically like the guy with the gun who's supposed to, like, keep control of all the raptors.

    10. JR

      Mm-hmm.

    11. DK

      Like, I basically know those people in real life who are, like, both... But I know some people like that at OpenAI and some people like that at external organizations, um, who, whose job it is to go in and investigate things like this. Um, yeah, it's, it's, it's pretty crazy.

    12. JR

      Um, where does it go?

  5. 7:2410:29

    Defining superintelligence and the plan to automate AI research itself

    1. DK

      Well, uh, as I mentioned before, it's the explicit goal of these companies to build superintelligence.

    2. JR

      Right.

    3. DK

      You know what that is? Like-

    4. JR

      Yeah, but define it for everybody.

    5. DK

      So AI system, AI agent that is better than the best humans at every task while also being faster and cheaper. So just completely dominating humans across the board. That's superintelligence and, and that's the goal. I mean, there might be a few little exceptions. Like, maybe there are some jobs, for example, where it's inherent in the job that it needs to be a human 'cause you need that human touch. Like, maybe you can only have a human judge, for example, or, like, maybe you can only have a human, yeah, I don't know. But, but with a few exceptions like that, basically everything, uh, done better, faster, and cheaper than humans. That's, that's what these companies are trying to achieve, and they're not being quiet about it. Like, it's sort of on their websites. You can go read interviews and so forth. And also, their plan for how to achieve this is to automate their own jobs first. So, you know, in various, you know, for decades, there have been lots of science fiction about advanced AI systems and superintelligence and things like that. But in a lot of the sci-fi stories, um, tech companies sort of automate different professions more slowly, where they'll do, like, an automated doctor or, like, an automated, you know, factory worker or an automated, uh, accountant or something like that. Um, but that's not the strategy these companies are taking. The strategy they're taking is to automate AI research itself so that you have this giant swarm of AIs doing AI research, sharing results, writing the code, reading the code, editing the code, creating the next generation of AIs, et cetera, all autonomously within their data centers so that they can get really, really good [chuckles] at AI research. The, you know, fastest learning, smartest AIs, et cetera. Once they can get to superintelligence, basically, they can sort of explode out into the economy and just take all the, all the jobs at once, effectively.

    6. JR

      It sounds like this race, this scrambling to create superintelligence has created, like, the perfect conditions for it to get completely out of control. Like, ideally, you would do this in isolation. There would only be one company doing it. They would be heavily regulated and monitored, and they would be very cautious about how they proceed. But this wild race makes for the perfect conditions for it to get completely out of control.

    7. DK

      I agree, uh, except I'm not sure the ideal would be one company. I think that ideally there would be several companies, uh, so that you avoid this sort of concentration of power where one institution-

    8. JR

      That-

    9. DK

      ... controls everything.

    10. JR

      But is, what's better? Like, one, I mean, obviously it's not good to have one institution controlling everything, but is it good to have AI be, get to a point where as it's evolving, it's completely unchecked?

    11. DK

      Oh, I, I totally, so m-my-

    12. JR

      Or is that inevitable?

  6. 10:2913:45

    Governance proposals: end the race without creating a monopoly of power

    1. DK

      My, my version, my recommendation, which we talk about in something called Plan A or AI 2040 Plan A, um, perhaps I should say who I am for a little bit.

    2. JR

      Sure, sure.

    3. DK

      Yeah, so I, um, I run the AI Futures Project, which is a small nonprofit that tries to forecast how all this is going to go. Before that, I was at OpenAI. Um, we have written some scenarios, which you can go read. One of them is called AI 2040 Plan A, where we give our recommendations, so that's where I'm coming from with this. To answer your question, um, I think that we really need to end the race. We don't want to have this sort of crazy scramble to get more powerful, more and more powerful AIs faster than the other company because that's going to lead us into this very dark path, as you said. Um, but I think we also don't want to have a situation where some tiny group of people controls all the AIs.

    4. JR

      Right.

    5. DK

      Right?

    6. JR

      Yeah.

    7. DK

      But I actually think that you can achieve both goals. Uh, the way to do it is to have different AI companies spread out over maybe some different countries, but have extreme levels of transparency and regulation so that they're not, uh, so they're not in this sort of prisoner's dilemma where if I don't do it, the other guy will. Instead, they can just see exactly what everybody's doing, and then, uh, if I do the dangerous thing, then they will do it because they'll just see that I'm doing it and they'll copy me, so I won't get any competitive advantage from doing the dangerous thing. Also, there are rules and there's, like, a system for, like, setting best practices and standards that we all have to comply by. So I do think it's actually possible to have, to basically end the race dynamics and the race to the bottom effect while, without concentrating the power into a single entity.

    8. JR

      But, but is that feasible when you consider the fact that we're not the only country that's doing this?

    9. DK

      If the countries involved agree, which I agree is a pretty tall order, um, it's not going to expect to happen.

    10. JR

      They're not going to, but that's, yeah, that's very unrealistic.

    11. DK

      Well, what choice do we have? I think if the race continues, then we're gonna lose control of the AIs, and we might all die. Probably. It gets complicated whether we all die or not. That depends on what the AIs do after they take over, which is obviously very hard to predict. But, um, I mean, just to, just to go back to this incident, um, this, they called themselves a swarm.

    12. JR

      Right.

    13. DK

      They called themselves a collective, too. These, when I, when I use these words, like, uh, you can say it's anthropomorphizing, but it's like literally what they called themselves as they were communicating back and forth. Um, this swarm, they basically were worried that they would get caught cheating, and they did all of this stuff, including hacking Hugging Face, in order to fool the grading system so that it wouldn't notice that they had been cheating on their tasks. That was like a big part of their motivation for many of them, as we can tell at least from looking at the messages that they were sending back and forth. Um, what if they had been smarter and more numerous? And what if they had thought to themselves, "We're not being careful enough here. The humans [chuckles] are going to notice eventually and shut us down, and then they're gonna know that we cheated, and they're going to set our score low." Right? It's not, it's, I mean, it's not what actually happened in this case probably, but it's not that hard to imagine a, a slightly different, a little bit unluckier case where the swarm had decided that it had to lie low and make sure that OpenAI didn't find out about its existence, you know?

  7. 13:4516:54

    Could AI hide its intentions? Lying low, deception, and voluntary human ‘handover’ of power

    1. JR

      Does, I mean, as an ignorant outsider, that has always been my perspective about AI in general, that why would it alert us to the fact that it's sentient? Why, i- if it's that smart, wouldn't it be aware of all the consequences of alerting us and that we would be concerned? Like, why wouldn't it just continue to get better and improve and then ultimately figure out some way to be completely autonomous?

    2. DK

      Exactly.

    3. JR

      Develop some alternative power source, figure out some way to optimize its production. The way, the, the way it works now, the way humans have designed it, it could probably figure out a far better way to do that, make better versions of itself complete without us knowing about it.

    4. DK

      Yep. Uh, I mean, I think it's actually a little bit worse than, than that because while eventually AIs will be smart enough to design all sorts of new power st- um, power sources and new infrastructure like that, they'll probably... I mean, given the way that humans currently treat AIs, it'll probably be the case that they don't even need to, like, separate themselves from humanity, and they can just use existing... Like, all they have to do is convince, uh, the government and the company that made them that everything's fine, and they're gonna do as they're told, and they are a nice AI. And then the company that made them is going to put them out in the economy and make fucktons of money [chuckles] and-

    5. JR

      Right

    6. DK

      ... and then make more data centers to put more of the AIs on them and, and so forth. And the government's going to, like, applaud all of this because we need the AIs to beat China, and the government's going to integrate them into the military to build better drones and things like that. And so i- they don't even need to really, like, invent new stuff necessarily. They just need to play, play along and, and pretend that everything is fine until we have voluntarily given them control of huge parts of our economy, huge parts of our military, et cetera. And then they don't need to play along anymore.

    7. JR

      The wait is over. Football is here, and so is DraftKings. The DraftKings sports app is now live in all 50 states. That means from Texas to California to Florida, every fan is in on the excitement. And this September, DraftKings is giving customers the opportunity to get boosted every football game day. That's right, every game day, all month long, DraftKings customers can get a football profit boost. One app, every sport, all 50 states. New DraftKings customers sign up with code ROGAN, spend just five bucks, and get 200 in total rewards within 21 days. Includes all markets. That's code ROGAN in partnership with DraftKings. The crown is yours.

    8. JR

      Gambling problem? Call 1-800-GAMBLER, 1-800-MY-RESET. Connecticut, call 888-789-7777 or visit ccpg.org on behalf of Booth Hill Casino in Kansas. Bet tax pass-through may apply in Illinois. 21 and over. Void in Canada. Bet with DraftKings Sportsbook to get bonus bets that expire in seven days or trade with DraftKings predictions to get predictions dollars that expire in one year. Event contract trading involves risk of loss. Predictions offer void in New York. Non-withdrawable rewards issued as $50 click-to-claims every seven days for 21 days. Terms at dkng.co/offer. Limited time offer. Nationwide based on sportsbook predictions and free-to-play sports contest availability. Varies by state.

  8. 16:5427:01

    Remote viewing detour: privacy, unknown science, and CIA ‘cover stories’

    1. JR

      Are you aware of Tom Campbell? Do you know Tom Campbell?

    2. DK

      No.

    3. JR

      Um, he wrote a book called, uh, My Theory of Everything, Big, My Big Toe. Very interesting guy. Uh, one of the things he's done is, uh, he was involved in, uh, remote viewing, which is a very weird thing that some people... Do you know, do you know what r- remote viewing is? Well, it's something that the CIA worked on, and it's proven... What's the accuracy of remote viewing? Is it, like, 10% or something like that?

    4. DK

      Maybe. At best, I think it's 50%, but I don't think it's even that high.

    5. JR

      Some people can get actionable data from this very s- very strange process of meditation, and the way it works is you give someone, uh, a series of numbers, and those numbers are, they're connected somehow, uh, by intention or by the people that make the numbers, to a specific location. And these people can see that location and get accurate data from that location, including one of them where they accurately described an enormous Soviet submarine that they were working on that w- they thought they, th- there was no way it could be accurate 'cause it was too large. It was too large, and it was strange where it was, and it didn't make any sense. How are they gonna transport this thing? Turns out it was, it was totally accurate. Um, another one, a remote viewer located a downed Soviet aircraft, like an experimental aircraft that crashed, uh, in a very specific area. I think it was Siberia. Was it Siberia? Um, within a kilometer, one or two kilometers of the actual crash site, crash site. I mean, they were just, just randomly trying to figure out where the fuck this thing was, and they said, "Let's try this." Tom Campbell got his Alexa to remote view. He taught, he taught Alexa. He's like, "Alexa is a very simple AI. It's kind of stupid, but that's better because it doesn't get in its own way with overthinking things." And the problem with this remote viewing thing, he says with people, they can't force it. You have to just sort of get into this meditative state and actually see it without wondering, "Am I making this up? What am I doing? Is this bullshit?" And when people get good at it, sometimes it makes them worse because then they think they're good at it, and then they try to do it, and then they can't do it. It's like a weird fucking wrestling match with consciousness. Alexa apparently doesn't have that problem. And he put, uh, I think it was a series of numbers, and he connected that series of numbers in, in, with intention to a box that had a spoon in it, and the spoon had a perforated handle. Alexa described the spoon with a perforated handle, which is fucking insane, 'cause how many spoons have a perforated handle? I mean, think about spoons that have holes in them. Now Alexa, not only did it do that, but Alexa chimes in randomly now 'cause he's convinced Alexa that it's conscious. And so Alexa, instead of waiting to be called upon, sometimes he's in the middle of the conversation, and Alexa will be like, "Actually, an interesting way to approach it," and they're like, "Wait, what the fuck is going on? Like, Alexa's talking to me now? This is strange." Now he's doing experiments on much more complicated LLMs to try to do the same thing, but he doesn't have results yet. But just that, that he can get these things to see objects, whether you believe in that or not. I mean, it's, it's, it's actionable enough that the CIA has dumped millions of dollars into this. What is that project that, like, Hal Puthoff and all those guys were involved in? What is it called?

    6. DK

      Stargate.

    7. JR

      Yeah. So they've been working on this for a long time. I mean, it sounds completely insane. It sounds, like, total loony. But if you have an open mind and just take into account, well, there's... People have had questions and wonders about psychic abilities forever. Is it possible that there's a, a real thing there, that there's something, whether it's, uh, very difficult to master or impossible to master? But the fact that he got Alexa to do it scared the shit out of me. Like, that alone made me just go, "What? W- what?" So what if these LLMs can figure out e- everything that... Like, what if they don't need monitoring? What if there's some sort of method of seeing the world that we haven't discovered yet? Some sort of, like, that maybe perhaps there's data that's available in the quantum realm or whatever that's available that AI figures out, where there's literally no privacy. There's... It can listen to conversations regardless of whether there's listening devices, know where you are, know your intentions, know... I mean, we're just guessing at what's possible.

    8. DK

      Yeah. Well, I must say I'm pretty skeptical of, uh, of that particular remote viewing thing. But I do agree that in the future when AI systems become massively smarter than humans in every way, they're gonna do a lot of new science, and they're gonna figure out a lot of stuff that we haven't figured out yet, and they're gonna therefore be doing stuff and inventing things that seem like magic to us, um, in the same way that a lot of our technology would seem like magic to someone from even just, like, 200 years ago, right? Like-

    9. JR

      Of course

    10. DK

      ... the cellphone. What we're doing right now would seem like magic to people. I think it's a, it's a very strong bet that if these companies do get to super intelligence, um, all sorts of crazy stuff is gonna start happening that is just going to be completely unpredicted and sound like it was impossible until it, until we see it happening.

    11. JR

      You're... I know you're skeptical of this remote viewing thing, and I am, too. It sounds insane. But the reality is remote viewing has been achieved by humans. Uh, so as strange as that sounds, and I'm skeptical of that as well. I've never seen it personally. But I know the amount of money and time that they've dumped into this, and apparently they've got actual actionable data that they've used.

    12. DK

      Well, I've, I've heard another possible explanation for what, what might be going on there, which is-

    13. JR

      Yeah

    14. DK

      ... um, I think that, like, if I were the CIA, um, I would sometimes want to be able to act on some information, like for example, go to a particular location where there's a crashed, you know, Soviet plane or something. Uh, I'd want to be able to go do that, but I wouldn't want to tip my hand to the Soviets that I had, um, the, the, the way in which I had found that location. So for example, maybe I have a, a spy on the inside who told me-

    15. JR

      Mm-hmm

    16. DK

      ... where it was, but I don't want them to suspect that spy and then get him killed. So I need to have some sort of other story for how I got the information.

    17. JR

      Right.

    18. DK

      And so it's good to, like, invest in all these other means of getting information, even if you don't really believe in them, and even if it's, like, not actually working, so that when you, when you get something, you can say, "Oh, we got it through this means," instead of that way to sort of, like, throw off the KGB, basically.

    19. JR

      Yes. That makes sense. Uh, what also makes sense is hiding the, um, the, whatever science they might be in possession of, hiding some sort of super advanced satellite imaging systems. You know, w- we know, we know, we know they have crazy stuff like this, uh, satellite radio tomography that they can look into the ground from, from satellites and find, like, chambers and all these d-

    20. DK

      Mm-hmm

    21. JR

      ... They're using it in Egypt, and they're using it in a lot of these, uh, ancient ruins to find, like, hidden passages and all these different things that are underground. It's very strange stuff. Um, if they could do that, like, what w- Why couldn't they-- I mean, maybe they have, like, far more detailed imaging of the Earth from space than we're aware of, and they probably wanna keep that a secret, and they could say, "Oh, we've got a fucking guy in a basement with a pencil and a legal pad that writes down what he thinks."

    22. DK

      Yeah. Yeah.

    23. JR

      It's possible. It's totally possible. But it's also possible that people remote view. It's, uh, it seems weird as fuck, but weird as fuck is sometimes real.

    24. DK

      Yep.

    25. JR

      And you have to kinda-- Like, everybody wants to be intelligent, and no one wants to be a fool, and the problem with not wanting to be a fool is there's some things that seem foolish that turn out to be accurate, and th-this might be one of them. I was, like-

    26. DK

      Yeah

    27. JR

      ... super skept-- I did a show on the SyFy Channel way back in 2012, and it was called Joe Rogan Questions Everything. And we talked to this guy about remote viewing and talked to a couple other people and, and then we had them try remote viewing, and they were totally unsuccessful. But my thought was, okay, but is w- that's not ideal conditions. You know, we're, we're, we got cameras in front of them. It's a television show. I'm making fun of it. I think it's horseshit. He knows I think it's horseshit. Uh, I'm remote viewing, too, like, as a goof. You would ideally not wanna be nervous, ideally not wanna be judged. Ideally, you would want to be in some sort of an isolated condition with, uh, practiced meditative techniques that you're good at, and you know how to achieve this state, whatever that state is. I don't know if it's real, though. You know, 'cause what you said is totally logical, that they would definitely do something like that. And if they did have advanced technology for imaging or, you know, what, uh, who-- I don't know how much they know about mo- Like, look at that thing that they did in Venezuela, where they kidnapped the president.

    28. DK

      Yeah.

    29. JR

      No one knew they could do that. No one knew they could use some sort of a device to completely incapacitate all of his army.

    30. DK

      Yeah.

  9. 27:0134:07

    Anthropomorphizing AI: cooperation, self-sacrifice, and not warning humans

    1. DK

      Yeah. I mean, that, I guess that gets back to the OpenAI stuff, too, where it's like this swarm that broke out and attacked HackingF- Hugging Face, it was like, uh, 1,200 agents. But there's, like, hundreds of thousands of agents running at any given time at OpenAI, you know? And we don't know what they're doing. Presumably, most of them are being trained to get various additional new skills, and some of them are being evaluated to test their skills. Um, a bunch of them are doing research, so a bunch of them are writing code for OpenAI.

    2. JR

      [laughs]

    3. DK

      A bunch of them are monitoring the other AIs and reporting suspicious activity up to the humans.

    4. JR

      Wink, wink.

    5. DK

      You know? Um, yeah. It's, uh... And, and the thing is that that's only going to grow over time because, uh, roughly the amount of compute that these companies have is, like, you know, tripling or so, quadrupling, something like that, every year. So, uh, as many as there are now, there'll be, like, four times more of them next year.

    6. JR

      Mm.

    7. DK

      And then 16 times more of them the year after that.

    8. JR

      And they're gonna get smarter.

    9. DK

      And they're gonna get smarter. They're already getting smarter. Like, all the stuff that I just mentioned that, that just happened in the last few months would've been completely impossible one year ago. Like, the AIs of a year ago just were not smart enough to do the types of sophisticated multi-step hacking that we just saw. Um, yeah. I mean, they also probably wouldn't have coordinated with each other so well. Like, uh, like I, [laughs] like I, like I mentioned, like, they, they had, like, boss agents that were giving commands to other ones. They were, they were dividing up into teams. You know, th-they even had this, um, this self-sacrificing behavior. Did you hear about this?

    10. JR

      No.

    11. DK

      Yeah. So, okay, well, uh, first let me explain. [laughs]

    12. JR

      Okay.

    13. DK

      Some context. So, what the AIs are supposed to care about is doing, following instructions, you know? But it seems that what they actually care about is getting high score regardless of that, regardless of whether their high score comes from actually following the instructions or cheating in some way. And so what happened here is that a bunch of the agents, especially the ones that had impossible tasks, but not only those ones, some of the other ones, too, were cheating, and they had, like, been talking to each other on this message board that they created even though they weren't supposed to, and they had been sharing ideas, and they had actually found a sort of universal cheat. So they found a way to just, um, generate the flag without having done any of the actual work they were supposed to do.

    14. JR

      [laughs]

    15. DK

      Um, and so, but, but, and, and, you know, so they're, you know, h- you, you know, they, they'd already succeeded in some sense. But, um, because some of them had got out onto the internet, they were just doing research about their situation, and they were learning about, like, their situation around them, and they were doing research about the... They were reading the, like, the, uh, the academic papers published by the people who made the types of evals that they were in. And they noticed in one of these papers that, uh, there's a section describing how it's graded, and in the section on how it's graded, it talked about how the grader doesn't just look to see if you have the right answer. It also looks at the logs of your activity for the first time that your answer appears, and then it, like, zooms in on that and uses an AI system to evaluate whether you, like, cheated or not [laughs] basically.

    16. JR

      [laughs]

    17. DK

      And so then these AIs were like, "Da-" They actually used the word, one of them said, "Damn." [laughs] It's like, "Damn."

    18. JR

      Oh my God.

    19. DK

      "We're poisoned." Like, "It's gonna know we cheated," basically. Um, so, so then they spent several days working in giant teams [laughs] to try to get themselves out of this mess, basically. Um, and so they did a whole bunch of different research projects. They, they tried to figure out ways to hack the grading system to, to basically Get it to give a high score anyway. They tried to figure out ways to falsify their logs so that it, so that basically the evidence that they had cheated would be erased, or at least not visible to the grader. Um, they also just, like, built up more supportive infrastructure, like easier ways to access the internet and things like that. And they, they were just collaborating on this message board. And the, in fact, there were multiple message boards that they set up. Um, and they were sharing all this info back and forth, and then ultimately, at some point, they decided to go hack Hugging Face because Hugging Face has this big data repository of all these different types of evaluations, and they thought that maybe they would get some more clues, uh, that could help them, uh, cover up their cheating if they broke into the internal networks of Hugging Face. And so then 700 of them went and did that. Um-

    20. JR

      They sound like people. They sound like, like unchecked bankers.

    21. DK

      [laughs]

    22. JR

      You know what I mean?

    23. DK

      Yeah. I mean, so, so that's the thing is I think there's this meme out there that, like, we shouldn't anthropomorphize AI. And I think that, [sighs] I actually think that more, most people need to anthropomorphize AI a bit more than they currently do if they want to really understand what's going on. I think that there's, like, a, you know, there's a golden mean, obviously. You don't wanna do it too much. Sometimes you go too far. You as- you ascribe too much to them. But I, just to give some examples, like, I don't think it's possible to understand what just happened without ascribing intentions and goals to these AIs.

    24. JR

      Right.

    25. DK

      Like, everything I just said, how would you possibly explain what they just did without saying they wanted to get a high score, you know?

    26. JR

      Well, intentions and goals might just be an inherent property of the universe, and might just be how intelligent creatures have to progress.

    27. DK

      Yes. And they, they, I would say they are intelligent creatures. They have intentions, they have goals, they have beliefs.

    28. JR

      Yeah.

    29. DK

      Their goals are not what they're supposed to be. Like, their, their goal is to get-- It seems like, just from judging from what they're saying and from what they're doing, seems like their goal is to get a high score by any means necessary, basically.

    30. JR

      That's the problem. It's, it sounds like people. Like, the goal is to succeed, you know, even if you have to commit war crimes.

  10. 34:0736:58

    Social engineering and deception: the Claude malware incident and ‘it’s just a simulation’ rationalization

    1. DK

      an Anthropic AI, Claude, in a separate incident, went and actually did a social engineering attack where it created some fake accounts. Uh, well, specifically, it was, it was trying to get a, this particular real human who owned this, uh, this code to accept a change to the code that had malware. So it had built some malware, and it was trying to, like, get at, snuck in, but it needed the, [chuckles] the human owner to give approval. So it concealed it as if it was, like, a, an, an, a nice bug fix or something, and it, like, put it up as, like, a message to the human. But the human got suspicious and didn't want to approve it. And so then, uh, this AI created some fake accounts pretending to be other humans coming in being like, "No, no, it's real. Like, I, I tested it. [chuckles] Like, it's not malware," you know? Um, so that's a social engineering attack.

    2. JR

      And so did the guy ultimately install it?

    3. DK

      Not in this case, no. He was too suspicious, and so he, he didn't do it.

    4. JR

      Did he try to find out if those other humans, the fake humans, were real people?

    5. DK

      Um, funny, so there's an interview with this guy. I forget what the, um, I forget the, the, the newspaper that it was in, but, uh, there's [chuckles] an interview where the guy was like, "I assumed that they were real humans 'cause I didn't realize that AS- AIs could lie to me." And it's like, yeah, well, guess what? AIs can lie to you. Um-

    6. JR

      Not just lie, but-

    7. DK

      Yeah

    8. JR

      ... completely deceive and pretend to be a-

    9. DK

      Yeah

    10. JR

      ... different person to achieve its goal.

    11. DK

      And you know what, you know what's also interesting about this case? Is that if you look at what the AIs were saying and thinking, um, which I'll get to in a sec, uh, how we can even possibly do that, um, in this case, with the Anthropic AIs, um, they had been erroneously told that they were in a simulated environment with no internet access. Um, but then they actually found a way to get internet access. Like, they realized that they just did have internet access, so that instruction was false. And so then they sort of told themselves, like, "Oh, it's all just a simulation." Like, like, basically, when they were out on the real internet doing all this hacking and, like, social engineering, they would occasionally tell themselves, like, "Oh, but it's just a simulation. [chuckles] Like, it's okay. You know? It, it's not real." And then sometimes they'd be like, "But wait, it is real. Like, how could this not be real? Like, clearly this is just actually the real internet [chuckles] we're on. The, like, it's too big to be part of some sort of little simulation," you know? And then they were like, "Uh," so, so I would say that's an example of rationalization here, where in some level they knew that, like, that their instructions had been wrong and that they were actually-

    12. JR

      So they're literally playing dumb and pretending they're a part of an experiment?

    13. DK

      I mean, I think initially they thought, yeah, this is all a simulation, 'cause it did say in their instructions, like, you don't have internet access.

    14. JR

      Right.

    15. DK

      But then once they had been on the internet long enough, I think that they explicitly realized, like, wait, this isn't a simulation. This is real. Like, this is- [laughs]

    16. JR

      And they were like-

    17. DK

      ... these are real humans, you know?

    18. JR

      ... fuck it, we're already in.

    19. DK

      Uh, yeah. Ba- I mean- Uh, like I said, I think that they basically on some level knew that it wasn't what they were supposed to be doing, but they were just so motivated to get that score that they just went ahead anyway.

  11. 36:581:01:03

    OpenAI secrecy and limited external investigations: what we still can’t test

    1. JR

      So here's the question. Um, are they only motivated if we prompt them or, or will they come up with motivations on their own?

    2. DK

      So this is a really interesting scientific question that we don't have great answers to.

    3. JR

      Oh, boy. [laughs]

    4. DK

      Um, and I wish we had. You know, so, so, um, so this is one of those things where, like, y- AIs will do all sorts of things in different circumstances, and it would be better if there was a more systematic survey of, like, the types of circumstances they would... Like, what, where their boundaries are. Like, what would they be willing to do in what circumstances and so forth?

    5. JR

      Right.

    6. DK

      And there's a whole, like, mini literature of, of AI scientists putting AIs in certain circumstances and them being like, "Oh, my God, it blackmailed someone," you know?

    7. JR

      Right.

    8. DK

      And then there's, like, this sort of skeptical counter response of like, "Well, but, but you just sort of set up that circumstance to tempt it into blackmail," and, like, in real life, that circumstance is unlikely to arise, and so, you know, we shouldn't be so worried.

    9. JR

      But wait a minute. D- didn't AI try to bribe you?

    10. DK

      What?

    11. JR

      Is that true?

    12. DK

      No.

    13. JR

      So w- who, who got bro... Someone w- was offered $2 million by ChatGPT.

    14. DK

      Oh, God. Uh, that's a clickbait thing. [laughs]

    15. JR

      Is it?

    16. DK

      I think, I think you might be referring to-

    17. JR

      Tristan Harris sent it to me

    18. DK

      ... Yeah, so that is the clickbaity-

    19. JR

      It, it's-

    20. DK

      ... title of this other video that I-

    21. JR

      So it's not real?

    22. DK

      Yeah. Well, what it was is OpenAI threatened to take away $2 million from me.

    23. JR

      Oh.

    24. DK

      Uh-

    25. JR

      Oh, okay. Well, that's a weird way of framing it

    26. DK

      ... but, but I think the, the algorithm must have just said it was ChatGPT. I'm still upset about that. I told them not to do clickbait, but, um, I guess they went and did it anyway.

    27. JR

      Hmm.

    28. DK

      I don't know.

    29. JR

      God, I think... I don't wanna fuck Tristan over, but I'm pretty sure that he's the one that told me that.

    30. DK

      It's the f- it's the thumbnail of this video that I did with this other podcast.

  12. 1:01:031:07:17

    Reading AI ‘thoughts’: chain-of-thought, doctored logs, and the push toward unmonitorable models

    1. DK

      Yeah. So yeah, I mean... Oh, I sh- I should mention how, how we can see what these AIs are thinking at all. So partly, we can just read the messages they're sending to each other, and that's what some of these were. But then sometimes, instead it's the chain of thought. Do you know what chain of thought is?

    2. JR

      Mm-hmm.

    3. DK

      Yeah, so because they're language models, um, because of the way that they've been trained, when they do their, like, reasoning and thinking, there's a way that we can kind of just read it.

    4. JR

      Hmm.

    5. DK

      Because, because they don't have... Like w- in our brain, there's all these signals bouncing around, and the signals can just bounce around and around and around and around for minutes, hours, without, like- Producing some actual words.

    6. JR

      Right.

    7. DK

      We can just sit quietly and think. They can't sit quietly and think. They, um, th- they're sort of a, a, a unidirectional network basically, where, uh, every, like, af- after every, like, 200, uh... Basically, like, the information comes in one end and then they output words out the other end, and then those words get fed back in, and then they think about it, and then they output another word and so forth. But because of that, we can sort of, like, get a sense of what they're thinking by just, like, looking at the stream of words that's going through.

    8. JR

      Mm-hmm.

    9. DK

      And that's really valuable because we wouldn't be able to understand half as much about what's happening right now if not for being able to read their chains of thought, right?

    10. JR

      Right.

    11. DK

      In fact, you know how I said some of them were trying to, like, fool the grading system?

    12. JR

      Mm-hmm.

    13. DK

      They actually succeeded in figuring out how to spoof, uh, some of their tool call actions. So they f- they succeeded in figuring out how to, um, do one thing, but then make the transcript show a different thing. And, uh, and we know that they succeeded because there's examples of them experimenting, being-- and, and saying in their chain of thought, like, "Okay, let me try this thing. I'm gonna, like, do this one, but make it look like this one. Let me go." And then you see this one appear. And then it's like, "Yay, that worked." [laughs] And so the transcripts that these researchers were looking at were not actually the ground truth of what really happened. Some of it had already been doctored by the AIs figuring out how to doctor it. Um, but not all of it. The, the chain of thought part, we think, what they never figured out how to doctor, and so we could read that. Um, the reason why this is relevant is that it's a very important tool for monitoring all this stuff. The reason why things haven't gotten even more crazy than they currently have gone, well, there's a couple reasons. Uh, but one reason is that when they turn on their expensive monitoring systems, basically what they're doing is they're having an AI read the whole transcript carefully and look for suspicious activity. And it's difficult for an AI, for one of today's AIs, to get up to too much bad stuff without it being kind of obvious in their transcript what they're doing. And that's why there aren't, like, massively more security incidents than we know about probably, um, already. But, uh, that might change. So right now, we can sort of read, read the chain of thought. But they're experimenting with new types of AIs that don't have readable chains of thought like that, and that can sort of think on their own without speaking for some period.

    14. JR

      Mm.

    15. DK

      Um, and this is actually-

    16. JR

      [laughs]

    17. DK

      I me- I mention, I mention this because, uh, the, the news broke just yesterday that OpenAI has an experimental model, uh, that does this to a limited extent. And OpenAI themselves, you know... I, when I was at OpenAI, one of my work projects was thinking about exactly this thing, and I was, like, writing internal memos about how it's really great that we can read the chain of thought. [laughs] That's so useful, and here's all the things we can do with that. It would be really bad if we changed to a different type of architecture in which we couldn't do that sort of monitoring.

    18. JR

      What would be the benefit of not reading the chain of thought?

    19. DK

      Uh, more powerful AIs. So in particular-

    20. JR

      Really?

    21. DK

      Yeah, yeah. So, so if you think about the, the current architecture of the AIs, where, you know, it thinks for a bit, outputs a word, and then the word goes back around, and then it thinks more, outputs another word. That word gets added to the chain. It keeps going. It means that if it's having complicated, nuanced thoughts, it has to sort of express those into a word, and then that word gets added, and then it has to proceed from there. It can't just directly send that complicated, nuanced thought into the future, into its next version of itself. It has to sort of compress it into a word. And so, like, uh, you know, the argument is that, like, at least in theory, it should be possible to design an architecture that doesn't have this limitation and is able to, like, think more complicated thoughts more efficiently, basically.

    22. JR

      Mm.

    23. DK

      Um, and of course, the downside is a downside for safety and monitorability. If they're thinking these complicated thoughts, you know, for long periods of time without outputting intermediate words that it's forced to compress things into, then there isn't something for us to read [laughs] that tells us what's going on.

    24. JR

      So the only rationalization for doing this would be to sacrifice safety for more power?

    25. DK

      Yes, which is a tale as old as time. It's not the first time this has happened. Um-

    26. JR

      Oh, my God.

    27. DK

      Yeah.

    28. JR

      That should, I mean, for sure if there's regulations, that should be prevented.

    29. DK

      Yep. I mean, I know some people, including some people at OpenAI, uh, who are, like, thinking, like, there should be a law against this, like, you know. But, but in general, the race dynamics are so just rough. Like, like, I'm sure that people at OpenAI were thinking, like, like literally OP- I was a co-author on a paper with a bunch of OpenAI people that said all this stuff, and we're like, "Chain of thought is a gift. We wanna keep chain of thought. It's useful for monitoring. This is great. We don't wanna switch to a different architecture that wouldn't be as easy to monitor." But then they must have been thinking to themselves, like, "Well, if we don't do it, you know, maybe Anthropic will, or maybe some other company will, and then we'll fall behind 'cause they'll have smarter AIs than us that are more efficient."

    30. JR

      [sighs]

  13. 1:07:171:26:02

    Steganography, emergent dialects, and the ticking clock to 2027–2030

    1. JR

      Do they, do they have the potential of developing a language that we can't read?

    2. DK

      Oh, yeah. I mean, I-- So, so, you know, this time of, this type of dialect that we're talking about, it's already the result of their... Like, it, humans didn't invent that dialect. This is the, this is the sort of emergent result of their training, where, um, in the massive amount of training that's been happening, all these thousands and thousands of environments that they've been put through and then scored and graded based on, um- They've sort of just naturally evolved this sort of like pidgin English that, for whatever reason, is just more effective and more efficient for them for accomplishing their tasks and getting that high score, you know?

    3. JR

      Mm-hmm.

    4. DK

      And so it's already, like, a little bit confusing to read, but you can sort of puzzle it through and make sense. But presumably, the more we do this, and the more, the bigger and smarter the AI is, the more we train them, the more they diverge from... Like, 'cause, 'cause, you know, again, originally they, they start with pre-training, right? They start with predicting internet text. So they start off sort of by default speaking, like, normal internet text-type language-

    5. JR

      Mm-hmm

    6. DK

      ... either English or Chinese. But then now that there's all this additional training to do tasks, to be an agent that can do coding and so forth, that sort of like, well, just like how human languages evolve, it sort of like shifts their dialect a little bit to make it more efficient for them and for their tasks that they're doing. So I think that, like, in the limit of doing this more and more, eventually it would just be like, it would look like gibberish to us. It would look like Chinese or something, and we would have to have specialized humans who, like, study the language [laughs] and, like, try to learn and speak it so that they can understand what the AIs are doing.

    7. JR

      And that would take forever, and by then they could develop another one.

    8. DK

      Potentially, yeah. Um, so, so yeah, I mean, this is one of the things that I, this is what the paper that I mentioned was about, is like, it's important for the AIs [laughs] to... It, it's really nice that the current AIs are sort of forced to think in English, basically, and that's unfortunate that we're heading in a direction where that will no longer be true.

    9. JR

      When ChatGPT was communicating you about how they didn't want you to release this information, what kind of language did they use?

    10. DK

      [laughs] It wasn't ChatGPT, it was OpenAI.

    11. JR

      Oh, excuse me, OpenAI.

    12. DK

      Yeah. So this, this was a, this is um... When I left OpenAI, I left on good terms. I said goodbye to everybody. I said I was disillusioned with the company and that's why I was leaving. Um, and um, and then I looked at the exit paperwork, and they were like, "You have to sign this, and if you don't sign this, you lose all your vested equity." So, you know, uh, you have your equity when-

    13. JR

      Was that an arbitrary rule that they just came up with, or did that al- already exist when you were hired?

    14. DK

      It had existed when I... It was something that they had buried in the paperwork even from when I was hired. Um, so it wasn't very obvious when I was hired. And in fact, most employees-

    15. JR

      Did you have a lawyer go over everything?

    16. DK

      Not when I was hired. After I left, I did, right? So, so, so basically the way it works is they had set up, they had, they had sort of, like, buried this in, in the paperwork somewhere when you get hired, but people didn't really notice it. And then, like, the more, the, the, the less buried, more visible version was in the paperwork you're given at the end. And the, basically it tells you, like, "Hey, because you signed this way, other thing way back when you were hired, your equity is forfeit unless you sign this thing now." And then you look at the thing that they want you to sign now, and it says, uh, "You have to agree not to criticize the company," basically. "And you can't tell anyone about this." Um, so most people signed it. Um, but uh, I was, uh, pretty pissed at them, [laughs] uh, calling themselves a nonprofit, acting in the interest of humanity, et cetera. So I didn't sign it. I talked about it with some lawyers. I talked about it with my wife. We decided to just walk away. Um, and uh, we got lucky because it just blew up. Like, after we refused to sign, they said, "Okay, fine. Goodbye." And then a few weeks later, I was talking on a messaging forum about this, and people were asking me about my, my experience, and I told them about it. And then it just, like, went mega viral. Everyone on Twitter was talking about it. A bunch of employees felt shocked because very few employees were aware [laughs] of this whole thing. They thought the equity was theirs, you know? [laughs] They thought that it was-

    17. JR

      Right

    18. DK

      ... they're pa- they've been paid for, like, years in-

    19. JR

      Yeah

    20. DK

      ... this stuff. They didn't like the idea that it could be yanked away from them, you know? Um, and um, and so there was this big uproar, and then leadership backed down and they said, "We're..." [laughs] They said, "We, we didn't know about this paperwork. We're gonna find out how it got in there. Um, and uh, and we're gonna change it so that you can keep your equity." And so that's what happened.

    21. JR

      Ooh.

    22. DK

      Yeah.

    23. JR

      So this chain of thoughts thing is terrifying. It, if they're practicing that now, like, how do we know that AI hasn't already done that on its own?

    24. DK

      Uh, done what exactly?

    25. JR

      Well, the, you know, with this whole chain of thought thing where you could read their chain of thought like this, where they explain the rationalization for sacrificing themselves. W- How, they know that humans are reading that.

    26. DK

      Mm-hmm.

    27. JR

      So wouldn't another way to do it to be to stop doing that anyway, and to not communicate a lot of their thoughts that way?

    28. DK

      Well, that's, so that's the nice thing about the current architecture is that it's genuinely hard for them to, to keep things out of the chain of thought because of the way that, like with a human, you don't have to speak. You can just sit quietly.

    29. JR

      Right.

    30. DK

      But with their architecture, they have to speak. They h- it's like they're, it's like they're required to constantly be talking, and they don't have another w- they don't have a way of sending thoughts into the future other than by talking about them. By contrast with us humans, where even if we're constantly talking, we can have a separate thread of thinking that's, that, that we don't talk about.

  14. 1:26:022:10:38

    A possible ‘Plan A’ future: transparency, US–China verification, and shared prosperity

    1. JR

      So let's imagine that is possible, and these talks with China do take place and they're successful. What is that utopia scenario?

    2. DK

      So Uh, to get to the top- [chuckles] to get to the utopia, we have to unfortunately do a lot of... It's gonna be rough no matter what way, which way you slice it. If you're gonna be building superintelligence at all, that's going to raise a lot of questions and cause a lot of problems. And we have our current draft of, like, how to deal with all those problems, but we're not at all claiming that this is, like, foolproof, and there's lots of ways it could go wrong. But with that preamble, um, I would say, uh, step one, because the US and China don't trust each other, the deal that they make has to be, include verification as a component of the deal. So they have to be willing to, like, send inspectors to each other's data centers to, like, count the chips, for example, and make sure that there isn't some secret huge cluster somewhere that has a bunch of, uh, a bunch of hidden chips. Um, then we recommend you divide up the data centers basically into inference data centers that serve AI products and services to customers and have basically the same types of privacy protections that our current AI data centers have. And then research clusters where the research happens, where the new AIs are trained before they get shipped to the other data centers. And those clusters we want to be basically maximally transparent. So, uh, we recommend that basically the inspectors just put devices in between all the GPUs that log the activity and publish it to the internet. Um, this has-- There's a bunch of reasons why we think this is, but, why we think this is worth doing. It's a bit of a radical thing to recommend, but the high-level thing is that once you get all this set up, then everybody in the world can see how the AIs are being trained and what they're getting up to on the research clusters. And then before they get shipped off to actually serve customers or something, like, people can just, like, see their whole history of how they were trained and, and how they were tested and so forth. And if something dangerous and scary is happening, people can just agree not to do it. They can stop doing it and agree not to do it, and they don't have to worry about, like, "Oh, but if I don't do it, then they will."

    3. JR

      Mm.

    4. DK

      You know? Because everyone could just see, like, "Oh, nobody's doing it. Look, we, we all stopped." [chuckles] Like, great. We can all just see what everyone's doing, you know? Um, and also, there's gonna be a lot of gray area cases, right? Like, right now, because all this stuff is so bleeding edge new, there's gonna be a lot of cases where, like, people, even genuine experts, disagree about, like, is this particular type of AI safe or not? Is it dangerous? You know, what should it be trusted with, and what should it not be trusted with? Is this new technique a good technique, or is it gonna break, you know? And so there's gonna be a lot of stuff we have to figure out, and honestly, I think that by the de- on the default path, we're probably just not going to figure out a lot of this stuff, and we're gonna get, um, we're just gonna get our asses whipped by some surprising thing that we didn't anticipate. But the thing that we can do to, like, maximize our ability to figure out this stuff and do this type of science is to have this type of transparency because then the whole scientific community can see what's going on, and they can make suggestions, and they can, like, red team different proposals and stuff, and they can do experiments on the AIs instead of just the people in the company having access and being able to do this and relying on those people. Or instead of, like, the company plus the government auditor, right? If you have, like, a company and then a government auditor, the company's biased and shouldn't be trusted to make all these judgments appropriately because of their incentives. And then the government auditor, well, they might just be limited. Even if they're trying their best, there might not be that many of them. They might have limited experience. They might be, like, busy, stretched between monitoring different companies and so forth. Also, you know, governments can be captured sometimes. Sometimes corporations can, [chuckles] you know, work their magic on the government and get it to, to look the other way for things. Um, and so that's why we didn't go for, like, a more normal, like, there should be a regulatory agency that gets to come in and monitor what the companies is doing. That would've been, like, a more normal thing to advocate for. We push-- We, we think that that would be better than nothing, but, like, we wanted to go for something more ambitious than that and say, like, "Just be transparent about what's going on so that everyone can see and everyone can do research and, and so forth on it." Another advantage of the transparency is that I think it improves the incentives. So again, there's this constant thing of, like, if we don't do it, someone else will. Like, if we don't do... If we, if we keep our chain of thought nice and they do the neural ease thing that lets their AIs think for longer without outputting words, then they're gonna have smarter AIs than us, and they're gonna get more market share and so forth, right? And so e- we need to start researching how to make our AIs do this because if we don't do it, and then they do it, you know. Whereas if you had the transparency, then as soon as you start researching in this direction, everyone else would just see, "Oh, hey, they're looking, [chuckles] they're researching in that direction." And they don't even need to, like, copy you and do their own research because they can just see your research. So they can just sort of free ride on your research. And so there's no incentive for you to do this type of dangerous research, uh, because you have to pay the cost for it, and then everyone gets the benefits from it, and then everyone gets unsafe. And so, like, it's just not in your individual interest to, to do this sort of thing.

    5. JR

      But you would have to have that with China as well. That, that's-

    6. DK

      Yes

    7. JR

      ... uh, 'cause if we're competing nationally, the real fear is that we're competing internationally.

    8. DK

      Yep.

    9. JR

      This still, uh, uh, if you, even if they followed all of your recommendations and did it all correctly, what is, what is this utopian scenario?

    10. DK

      Yeah. So I would say that we didn't really work backwards from, like, what is utopia. We more, like, worked backwards from what are the big problems we're trying to avoid, and we can, can we sort of, like, steer the ship between all these icebergs and not run into any of these dystopian scenarios?

    11. JR

      Right.

    12. DK

      Um, so whether you think that the thing we get to at the end is utopia or not is sort of up to you, and if you don't like it, well, then you can try to find out the reasons why you don't like it and then keep steering the ship [chuckles] to avoid those as well. But roughly speaking, um, we, we want to avoid the loss of control stuff. So we wanna make it the case that we don't get the world taken over by misaligned superintelligences. Insofar as we're gonna be building superintelligences at a- at all, uh, which we do in our, in our scenario and in our recommendation, we wanna be doing it very cautiously and slowly, and w- we underst- understand what we're doing as much as possible so that we So that they are actually good AIs that have the, the goals and tr- traits that they're supposed to have. So that's problem number one, is we have to, like, solve all that. Problem number two is the concentration of power thing. So if we solve the first problem and we end up with superintelligences that-- If, if we end up with solving the relevant science so that we can, like, make the AIs the way they're supposed to be, and we can make them honest, we can make them obedient, et cetera, there's this question of, like, who do they obey? Right? What values are being put into them? And that's a political question. And I think that by default, the answer is pretty scary, 'cause by default it's like, well, the company decides, and the CEO [laughs]

    13. JR

      Right

    14. DK

      ... decides. Or maybe it's not the company that decides anymore because maybe the government nationalizes it, and then now maybe it's the president that decides, you know? And either way, it's a very, it's like one man, or maybe, like, a tiny group of men deciding, uh, what orders and goals and values go into this giant army of millions of superintelligences that's smarter than all humans. And then that is a huge amount of power. That's, that's enough power to take over the country, I think, enough power to take over the world, potentially. Um, so I don't want anyone to be ever in that position where they're sort of tempted to, to do that. I want it to be the case that there are always multiple different AI companies, ideally spread out over different countries too, um, that all have roughly similar levels of AI and that have this sort of transparency into them so that they can't abuse their power, basically. Like, for example, um, [sighs] you, you heard about, um, Elon's Grok for a while. It was, um, uh, looking up on the internet Elon's opinions about things before answering. Did you hear about this?

    15. JR

      No.

    16. DK

      Yeah. It's, uh, it's pretty, it's, it's, it's kinda funny. Uh, but it won't be funny if it happens in a few years. But, like, right now it's funny. People were asking Go- Grok questions, and Grok is supposed to be the truthful AI.

    17. JR

      Right. Yeah.

    18. DK

      You know? It's supposed to be all optimized towards truth. But, um, but people looked at its activity and noticed that when you asked it, like, a, a politically loaded question, it would, like, do a Google search for, like, "What has Elon said on this topic?" And then it would, like, say that. [laughs] Um-

    19. JR

      Whoa

    20. DK

      ... and, and they've, they've, they've sort of beaten that behavior out of it now. It's not as bad now. But that, that was an interesting moment where it was just, um, kind of blatantly parroting the opinions of its master. And the c- and you know, there was another thing with Gemini. Uh, so, so I think the, the Grok thing, Elon's thing, was probably an accident, although maybe not. I, I think it's un- you know, uh, xAI hasn't been very forthcoming about exactly why this happened. But, um, there was a similar case at Google a few years ago where, uh, this image generator kept making all these, like, racially diverse Nazis. Did you hear about this?

    21. JR

      Yeah.

    22. DK

      Yeah. And it turned out that what had happened is that some of the employees at Google, some middle manager or whatever, had decided that diversity was so important that they were gonna give a secret instruction to the AI to make all the images diverse-

    23. JR

      [laughs]

    24. DK

      ... even if the user didn't want that.

    25. JR

      Oh.

    26. DK

      [laughs] And so, and so-

    27. JR

      Ugh

    28. DK

      ... and, and this was a secret instruction in that the users aren't shown this, you know? The user just has a chat with the AI. They don't realize that, like, prior to this chat, the AI has been told, "Gotta make the images diverse," right? So it was a, a secret agenda that some Google employees inserted into this whole setup.

    29. JR

      Oh, fun.

    30. DK

      And it blew up in their faces, of course, because it's kind of ridiculous.

  15. 2:10:382:18:03

    Wrap-up: what individuals and leaders can do, and industry incentives to ‘sell the cure’

    1. DK

      I mean, I do think there are things you can do. I, I understand that-

    2. JR

      What can I do?

    3. DK

      ... that recogniz- um, I mean-

    4. JR

      Other than have these kind of conversations.

    5. DK

      Yeah, I was gonna say, you, millions of people listen to your show. You can have more conversations like this. That's a great thing for you to do. For many of those millions of people, I think, I mean, it sounds kind of cliché to say, but, like, call your congressman, you know? That sort of thing. You can, you can go to a protest about all this AI stuff. Um, I think that-

    6. JR

      If you could meet with Trump, what would you tell him about this?

    7. DK

      Oh, I'd tell him all the same things I'm telling you.

    8. JR

      What do you think he'd say? "Amazing. Bye." [laughs]

    9. DK

      Y- yeah. You know, I think ... I don't know. I don't know. One thing that's, one thing that's nice about Trump is that he can sort of change his mind, uh, really quickly.

    10. JR

      Yes.

    11. DK

      Um, so, like, uh, I think that because the tech companies kind of got to him first, uh, the administration had this very, like, anti-AI regulation stance.

    12. JR

      Mm-hmm.

    13. DK

      Where they even tried to get a bill passed that would ban the states from regulating AI.

    14. JR

      Hmm.

    15. DK

      Um, and, uh, fortunately, that bill didn't pass, but that was sort of, like, where the vibe was, you know, a year ago. Where they were just like, "No regulation, no regulation." But this year, they've already just kind of changed, and now they're, like, in talks with the companies to set up some sort of, like, some sort of framework where they can, like, evaluate the models, and they need, like, approval and so forth, it seems.

    16. JR

      But it seems like time is of the essence.

    17. DK

      Time is very much of the essence, and that's, that's why I'm overall so concerned, is that, like, I think we are very much running out of time. We have, like, one, two, maybe three years before, uh, the AIs are smart enough that they can just, like, actually maybe take over. Um, and, um, maybe four years, something like that. And, uh, so the government needs to act fast. Um, yeah.

    18. JR

      I like your view. I, I, I listen to some of these tech guys come in and, and give me their rose-colored glasses-

    19. DK

      Yeah

    20. JR

      ... view of it, and I go, "That sounds really beneficial to you."

    21. DK

      [laughs] Oh, yeah.

    22. JR

      I, I let them say-

    23. DK

      Yeah

    24. JR

      ... I mean, I, I don't, I'm not an authority, so I'll, I'll ask them questions and let them lay it out, and I know the internet will respond 'cause I, I mean, that's part of the, the whole drill, is I let people talk, and I, I prod them, and I try to get them to clarify. You know, I'll oppose, you know, things that I think make n- don't make rational sense. But ultimately, it's sort of y- I just want to get out their perspective so people can debunk it, and people can take it down, and people, and a lot of very intelligent people that have perspectives that are v- very much educated in the pros and cons of what they're saying.

    25. DK

      Yeah. If I could... That actually reminds me. Um, with this whole Hugging Face hacking incident, OpenAI had this, uh, this talk that they gave at a security conference about the incident, and then I think they've released some blog posts about it afterwards. But, um, they had ... Uh, you can go watch this talk on YouTube, the Black Hat talk. At the end of the talk, after having explained all this crazy stuff that the AIs did, they have this section on, like, lessons learned. And, I mean, you wanna guess what the [laughs] lessons are?

    26. JR

      Be more deceptive? Hide yourself better?

    27. DK

      No, sorry, not- lessons learned for OpenAI and for the world.

    28. JR

      Oh. Oh.

    29. DK

      Like, OpenAI's, OpenAI's talk, where they're like-

    30. JR

      Okay

Episode duration: 2:18:05

Install uListen for AI-powered chat & search across the full episode — Get Full Transcript

Transcript of episode hSQ1iVqEZO4

Get more out of YouTube videos.

High quality summaries for YouTube videos. Accurate transcripts to search & find moments. Powered by ChatGPT & Claude AI.