Lex Fridman PodcastTuomas Sandholm: Poker and Game Theory | Lex Fridman Podcast #12
EVERY SPOKEN WORD
115 min read · 23,353 words- 0:00 – 4:10
Heads-up No-Limit Texas Hold’em as an AI benchmark (rules + why it’s hard)
- LFLex Fridman
The following is a conversation with Tuomas Sandholm. He's a professor at CMU and co-creator of Libratus, which is the first AI system to beat top human players in the game of Heads-up No-Limit Texas Hold'Em. He has published over 450 papers on game theory and machine learning, including a best paper in 2017 at NIPS, now renamed to NeurIPS, which is where I caught up with him for this conversation. His research and companies have had wide-reaching impact in the real world, especially because he and his group not only propose new ideas, but also build systems to prove that these ideas work in the real world. This conversation is part of the MIT course on Artificial General Intelligence and the Artificial Intelligence podcast. If you enjoy it, subscribe on YouTube, iTunes, or simply connect with me on Twitter @lexfrid. And now here's my conversation with Tuomas Sandholm. Can you describe, at the high level, the game of poker, Texas Hold'Em, Heads-up Texas Hold'Em, for people who might not be familiar, uh, this card game?
- TSTuomas Sandholm
Yeah, happy to. So Heads-up No-Limit Texas Hold'Em has really emerged in the AI community as a main benchmark for testing these application-independent algorithms for imperfect information game solving. And this is a game, uh, that's actually played by humans. You don't see it that much on TV or casinos, because, uh, well, for various reasons, but, uh, uh, you do see it in some expert-level casinos, and you see it in the best poker movies of all time. It's actually an event in the World Series of Poker. But mostly it's played online and typically for pretty, uh, big sums of money. And this is a game that usually only experts play. So if you re- uh, go to your home game on a Friday night, it probably is not gonna be Heads-up No-Limit Texas Hold'Em. It might be, uh, No-Limit Texas Hold'Em in some cases, but typically for a, a big group when it's not as competitive. While heads-up means it's two players, so it's really like me against you. Am I better or are you better? Much like chess or, or, or Go in that sense, but an imperfect-information game, which makes it much harder, of course. I have to deal with issues of, uh, you knowing things that I don't know, and I know things that you don't know, instead of pieces being nicely laid on the board for both of us to see.
- LFLex Fridman
So in Texas Hold'Em, uh, there's, uh, two cards that you only see? The-
- TSTuomas Sandholm
Yes.
- LFLex Fridman
... they belong to you?
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
And then there is they gradually lay out some cards that add up overall to five cards that everybody can see.
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
So the imperfect nature of the information is the two cards that you're holding in your hand.
- TSTuomas Sandholm
Up front, yeah. So as you said, you know, you first get two cards in private each, and then you, uh, there's a betting round. Then you get three cards in public on the table, then there's a betting round. Then you get the fourth card in public on the table, there's a betting round.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Then you get the five- fifth card on the table, there's a betting round. So there's a total of four betting rounds and four tranches of information revelation, if you will. The only- the first tranche is private, and then it's public from there.
- LFLex Fridman
And this is probably, probably by far the most popular game in AI and, um, just the general public in terms of imperfect information. So it's probably the most popular spectator game to watch, right? So, uh, th- which is why it's a super exciting game to tackle. So it's, it's on the order of chess, I would say, in terms of popularity, in terms of AI setting it as the bar of what is intelligence. So in 2017, Libratus... How do you pronounce it? Libratus.
- TSTuomas Sandholm
Libratus.
- LFLex Fridman
Libratus. Libratus beats...
- TSTuomas Sandholm
Little Latin there.
- LFLex Fridman
A little bit of Latin. Uh, Libratus beat a few, uh, four expert human players. Can you describe that event? What you learned from it? What was it like? What was the process in general for people who have not read the papers and, uh-
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
... studied?
- 4:10 – 7:08
The Libratus vs. top pros match: setup, incentives, and interface
- TSTuomas Sandholm
Yeah, so the event was that, uh, we invited four of the top ten players. With these, uh, specialist players in Heads-up No-Limit Texas Hold'Em, which is very important, because this game is actually quite different than the, the multiplayer version. We brought them in to Pittsburgh to play at the Rivers Casino, uh, for 20 days. We wanted to get to 120,000 hands in, because, uh, we wanted to get statistical significance. Uh, so it's a lot of hands for humans to play, even for these top pros who play fairly quickly normally. Uh, so we, uh, couldn't just have one of them play so many hands. 20 days, they were playing basically morning to evening, and, um, I raised $200,000 as a little incentive for them to play. And w- the setting was so that they didn't all get $50,000. Um, we actually paid them out based on how they did against the AI each.
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
So they had an incentive to play, uh, as hard as they could, whether they're way ahead or way behind or right at the mark of beating the AI.
- LFLex Fridman
And you don't make any money, unfortunately.
- TSTuomas Sandholm
Right, no.
- LFLex Fridman
(laughs)
- TSTuomas Sandholm
We can't make any money, so, so originally, a couple of years earlier, I, uh, well, uh, I actually explored whether we could actually play for money, because that would be, of course, uh, interesting as well, uh, to play against the top people for money, but the Pennsylvania Gaming Board said no. So, so we, we couldn't. So this is much like an exhibit, like, uh, like for a musician or a boxer or something like that.
- LFLex Fridman
Nevertheless, you were keeping track of the money and Li- Li- Libratus, uh, won close to $2 million, I think. Uh, so (laughs) so if that- if it was for real money, uh, if you were able to earn money, that was a quite impressive and in- inspiring achievement. Just, uh, a few details. What, what were the players looking at? I mean, uh, were they behind a computer? What, what was the interface like?
- TSTuomas Sandholm
Yes, uh, they were playing much like they normally do. These top players, when they play this game, they play mostly online, so they're used to playing through a UI.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
And they did the same thing here. So there was this layout. You could imagine, there's a table-
- LFLex Fridman
Yes.
- TSTuomas Sandholm
... uh, on a screen. There's...... the, the, the human sitting there, and then there's the AI sitting there, and the, the, uh, the screen shows everything that's happening. The cards coming out, and shows the bets being made. And we also had the betting history for the human, so if the human forgot what had happened in the hand so far, they could actually reference back and, and, and, and so forth.
- LFLex Fridman
Is there a reason they were given access to the betting history for-
- TSTuomas Sandholm
Well, we, we just, uh, uh-
- LFLex Fridman
... some kind of ...
- TSTuomas Sandholm
... it, it's a, it, it didn't really matter-
- LFLex Fridman
Okay.
- TSTuomas Sandholm
... that they wouldn't have forgotten anyway. These are top-quality people. But, uh, the, uh, we just wanted to put out there, so it's not a question of the human forgetting and the AI somehow trying to get advantage of better memory.
- LFLex Fridman
So, what was that like? I mean, that was an incredible accomplishment. So, what did it feel like before the event? Did you have doubt? Hope? Or what, where was your confidence at?
- 7:08 – 10:24
Confidence, betting markets, and the myth of “poker tells”
- TSTuomas Sandholm
Yeah, that's great. So, uh, great question. So, uh, 18 months earlier, I had organized a similar brains-versus-AI competition-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... with a previous AI called Cloudico, and we couldn't beat the humans. Uh, so, uh, this time around, it was only 18 months later, and I knew that this new AI, Libratus, was way stronger, but it's hard to say how you'll do against the top humans before you try. So, I thought we had about a 50/50 shot. And the international betting sites put us as a, us as a four-to-one or five-to-one underdog.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So it is kinda interesting that people really believe in people and over AI. Uh, not just, pe- people don't just believe, over-believe in themselves, but they have overconfidence in other people as well, compared to the performance of AI. And, uh, yeah, so we were a four-to-one or five-to-one underdog, and even after three days of beating the humans in a row, we were still 50/50 on the international betting sites.
- LFLex Fridman
Do you think there's something special and magical about poker in, in the way people think about it? In a sense, you have, I mean, even in chess, there's no Hollywood movies. Poker is, uh, the s- the star of many movies, and there's this feeling that, uh, certain human em- facial expressions and body language, eye movement, all these tells are critical to poker. Like, you can look into somebody's soul and understand-
- TSTuomas Sandholm
(laughs)
- LFLex Fridman
... their betting strategy and so on. And so, that's probably why, p- possibly, do you think that is why people have a confidence that humans will outper- because AI systems cannot, in this construct, perceive these kinds of tells? They're only looking at betting patterns and, uh, th- and nothing else. (laughs) The betting patterns and, and statistics. So, th- what's more important to you if you step back on human players, human-versus-human? What's the role of these tells, of these, uh, ideas that we romanticize?
- TSTuomas Sandholm
Yeah. So, I, I'll, I'll split it into two parts. So, one is, why do humans trust humans more than AI-
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
... and o- have overconfidence in humans?
- LFLex Fridman
Yes.
- TSTuomas Sandholm
I think that's, that's not really related to, to the tell question.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
It's just that, uh, they've seen these top players, how good they are, and they're really fantastic. So, it's just hard to believe-
- LFLex Fridman
That you can beat them.
- TSTuomas Sandholm
... yeah, or that an AI could beat them.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
So I think that's where that comes from. And, and that's actually maybe a more general lesson about AI, that un- until you've seen it over-perform a human, it's hard to believe that it could. But, um, then the tells, um, a lot of these top players, they're so good at hiding tells that among the top players, it's actually not really-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... worth it for them to invest a lot of effort trying to find tells in each other because they're so, so good at hiding them. So, uh, yes, at the kind of Friday evening game, tells are gonna be a huge thing.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
You can read other people and if you're a good reader, you'll, you'll read them like an open book.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
But at the top levels of poker now, the tells become a less, a much, much smaller and smaller aspect of the game as you go to the top levels.
- 10:24 – 12:59
Scaling to 10^161: information vs action abstraction (and abstraction pitfalls)
- LFLex Fridman
The, the amount of strategies, the amounts of possible actions is, um, is very large, uh, 10 to the power of 100 plus, uh, so there has to be some... I've read a few of the papers related, um, there has, it has to form some abstractions of various hands and actions. So, what kind of abstractions are effective for the game of poker?
- TSTuomas Sandholm
Yeah. So, you're exactly right. So, uh, when you go from a game tree that's 10 to the 161-
- LFLex Fridman
161.
- TSTuomas Sandholm
... especially in an imperfect information game, it's way too large to solve directly, even with, with our fastest equilibrium-finding algorithms. So, uh, you wanna abstract it first. And abstraction in games is much trickier than abstraction in MDPs or other single-agent settings, 'cause you have these abstraction pathologies, that if I have a finer-grained abstraction-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... the strategy that I can get from that for the real game might actually be worse than the strategy I can get from the coarse-grained abstraction. So you have to be very careful.
- LFLex Fridman
Now, the, the kinds of abstractions, just to zoom out, we're talking about, there's the ha- hands abstractions and then there is betting strategies. What, what-
- TSTuomas Sandholm
Yeah. Betting actions, yeah.
- LFLex Fridman
B- betting actions?
- TSTuomas Sandholm
So, so, there's information abstraction to talk about general games. Information abstraction, which is the abstraction of what chance does. And this would be the cards in the case of poker.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And then there's action abstraction, which is abstracting the actions of the actual players-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... which would be bets in the case of poker.
- LFLex Fridman
Yourself and the other players?
- TSTuomas Sandholm
Yes. Yourself and the other players. And, uh, for information abstraction, we were completely automated. So these wa- these are algorithms, uh, but they do what we call potential-aware abstraction where we don't just look at the valley of the hand, but also how it might materialize into good or bad hands over time. And it's a certain kind of bottom-up process, uh, with integer programming there and clustering and various aspects, how do you build, build this abstraction.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And then, in the action abstraction, there, it's largely based on how, um...... humans, other un- and other AIs have played this game in the past, but in the beginning, we actually used an automated action abstraction technology, which is provably convergent, that it finds optimal combination of bet sizes. But it's not very scalable, so we couldn't use it for the whole game, but we used it for the first couple of betting actions.
- 12:59 – 14:27
Luck, variance, and why you need 100,000+ hands
- LFLex Fridman
So what's more important, uh, the strength of the hand, so the, the information abstraction, or the, uh, how you play them? Uh, the actions. Does it... You know, the romanticized notion, again, is that it doesn't matter what hands you have, that the actions, the betting, uh, may be the way you win, no matter what hands you have.
- TSTuomas Sandholm
Yeah, so that's why you have to play a lot of hands, so that the role of luck gets smaller. So, uh, you c- you could otherwise get lucky and get some good hands, and then you're gonna win the match. Even with thousands of hands, you can get lucky, uh, because there's so much variance in no-limit Texas hold 'em, because if we both go all in, it's a huge stack, uh, of variance. So there are these massive swings in, uh, no-limit Texas hold 'em.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So that's why you have to play not just thousands, but, uh, over 100,000 hands to get statistical significance.
- LFLex Fridman
So, so let me, let me ask another way this question. If you didn't even look at your hands, but they didn't know that, your opponents didn't know that, uh, how well would you be able to do?
- TSTuomas Sandholm
Uh, that's a good question. There's actually... I heard this story that there's this Norwegian female poker player called Anette Oberstad-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... who's actually won a tournament by doing exactly that. But that would be extremely rare. So, so I, uh, I... Y- y- y- you can-
- LFLex Fridman
(laughs)
- TSTuomas Sandholm
... you cannot really play well that way. (laughs)
- LFLex Fridman
Oh, okay. So, (laughs) so the, the hands do have some role to play, okay.
- TSTuomas Sandholm
Yes.
- 14:27 – 18:57
Learning vs game-theoretic solving: why imperfect information complicates value functions
- LFLex Fridman
So, Libratus does not, uh, use... As, as far as I understand, uh, use learning methods, uh, deep learning. Is there room for learning in, uh... You know, there, there's no reason why Libratus doesn't, you know, combine with an AlphaGo-type approach for estimating the quality, uh, for a function estimator. What are your thoughts on this, maybe as compared to another algorithm which I'm not that familiar with, DeepStack, the, the engine that does use deep learning, that it's unclear how well it does, but nevertheless uses deep learning? So what are your thoughts about learning methods to aid in the way that, uh, Libratus plays the game of poker?
- TSTuomas Sandholm
Yeah, so as you said, Libratus did not use learning methods, and, um, played very well without them. Since then, we have actually...
- LFLex Fridman
Awesome.
- TSTuomas Sandholm
Actually here, we have, uh, a couple of papers on things that do use learning techniques.
- LFLex Fridman
Excellent.
- TSTuomas Sandholm
Uh, so, um, and, and deep learning in particular. And th- uh, sort of the way you're talking about, where it's learning an evaluation function.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
But, um, in imperfect information games, unlike, let's say, in Go, or now, now also in chess and shogi-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... it's not, um, sufficient to learn an evaluation for a state, because the value of an information set, uh, depends not only on the exact state, but it also depends on both players' beliefs.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Like if I have a bad hand, I'm much better off if the opponent thinks I'm, I have a good hand, and vice versa. If I have a good hand, I'm much better off if the opponent believes I have a bad hand.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Uh, so the value of a state is not just a function of the cards. Uh, it depends on, uh, if you will, the path of play, but only to the extent that it's captured in the belief distributions. So, uh, so that's why it's not as simple as, uh, uh, as it is in perfect information games. And I don't wanna say it's simple there either. It's, of course-
- LFLex Fridman
Right.
- TSTuomas Sandholm
... very complicated computationally there, too. But at least conceptually it's very straightforward. There's a state, there's an evaluation function, you can try to learn it. Here, you have to do something more. Uh, uh, and, uh, w- what we do is in one of these papers we're looking at allowing... where, where we allow the opponent to actually take different strategies-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... at the leaf of the search tree, as, as... if you will. And, and that, uh, is a different way of doing it, and it doesn't assume, therefore, a particular way that the opponent plays. But it allows opponent to choose from a s- uh, set of different continuation strategies, and that forces us to not be too optimistic in a lookahead search. And that's o- that's one way you can do sound lookahead search in imperfect information games, which is very diff- difficult. And in, in, in... You asked, you were asking about DeepStack.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
What they did, uh, it was very different than what we do, either in Libratus or in this new work. They were gen- randomly generating various situations in the game, then they were doing lookahead from there to the end of the game as if that was the start of a different game.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And then they were using deep learning to learn those, uh, values of those states, but the states were not just the physical states. They include the belief distributions.
- LFLex Fridman
When you talk about lookahead, uh, for DeepStack or with Libratus, does it mean considering every possibility that the game can involve? Is... Are we talking about extremely sort of ex- this exponentially growth of a tree?
- TSTuomas Sandholm
Yes. So we, we, we're talking about exactly that, much like you do in alpha-beta search or Monte Carlo tree search, but with different techniques. So there's a different search algorithm, and then we have to deal with the leaves differently. So if you think about what Libratus did, we didn't have to worry about this, because we only did it at the end of the game, so we would always terminate into a real situation, and we would know what the payout is. It didn't do these depth-limited lookaheads. But now in this new paper, uh-... which is called depth-limited, I think it's called depth-limited search for imperfect information games. We can actually do sound depth-limited lookahead. So we can actually start to do the lookahead from the beginning of the game on because that's too complicated to do for this whole long game. So when we brought this, we were just doing it for the end.
- 18:57 – 22:43
Beliefs, Bayes, and Nash equilibrium: ‘beliefs are output, not input’
- LFLex Fridman
So, and then the other side, this belief distribution, so is it explicitly modeled what kind of beliefs that the opponent might have?
- TSTuomas Sandholm
Yeah. Yeah, it is explicitly modeled, but it's not assumed. The beliefs are actually output, not input. Of course, the starting beliefs are input, but they just fall from the rules of the game, because we know that the dealer deals uniformly from the deck.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So I know that every pair of cards that you might have is equally likely.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
I know that for a fact. That just follows from the rules of the game. Of course, except the two cards that I have. I know you don't have those.
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
Uh, you have to take that into account. That's called card removal, and that's very important.
- LFLex Fridman
Is, is the dealing always co- coming from a single deck in heads-up?
- TSTuomas Sandholm
Yes.
- LFLex Fridman
So you can assume-
- TSTuomas Sandholm
Single deck. That's called card removal. It's- So you, you, you know that if some, uh, if, if I have the ace of spades, I know you don't have-
- LFLex Fridman
Right.
- TSTuomas Sandholm
... an ace of spades.
- LFLex Fridman
Okay, great. So in the beginning, your belief is basically the fact that it's a fair dealing of hands, but how do you adjust, start to adjust that belief?
- TSTuomas Sandholm
Well, that's, uh, where this beauty of game theory comes. So Nash equilibrium, which John Nash introduced in 1950-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... introduces what rational play is when you have more than one player. And these are pairs of strategies where strategies are contingency plans, one for each player, uh, so that neither player wants to deviate to a different strategy-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... given that the other doesn't deviate. But as a side effect, you get the beliefs from Bayes' rule. So Nash equilibrium really isn't just deriving... In these imperfect information games, Nash equilibrium d- doesn't just define strategies, it, it also defines beliefs for both of us, and it de- refines beliefs for each state. So at each state, uh, each... If they take all information sets.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
At each information set in the game, there's a set of different states that we might be in, but w- I don't know which one we're in.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Nash equilibrium tells me exactly what is the probability distribution over those real-world states in my mind.
- LFLex Fridman
How does Nash equilibrium give you that distribution? So why...
- TSTuomas Sandholm
So I'll, I'll do a simple example.
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
So, you know the game rock, piece, p- uh, paper, scissors? So we c- we can draw it as player one moves first and then player two moves, but of course, it's important that player two doesn't know what player one moved.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Otherwise, player two would win every time. So we can draw that as an information set, where player one makes one or three moves first, and then there's an information set for player two.
- 22:43 – 25:13
Opponent exploitation: hybridizing equilibrium play with data (and why Libratus avoided it)
- TSTuomas Sandholm
I mean, we've done some work on combining opponent modeling with game theory so you can exploit weak players even more. But that's another strand, and in Libratus we didn't turn that on 'cause I decided that these players are too good. And when you start to exploit an opponent, you typically open yourselves up, self up to exploitation. And these guys have so few holes to exploit, and they're world's leading experts in counter-exploitation. So I decided that we're not gonna turn that stuff on.
- LFLex Fridman
Actually, I saw a few of your papers about exploiting opponents. It sound very interesting to, uh, explore. Uh, do you think there's room for exploitation g- generally outside of libratus? Is, is there a subject, uh, or people differences that could be exploited? Maybe not just in poker, but in general interactions, negotiations, all these other domains that you're considering?
- TSTuomas Sandholm
Yeah, definitely. We've done some work on that, and I really like the work that hybridizes the two.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So you figure out what would a rational opponent do. And by the way, that's safe in these zero-sum games, two-player zero-sum games, because if the opponent does something irrational, yes, it might show, uh throw off my beliefs-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... um, uh, but the amount that the player can gain by throwing o- off my belief is always less than they lose by playing poorly.
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
So, so it's safe. But, uh, still, if somebody's weak as a player, you might wanna play differently to exploit them more.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So that you can think about it this way. A game theoretic strategy is un- unbeatable, but it doesn't maximally beat the other opponents. So the winnings per hand might be better with a different strategy. And the hybrid is that you start from a game theoretic approach, and then as you gain data from, uh, th- about the opponent in certain parts of the game tree, then in those parts of the game tree, you start to tweak your strategy more and more-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... uh, towards exploitation while still staying fairly close to the game theoretic strategy so as to not open yourself up to exploitation too much.
- LFLex Fridman
... h- how, how do you do that? Do you, um, try to vary up strategies, make it unpredictable? It's like, uh, uh, what is it, uh, tit-for-tat strategies and, um, prisoner's dilemma or...
- TSTuomas Sandholm
Well, it doesn't r- that, that's a repeated game, kind of-
- LFLex Fridman
Repeated games. That's right.
- TSTuomas Sandholm
... s- s- simple prisoner's dilemma, repeated games.
- LFLex Fridman
Games, yeah.
- TSTuomas Sandholm
But, but even there, there is no proof that says that that's the best thing. But experimentally, it actually does...
- LFLex Fridman
It does well.
- TSTuomas Sandholm
Does well.
- 25:13 – 30:22
Taxonomy of games and why multiplayer/general-sum is a major leap
- LFLex Fridman
So what kind of games are there, first of all? Y- I don't know if this is something that you could just summarize. There's perfect information games where all the information is on the table. There is imperfect information games. There is repeated games that you play over and over. Uh, there's f- uh, zero sum games.
- TSTuomas Sandholm
Mm-hmm.
- LFLex Fridman
Uh, there's non-zero sum games.
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
(laughs) And then there's a really important distinction you're making, two-player versus more players.
- TSTuomas Sandholm
Mm-hmm.
- LFLex Fridman
So what are... O- what other games are there, and what's the difference, for example, with this two-player game versus more players?
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
What, what are the key differences in your view?
- TSTuomas Sandholm
Right. So let me start from the, the basics. So a repeated game is a game where the same exact game is played over and over. In these extensive form games, uh, whereas... K- think about tree form, maybe with these information sets to represent incomplete information. You can have kind of repetitive interactions. E- even repeated games are a special case of that, by the way.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
But, uh, i- the game doesn't have to be exactly the same. It's like in sourcing auctions. Yes, we're gonna see the same supply base year to year, but what I'm buying is a little different every time, and the supply base is a little different every time, and so on. So it's not really repeated.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So to find a purely repeated game is actually very rare in the world. So they're really a very, uh, coarse model of what's going on.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Then if you move up from r- uh, just repeated, simple repeated matrix games, uh, not all the way to extensive form games, but in between, they're stochastic games, where, you know, there's these, uh... Y- think about it like li- uh, these little matrix games.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And when you take an action and your opponent takes an action, they determine not which next state I'm going to, next game-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... I'm going to, but the distribution of our next games where I might be going to.
- LFLex Fridman
Interesting.
- TSTuomas Sandholm
So, so that's the stochastic game. But it's like matrix games repeated, stochastic games, extensive form games.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
That is, uh, from less to more general. And, and, uh, poker is an example of the last one, so it's really in the most general setting. Uh, extensive form games, and that's kind of what the AI community has been working on and being benchmarked on with this heads-up no-limit Texas hold 'em.
- LFLex Fridman
Can you describe extensive form games? What's, what's the motto here? So, uh, how, how do you describe-
- TSTuomas Sandholm
Yeah, so if you, if you're familiar with the, the tree form... So it's really the tree form. Like in chess, there's a search tree.
- LFLex Fridman
Versus a matrix that-
- TSTuomas Sandholm
Versus a matrix, yeah. And, uh, that's... The matrix is called the matrix form or bi-matrix form or normal form game.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And here, you have the tree form, so you can actually do certain types of reasoning there that you lose the information when you go to normal form.
- 30:22 – 32:45
Collusion and coordination: why cooperative poker variants explode in difficulty
- LFLex Fridman
So, uh, No Brown, the s- student you worked with on this, has mentioned... Uh, I looked through the AMA on Reddit. He mentioned that the ability of poker players to collaborate would make the game... He was asked the question of how would you make the game of poker... Uh, or both of you were asked the question, how would you make the game of poker, uh, beyond being solvable by current AI methods? And he said that, uh, there's not many (laughs)
- TSTuomas Sandholm
Yes.
- LFLex Fridman
... ways of making poker more difficult, uh-... but a collaboration or cooperation between players, uh, would make it extremely difficult. So can you provide the intuition behind why that is, uh, if you agree with that idea?
- TSTuomas Sandholm
Yeah, yeah. So, uh, we've d- uh, I've done a lot of work on, uh, coalitional games.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And we actually have a paper here with, uh, my other student, Gabriele Farina and some other collaborators on, uh, at the, at NIPS on that.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
Actually just came back from the poster session where we presented this.
- LFLex Fridman
(laughs)
- TSTuomas Sandholm
But, uh, so, uh, w- when you have collusion, it's a, it's a different problem.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
And it typically gets even harder then. Even the game representations, some of the game representations don't really allow good computation, so we actually introduced a new game representation for, for that.
- LFLex Fridman
Is that kinda cooperation part of the model? Is, are you, do you have w- uh, do you have information about the fact that other players are cooperating or is it just this chaos that where nothing is known?
- TSTuomas Sandholm
So, so there's some- some things unknown.
- LFLex Fridman
Can you give an example of a, a collusion type game? Or, or is it usually-
- TSTuomas Sandholm
So like bridge.
- LFLex Fridman
Right.
- TSTuomas Sandholm
Yeah. So think about bridge. It's like when you and I are on a team, our payoffs are the same.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
The problem is that we can't talk. So, so when I get my cards, I can't whisper to you what my cards are.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
That would not be allowed.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So, uh, we have to somehow coordinate our strategies ahead of time, and only ahead of time, and then there are certain signals we can talk about-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... but they have to be such that the other team also understands them.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So, uh, so that's, that's, that's an ex- example where the coordination is already built into the rules of the game.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
But in many other situations like auctions, or negotiations, or diplomatic relationships, poker, it's not really built in, but it still can be very helpful for the colluders.
- 32:45 – 37:55
From poker to practice: startups, negotiation, and autonomous-vehicle coordination
- LFLex Fridman
I've, I've read you write somewhere the, when negotiations you come to the table with prior, uh, like a strategy that, mm, like that you're willing to do and not willing to do, those kinds of things. So how do you start to, now moving away from poker, moving beyond poker into other applications like negotiations, how do you start applying this to other, to other domains?
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
Maybe even real world domains that you have worked on?
- TSTuomas Sandholm
Yeah. I actually have two startup companies doing exactly that. One is called Strategic Machine, and that's for kind of business applications, gaming, sports, all sorts of things like that. Any applications of this, to business and to sports and to, uh, gaming-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... to va- various types of things, f- in finance, electricity markets, and so on. And the other is called Strategy Robot, where we are taking this to, uh, military, security, cyber security, and intelligence applications.
- LFLex Fridman
I think you worked, uh, a little bit in, um... how, how do you put it? Advertisement, uh, sort, sort of suggesting, um, ads kind of thing, um, auctioning-
- TSTuomas Sandholm
Yes, that's another company. Optimized Markets.
- LFLex Fridman
Optimized Markets.
- TSTuomas Sandholm
But that's much more about like combinatorial market and optimization based technology that's not using these, uh, game theoretic reasoning technologies.
- LFLex Fridman
I see. Okay, so what high l- sort of high level do you think about our ability to use game theoretic concepts to model human behavior? Do you think, do you think human behavior is amenable to this kinda modeling? So outside of the poker games, and where have you seen it done successfully in your work?
- TSTuomas Sandholm
I'm not sure the goal really is modeling humans. Uh, uh, like for example, if I'm playing a zero sum game-
- LFLex Fridman
Yes.
- TSTuomas Sandholm
... I don't really care that the opponent is actually following my model of rational behavior-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... because if they're not, that's even better for me.
- LFLex Fridman
Right. So, so the, see with the opponents in games, there's a, the prerequisite is that you f- formalize the interaction in some way that can be amenable to analysis. And you've done this w- um, amazing work with mechanism design, designing games that have certain outcomes. Uh, but so I- I'll tell you an example from my, uh, from my world of autonomous vehicles, right?
- TSTuomas Sandholm
Okay.
- LFLex Fridman
We're studying pedestrians, and pedestrians and cars negotiate in this non-verbal communication. There's this weird in- game dance of tension where pedestrians are basically saying, "I trust that you won't kill me, and so as a jaywalker I will step onto the road even though I'm breaking the law," and there's this tension. And the question is we really don't know how to model that well, uh, in, in trying to model intent. And so people sometimes bring up ideas of game theory and so on.
- TSTuomas Sandholm
Mm-hmm.
- LFLex Fridman
Do you think that aspect of human behavior can use these kinds of imperfect information apro- approaches, modeling? How do we s- how do you start to attack a problem like that, uh, when you don't even know how the game, to design the game to describe the situation in order to solve it?
- TSTuomas Sandholm
Okay, so I haven't really thought about jaywalking-
- LFLex Fridman
(laughs)
- TSTuomas Sandholm
... but one thing that, uh, I, I think could be a good application in, in, in autonomous vehicles is the following. So let's say that you have fleets of autonomous cars operated by different companies.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So maybe here's the Waymo fleet and here's the Uber fleet. Uh, if you think about the rules of the road, they define certain legal rules, but that still le- leaves a huge strategy space open.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Like as a simple example, when cars merge-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... you know how humans merge, you know, they slow down and look at each other and, uh, w- try to, uh, try to merge. Wouldn't it be better if these situations would already be pre-negotiated so we can actually merge at full speed and we know that this is the situation, this is how we do it, and it's all gonna be faster?
- 37:55 – 43:16
Performance-oriented research: why scaling experiments matter (and the poker backlash)
- LFLex Fridman
Great. So what lessons have you learned? The Annual Computer Poker Competition, uh, an incredible accomplishment of AI. You know, you look at the history of Deep Blue, AlphaGo, these kind of, um, moments when AI stepped up in an engineering effort and a scientific effort combined to, to beat the best of human players. So what, what do you take away from this whole experience? What have you learned about designing AI systems that play these kinds of games, and what does that mean for sort of, uh, AI in general, for the future of AI development?
- TSTuomas Sandholm
Yeah, so that's a good question. Uh, so there's so much to say about it. I do like this type of performance-oriented research.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Although in my group we go all the way from, like, idea, to theory, to experiments, to big system building, to commercialization. So we span that spectrum, bu- but I think that in a lot of situations in AI, you really have to build the big systems and evaluate them at s- at scale before you know what works and doesn't. And we've seen that in the, uh, computational game theory community-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... that there are a lot of techniques that look good in the small, but then they cease to look good in the large. And we've also seen that there are a lot of techniques that look superior in theory-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And I, I really mean in, in terms of convergence rates, better... Like, first, auto methods, better convergence rates, like the CFR-based, based algorithms, yet the CFR-ba- based algorithms are the fir- fastest in practice. So it really tells me that you have to test this in re- uh, reality. The theory isn't tight enough, if you will, to tell you which algorithms are better than the others. And, uh, you have to look at these things at... in the large, because w- any sort of projections you do from the small can, at least in this domain, be very misleading. So that, that's kind of from, from a kind of a science and engineering perspective. From personal perspective, it's been just a wild experience in that, uh, with the first poker competition, the first s- so first, uh, brains-versus-AI and machine poker competition that we organized. There had been, by the way, for other poker games, there had been previous competitions, but this was for heads-up no limit, this was the first. And, uh, I probably became the most hated person in the world of poker, and I didn't mean to. I, I, I s-
- LFLex Fridman
Why is that?
- TSTuomas Sandholm
Uh, the, the, uh-
- LFLex Fridman
For cracking the game, for solving-
- TSTuomas Sandholm
Yeah. Uh, it was qu- uh, a lot of people felt that it was a real threat to the whole game, the whole existence of the game. If, if, if AI becomes better than humans, people would be scared to play poker, because there are these superhuman AIs running around taking their money and, you know, all of that. So, so I just... uh, it was really aggressive. Uh-
- LFLex Fridman
Interesting.
- TSTuomas Sandholm
... the comments were super aggressive. I got everything s- just short of death threats. (laughs)
- LFLex Fridman
(laughs) Do, do you think the same was true for chess? 'Cause right now, I mean, they just completed the world championships in chess, and humans just started ignoring the fact that there's AI systems now that outperform humans, and they still enjoy the game. It's still a beautiful game.
- TSTuomas Sandholm
See, that's what I think.
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
Yeah, and I think the same thing happens in poker. And so, so I didn't think of myself as somebody who was gonna kill the game, and I don't think I did.
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
I've really lu- learned to love this game. I wasn't a poker player before, but learned so many nuances about the... from these AIs, and they've really changed how the game is played, by the way.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So they have these very Martian ways of playing poker, and the top humans are now incorporating those types of strategies into their own play. So if anything, to me, our work has made poker a richer, more interesting game for humans to play, not something that is gonna steer humans away from it entirely.
- LFLex Fridman
Just a quick comment on something you said, which is... if I may say so, in academia is a little bit rare sometimes, it, it's, it's pretty brave to put your ideas to the test in the way you described, uh, s- saying that sometimes good ideas don't work when you actually try to apply them at scale. And so where does that come from? I mean, what... uh, if you could do, um... advice for people, what, what drives you in that sense? Were you always this way? (laughs) I mean, it ta- it takes a brave person, I guess is what I'm saying, to test their ideas and to see if this thing actually works against human... top human players and so on.
- TSTuomas Sandholm
Yeah, I don't know about brave, but it takes a lot of work.
- LFLex Fridman
Yes, it takes a lot of work.
- TSTuomas Sandholm
Uh, it, it takes a lot of work and a lot of time, uh, to organize, do... make something big, and to organize an event, and stuff like that.
- LFLex Fridman
And what drives you in that effort? Because you could still, I would argue, get a best paper award at NIPS as you did in '17 without doing this.
- TSTuomas Sandholm
That's right, yes. (laughs)
- LFLex Fridman
And s- and so, uh... (laughs)
- TSTuomas Sandholm
So in, uh, in, in general, I believe it's very important to do things in, uh, i- i- in the real world and at scale, uh, and that's really where the, the, the, the pudding, if you will-
- 43:16 – 48:31
Automated mechanism design: impossibility results and ‘islands of possibility’
- LFLex Fridman
(laughs) So, uh, the topic of mechanism design, which is really interesting, also kind of new to me, uh, except as an observer of, I don't know, politics and any... I'm an obs- observer of mechanisms, but you, you write in your paper on automated mechanism design that, that I quickly read. So mech- mechanism design is designing the rules of the game so you get a certain desirable outcome, and you have this work on doing so in an automatic fashion, as opposed to fine-tuning it. So what have you learned from those efforts? Uh, if you look, s- say, I don't know, at a complex s- like, um, like our political system, c- can we design our political system to have, in an automated fashion, uh, to have outcomes that we want? Can we design something like, um, traffic lights, uh, to be smart, uh, where it gets outcomes that we want? So what, what are the lessons that you draw from that work?
- TSTuomas Sandholm
Yeah, so I still very much believe in the automated mechanism design direction.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
But it's not a panacea. There are impossibility results in mechanism design saying that there is no mechanism that accomplishes objective X in class C. So, so there is, it, it's not going to... There's no way, using any mechanism design tools, manual or automated, to do certain things in mechanism design.
- LFLex Fridman
Can you describe that again? So meaning, there... it's impossible to achieve that, uh-
- TSTuomas Sandholm
Yeah. Yeah, c- cert- there's a certain imposs-
- LFLex Fridman
... end or unlikely? Uh-
- TSTuomas Sandholm
Impossible.
- LFLex Fridman
Impossible.
- TSTuomas Sandholm
So, so, so these are, th- these are not statements about human ingenuity, that who might come up with something smart. These are proofs-
- LFLex Fridman
Yes.
- TSTuomas Sandholm
... that if you want to accomplish properties X in class C, that is not doable with any mechanism. The good thing about automated mechanism design is that we're not really designing for a class. We're designing for specific settings at a time.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
So even if there's an impossibility result for the whole class, it just me- doesn't mean that all of the cases in the class are impossible. It just means that some of the cases are impossible. So we can actually carve these islands of possibility within these known impossible classes, and we've actually done that. So one, one of the famous results in mechanism design is the Myerson-Satterthwaite theorem f- by Roger Myerson and Mark Satterthwaite from 1983. It's, it's an impossibility of efficient trade under imperfect information. We show that you can, in many settings, avoid that and get efficient trade anyway.
- LFLex Fridman
Depending on how they design the game. Okay, so-
- TSTuomas Sandholm
D- depending how you design the game. And of course, it's not... uh, it doesn't, in any way, any way, uh, contradict the impossibility result. The impossibility result is still there, but it just, uh, finds spots within this impossible class where, in those spots, you don't have the impossibility.
- LFLex Fridman
Sorry if I'm going a bit philosophical, but, uh, what lessons do you draw towards, like I mentioned, politics or the human interaction and designing mechanisms for, outside of just these kinds of trading or auctioning or, um, purely formal games, are human interaction, like a political system. What, how... can... do you think it- it's applicable to, yeah, politics, or to business, uh, to negotiations, these kinds of things, designing rules that have certain outcomes?
- TSTuomas Sandholm
Yeah, yeah. I, I do think so.
- LFLex Fridman
Have you seen success, that successfully done?
- TSTuomas Sandholm
There hasn't really... Oh, you mean mechanism design or automated mechanism?
- LFLex Fridman
Automated mechanism design.
- TSTuomas Sandholm
So, so mechanism design itself has had fairly limited success so far. There are certain cases, but most of the real-world situations are actually not sound from a mechanism design perspective.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Even in those cases where they've been designed by very knowledgeable mechanism design people-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... the people are typically just taking some insights from the theory and applying those insights into the real world-
- LFLex Fridman
Yes.
- TSTuomas Sandholm
... rather than applying-
- LFLex Fridman
Directly.
- TSTuomas Sandholm
... the mechanisms directly. So one famous example of, is the FCC spectrum auctions. So, um, um, I- I've also had a small role in that, and, uh, very good economists have been work- excellent economists have been working on that with no game theory. Yet, the rules that are designed in practice there, they're such that bidding truthfully is not the best strategy. Usually, mechanism design, we try to, uh, make things easy for the participants so telling the truth is the best strategy.
- 48:31 – 54:10
What’s next for AI/game solving: benchmarks, real-world strategy, and interpretability
- LFLex Fridman
If you look at the history of AI, it's marked by seminal events. Uh, AlphaGo being a world champion, human Go player. I would put Libratus winning the heads-up no-limit hold 'em as one of such event.
- TSTuomas Sandholm
Thank you.
- LFLex Fridman
Uh, uh, and what, what do you think is the next such event, uh, whether it's in your life or in the broadly AI community, that you think might be out there that would surprise the world?
- TSTuomas Sandholm
... so that's a great question, and I don't really know the answer. In terms of game solving, uh, Heads-up No Limit Texas Hold 'Em really was the one remaining widely agreed-upon benchmark, so that was the big milestone. Now, are there other things? Yeah, certainly there are, but there, uh, there is not one that the community has kind of focused on. So what could be other things? Uh, there are groups working on StarCraft. There are th- groups working on DOTA 2. These are video games.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
Or you could have, like, Diplomacy or Hanabi, you know, things like that. These are, like, recreational games, but none of them are, uh, really acknowledged as kind of the main next challenge problem, uh, like chess, or Go, or Heads-up No Limit Texas Hold 'Em was. So I, I don't really know in the game-solving space what is or what will, will be the next benchmark. I hope, kind of hope that there will be a next benchmark 'cause really, the, uh, different groups working on the same problem really drove these application-independent techniques forward very quickly over 10 years.
- LFLex Fridman
Do you think there's an open problem that excites you that you start moving away from games into real-world games, like say, the stock market trading? Uh, it's-
- TSTuomas Sandholm
Yeah, so that, that's kind of how I am, so I am probably not going to work, eh, as hard on these, uh, recreational benchmarks. Uh, I'm doing two startups on game-solving technology, Strategic Machine and Strategy Robot, and we're really interested in pushing this, uh, stuff into practice.
- LFLex Fridman
What, what do you think would be, uh, really, you know, uh, a s- uh, a powerful result that would be surprising, that would, would be, um, uh, if you can say... Uh, I mean, it's, uh, th- uh, you know, five years, ten years from now, something that statistically you would say is not very likely, but if there's a breakthrough, w- w- what achieve...
- TSTuomas Sandholm
Yeah, so I think that overall, we're in a very different situation in game theory than we are in, let's say, machine learning.
- LFLex Fridman
Yes.
- TSTuomas Sandholm
So in machine learning, it's a fairly mature technology, and it's very broadly applied and proven success in the real world. In game-solving, there are almost no applications yet. We have just become superhuman, which machine learning, you could argue happened in the '90s, if not earlier, and, uh, at least from supervised learning ap- certain complex supervised learning applications. Now, I think the next, uh, challenge problem... I know you're not asking about it this way. You're, you're asking about the technology breakthrough-
- LFLex Fridman
Yeah.
- TSTuomas Sandholm
... but I think the big, big breakthrough is to be able to show that, hey, maybe most of, let's say, military planning, or most of business strategy will actually be done strategically using computational game theory. That, that's what I would like to see as the next five or 10-year goal.
- LFLex Fridman
Maybe you can explain to me again... uh, forgive me if this is an obvious question, but, uh, you know, machine learning methods, neural networks are... suffer from not being transparent, not being explainable. Uh, game theoretic methods, you know, Nash equilibria, do they generally, when you see the different solutions, are they... uh, when you, when you talk about military operations, are they... once you see the strategies, do they make sense, are they explainable, uh, or do they suffer from the same problems as neural networks do?
- TSTuomas Sandholm
So that's, that's a good question. I would say a little bit yes and no, and, uh, well, what I mean by that is that these game-theoretic strategies, let's say Nash equilibrium-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... it has provable properties. So it's unlike, let's say, deep learning, where you kind of cross your fingers, hopefully it'll work, and then after the fact, when you have the weights, you're still crossing your fingers, and, "I hope it will work."
- LFLex Fridman
Okay.
- TSTuomas Sandholm
Uh, h- here, you know that the solution quality is there. There's provable solution-quality guarantees. Now, that doesn't necessarily mean that the strategies are human-understandable. That's a whole other problem. So it's, uh... so I think that deep learning and computational game theory are in the same boat in that sense, that both are difficult to understand. But at least the game-theoretic techniques, they have these g- guarantees of solution quality.
- LFLex Fridman
Guarantees. So d- do you see business operations, strategic operations, or even military in the future being at least the, the strong candidates being proposed by automated systems? Do you see that?
- TSTuomas Sandholm
Yeah, I do. I do. But that's more of a beli- belief than a, uh, uh-
- LFLex Fridman
(laughs)
- TSTuomas Sandholm
... and a substantiated fact.
- LFLex Fridman
Depending on where you land in optimism or pessimism, that's a really ex- to me, that's an exciting future, especially if there's, uh, uh, provable things, uh, in terms of optimality. So looking into the future, there's, uh, a few folks worried about th- uh, especially y- you look at the game of poker, which is probably one of the last benchmarks in terms of games being solved, they, they worry about the future and the existential threats of artificial
- 54:10 – 1:06:01
AI risk, societal impact, and the threats Tuomas worries about most
- LFLex Fridman
intelligence, so the negative impact in whatever form on society. Is that something that concerns you as much, or are you more optimistic about the positive impacts of AI, or...
- TSTuomas Sandholm
Oh, I am much more optimistic about the positive impacts. So just in my own work, what we've done so far, w- we run the nationwide kidney exchange. Hundreds of people are walking around, alive today. Who wouldn't be?
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
And it's increased employment. You ha- you have a lot of people now running kidney exchanges and at transplant centers, uh, interacting with the kidney, uh, exchange. You have s- extra surgeons, nurses, anesthesiologists, hospitals, all of that.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Uh, so, so employment is increasing from that, and the world is becoming a better place. Another example is, uh, combinatorial sourcing auctions. We, uh, did 800 large-scale combinatorial sourcing auctions from 2001 to 2010, uh, in a previous startup of mine called Combinet, and, um...... we increase the supply chain efficiency on that $60 billion of spend by 12.6%. So that's over $6 billion of efficiency improvement in the world.
- LFLex Fridman
That's-
- TSTuomas Sandholm
And this is not like shifting value from somebody to somebody else, just efficiency improvement, like in trucking, less empty driving, so there's less waste, less carbon footprint, and so on.
- LFLex Fridman
This is a huge positive impact in the near term, but sort of, to, to stay in it for, for a little longer, because I think game theory has a role to play here-
- TSTuomas Sandholm
Oh, le- let me actually come back on that. That-
- LFLex Fridman
... yes.
- TSTuomas Sandholm
... is one thing. I think AI is also going to make the world much safer. So, uh, so, uh, so that's another aspect that often gets overlooked.
- LFLex Fridman
Well, let me ask this question, maybe you can speak to the, the safer. So I talked to Max Tegmark and Stuart Russell, uh, who are very concerned about existential threats of AI, and often the concern is about value misalignment, so AI systems basically up, uh, working, operating towards goals that are not the same as h- human civilization, human beings. So, it seems like game theory has a role to play there, uh, to m- to, uh, make sure the values are aligned with human beings. I don't know if that's how you think about it. If not, how do you think AI might help with this problem? Uh, how, how do you think AI might make the world safer?
- TSTuomas Sandholm
Yeah. I, I think this value misalignment is a fairly theoretical worry, and I haven't really seen it any... 'Cause I do a lot of real applications, I don't see it anywhere. Uh, the closest I've seen it was the following type of mental exercise really, where I had this argument in the late '80s when we were building these transportation optimization systems, and somebody had heard that it's a good idea to have high utilization of assets. So they told me that, "Hey, why don't you put that as objective?" And we didn't even put it as an objective because I just showed him that, you know, if you had that as your objective, the solution would be to load your trucks full and drive in circles.
- LFLex Fridman
Okay.
- TSTuomas Sandholm
Nothing would ever get delivered, you'd have 100% utilization.
- LFLex Fridman
That's right.
- TSTuomas Sandholm
So yeah, I know this phenomenon, I've known this for over 30 years and, but, but I've never seen it actually be a problem reality- in reality. And yes, if you have the wrong objective, the AI will optimize that to the hilt, and it's gonna hurt more than some human who's kinda trying to solve it in a half-baked way with some human insight too, but I, I just haven't seen that materialize in practice.
- LFLex Fridman
There's this gap that you actually put your finger on, uh, very clearly just now between theory and reality that's very difficult to put into words, I think. It's what you can theoretically imagine, uh, the, the worst possible case or even, yeah, I mean, bad cases, and what usually happens in reality. So for example, to me, maybe it's something you can comment on, uh, f- having grown up in... I h- grew up in the Soviet Union, you know, there's currently 10,000 nuclear weapons in the world, and for many decades, it's, uh, theoretically, uh, surprising to me that a nuclear war has not broken out. Do you think about this aspect, from a game theoretic perspective in general, why is that true? Uh, why in theory you can see how things would go terribly wrong and somehow yet they have not?
- TSTuomas Sandholm
Yeah.
- LFLex Fridman
How do you think about that?
- TSTuomas Sandholm
So, so I do think that about that a lot. I think the biggest two threats that we're facing as mankind, one is climate change and the other is nuclear war. So as, so, so those are my main two worries that I worry about, and I, I've tried to do something about climate... Uh, thought about trying to do something for climate change twice, actually. For two of my startups I've actually commissioned studies of what we could do on those things, and we didn't really find a sweet spot but I'm still keeping an eye out on that if there's something where we could actually provide a market solution or optimization solution or some other technology solution to problems. Right now, um, like for example, pollution credit markets was what we were looking at then, and it was much more the lack of political will by those markets were not so successful rather than bad market design. So I could go in and make a better market design, but that wouldn't really move the needle on the world very much if there's no political will, and in the US, you know, the market, uh, at least the Chicago market was just shut down, uh, uh, and, and so on. So it, then it doesn't really help how great your market design was. So-
- LFLex Fridman
And on the nuclear side, it's more, so global warming is a more encroaching problem, you know, nuclear weapons (laughs) have been here, it's an obvious problem that's just been sitting there, so how do you think about... what is the mechanism design there that just made everything seem stable and are you still extremely worried?
- TSTuomas Sandholm
I am still extremely worried. So, uh, y- yeah, you probably know the simple game theory of MAD. So, so, so, uh, this was, uh, mutually assured destruction-
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
... and it's li- it doesn't require any computation, with small matrices you can actually convince yourself that the game is such that nobody wants to initiate.
- LFLex Fridman
Mm-hmm.
- TSTuomas Sandholm
Yeah, that's a very coarse-grained analysis, and it really works in a situation where you have two super powers or small number of super powers. Now things are very different. You have, uh, smaller nukes so the threshold o- of initiating is smaller and you have smaller countries and non-n- non-nation actors who may get h- nukes and so on, so it's, I, I think it's riskier now than it was maybe ever before.
- LFLex Fridman
And what idea, application of AI, you've talked about it a little bit but, what is the most exciting to you right now? I mean, you're here at NIPS, NeurIPS now-You, you ha-have a few excellent pieces of work, but what are you thinking into the future? With several companies you're doing, what's the most exciting thing or one of the exciting things?
- TSTuomas Sandholm
The number-one thing for me right now is coming up with these scalable techniques for game solving and applying them into the real world. Uh, I'm still very interested in market design as well, and we're doing that in the optimized markets. But I'm most interested if number one right now is strategic machine, strategy robot, getting that technology out there and seeing as you are in the trenches doing applications, what needs to be actually filled, what technology gaps still need to be, be filled. So it's so hard to just put your feet on the table and imagine what needs to be done. But when you're actually doing real applications, the applications tell you what needs to be done. And I really enjoy that interaction.
Episode duration: 1:06:17
Install uListen for AI-powered chat & search across the full episode — Get Full Transcript
Transcript of episode b7bStIQovcY