Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation | Lex Fridman Podcast #344
Read full transcript 99 segments
-
a lot of people were saying like oh this a lot of people were saying like oh this whole idea of Game Theory it's just whole idea of Game Theory it's just whole idea of Game Theory it's just nonsense and if you really want to make nonsense and if you really want to make nonsense and if you really want to make money you got to like look into the money you got to like look into the money you got to like look into the other person's eyes and read their soul other person's eyes and read their soul other person's eyes and read their soul and figure out what cards they have but and figure out what cards they have but and figure out what cards they have but what happened was where we played our what happened was where we played our what happened was where we played our bot against four top heads up no limit bot against four top heads up no limit bot against four top heads up no limit Hold'em poker players and the bot wasn't Hold'em poker players and the bot wasn't Hold'em poker players and the bot wasn't trying to adapt to them it wasn't trying trying to adapt to them it wasn't trying trying to adapt to them it wasn't trying to exploit them it wasn't trying to do to exploit them it wasn't trying to do to exploit them it wasn't trying to do these Mind Games it was just trying to these Mind Games it was just trying to these Mind Games it was just trying to approximate the Nash equilibrium and it approximate the Nash equilibrium and it approximate the Nash equilibrium and it crushed them the following is a conversation with no the following is a conversation with no Brown research scientists at Fair Brown research scientists at Fair Brown research scientists at Fair Facebook AI research group at meta AI he Facebook AI research group at meta AI he Facebook AI research group at meta AI he co-created the first AI system that co-created the first AI system that co-created the first AI system that achieved superhuman level performance in achieved superhuman level performance in achieved superhuman level performance in No Limit Texas Hold'em both heads up and No Limit Texas Hold'em both heads up and No Limit Texas Hold'em both heads up and multiplayer and now recently he multiplayer and now recently he multiplayer and now recently he co-created an AI system that can co-created an AI system that can co-created an AI system that can strategically out-negotiate humans using strategically out-negotiate humans using strategically out-negotiate humans using natural language in a popular board game natural language in a popular board game natural language in a popular board game called diplomacy which is a war game called diplomacy which is a war game called diplomacy which is a war game that emphasizes negotiation that emphasizes negotiation that emphasizes negotiation this is a Lex Friedman podcast to this is a Lex Friedman podcast to this is a Lex Friedman podcast to support it please check out our sponsors support it please check out our sponsors support it please check out our sponsors in the description and now dear friends in the description and now dear friends in the description and now dear friends here's gnome proud here's gnome proud here's gnome proud you've been a lead on three amazing AI you've been a lead on three amazing AI you've been a lead on three amazing AI projects so we've got libratus that projects so we've got libratus that projects so we've got libratus that solved or at least achieved human level solved or at least achieved human level solved or at least achieved human level performance on No Limit Texas Hold'em performance on No Limit Texas Hold'em performance on No Limit Texas Hold'em poker with two players heads up poker with two players heads up poker with two players heads up you got pluribus that solved No Limit you got pluribus that solved No Limit you got pluribus that solved No Limit Texas HoldEm Poker with six players and Texas HoldEm Poker with six players and Texas HoldEm Poker with six players and just now you have Cicero these are all just now you have Cicero these are all just now you have Cicero these are all names of systems that solved or achieved
-
names of systems that solved or achieved names of systems that solved or achieved human level performance on the game of human level performance on the game of human level performance on the game of diplomacy which for people who don't diplomacy which for people who don't diplomacy which for people who don't know is a popular strategy board game it know is a popular strategy board game it know is a popular strategy board game it was loved by JFK John F Kennedy and was loved by JFK John F Kennedy and was loved by JFK John F Kennedy and Henry Kissinger and many other big Henry Kissinger and many other big Henry Kissinger and many other big famous people in the decades since so famous people in the decades since so famous people in the decades since so let's talk about poker and diplomacy let's talk about poker and diplomacy let's talk about poker and diplomacy today first poker what is the game of No today first poker what is the game of No today first poker what is the game of No Limit Texas Hold'em and how's it Limit Texas Hold'em and how's it Limit Texas Hold'em and how's it different from chess well no limit Texas different from chess well no limit Texas different from chess well no limit Texas hold 'em poker is the most popular hold 'em poker is the most popular hold 'em poker is the most popular variant of Poker in the world variant of Poker in the world variant of Poker in the world so you know you go to a casino you play so you know you go to a casino you play so you know you go to a casino you play sit down at the poker table the game sit down at the poker table the game sit down at the poker table the game that you're playing is no limit Texas that you're playing is no limit Texas that you're playing is no limit Texas Hold'em Hold'em Hold'em if you watch movies about poker like if you watch movies about poker like if you watch movies about poker like Casino Royale or rounders the game that Casino Royale or rounders the game that Casino Royale or rounders the game that they're playing is no limit Texas hold they're playing is no limit Texas hold they're playing is no limit Texas hold in poker in poker in poker now it's very different from limit now it's very different from limit now it's very different from limit Hold'em in that you can bet any amount Hold'em in that you can bet any amount Hold'em in that you can bet any amount of chips that you want and so the stakes of chips that you want and so the stakes of chips that you want and so the stakes escalate really quickly you start out escalate really quickly you start out escalate really quickly you start out with like one or two dollars in the pot with like one or two dollars in the pot with like one or two dollars in the pot and then by the end of the hand you've and then by the end of the hand you've and then by the end of the hand you've got like a thousand dollars in there got like a thousand dollars in there got like a thousand dollars in there maybe so the option to increase the maybe so the option to increase the maybe so the option to increase the number of very aggressively very quickly number of very aggressively very quickly number of very aggressively very quickly is always there right the No Limit is always there right the No Limit is always there right the No Limit aspect is there's no limit to how much aspect is there's no limit to how much aspect is there's no limit to how much you can bet you know you in limit you can bet you know you in limit you can bet you know you in limit Hold'em there's like two dollars in the Hold'em there's like two dollars in the Hold'em there's like two dollars in the Pod you you can only bet like two Pod you you can only bet like two Pod you you can only bet like two dollars but if you got ten thousand dollars but if you got ten thousand dollars but if you got ten thousand dollars in front of you you're always dollars in front of you you're always dollars in front of you you're always welcome to put ten thousand dollars into welcome to put ten thousand dollars into welcome to put ten thousand dollars into the pot so I've got a chance to hang out the pot so I've got a chance to hang out the pot so I've got a chance to hang out with uh Phil Hellmuth who plays all with uh Phil Hellmuth who plays all with uh Phil Hellmuth who plays all these different variants of Poker and these different variants of Poker and these different variants of Poker and correct me if I'm wrong but it seems correct me if I'm wrong but it seems correct me if I'm wrong but it seems like no limit rewards crazy versus the like no limit rewards crazy versus the like no limit rewards crazy versus the other ones rewards more kind of
-
other ones rewards more kind of other ones rewards more kind of calculated strategy or or no because calculated strategy or or no because calculated strategy or or no because you're sort of looking from an from an you're sort of looking from an from an you're sort of looking from an from an analytic perspective is is strategy also analytic perspective is is strategy also analytic perspective is is strategy also rewarded in No Limit taxes hold on I rewarded in No Limit taxes hold on I rewarded in No Limit taxes hold on I think both variance reward strategy but think both variance reward strategy but think both variance reward strategy but I think what's different about No Limit I think what's different about No Limit I think what's different about No Limit Hold'em is it's it's much easier to get Hold'em is it's it's much easier to get Hold'em is it's it's much easier to get jumpy you know you go in there thinking jumpy you know you go in there thinking jumpy you know you go in there thinking you're going to lose you're going to you're going to lose you're going to you're going to lose you're going to play for like a hundred dollars or play for like a hundred dollars or play for like a hundred dollars or something and suddenly there's like you something and suddenly there's like you something and suddenly there's like you know a thousand dollars in the pot a lot know a thousand dollars in the pot a lot know a thousand dollars in the pot a lot of people can't handle that can you of people can't handle that can you of people can't handle that can you define jumpy when you're playing poker define jumpy when you're playing poker define jumpy when you're playing poker you always want to choose the action you always want to choose the action you always want to choose the action that's going to maximize your expected that's going to maximize your expected that's going to maximize your expected value it's kind of like kind of like value it's kind of like kind of like value it's kind of like kind of like with investing right like if you're ever with investing right like if you're ever with investing right like if you're ever in a situation where you're the amount in a situation where you're the amount in a situation where you're the amount of money that's at stake is of money that's at stake is of money that's at stake is um is going to have a material impact on um is going to have a material impact on um is going to have a material impact on your life then you're going to play in a your life then you're going to play in a your life then you're going to play in a more risk-averse style you know if more risk-averse style you know if more risk-averse style you know if somebody makes a huge bet you're gonna somebody makes a huge bet you're gonna somebody makes a huge bet you're gonna if you're playing no limit hold them and if you're playing no limit hold them and if you're playing no limit hold them and somebody makes a huge bet somebody makes a huge bet somebody makes a huge bet there might come a point where you're there might come a point where you're there might come a point where you're like this is too much money for me to like this is too much money for me to like this is too much money for me to handle like I can't risk this amount uh handle like I can't risk this amount uh handle like I can't risk this amount uh and that's what throws a lot of people and that's what throws a lot of people and that's what throws a lot of people off off off so that's the big difference I think so that's the big difference I think so that's the big difference I think between No Limit and limit what about on between No Limit and limit what about on between No Limit and limit what about on the action side when you're actually the action side when you're actually the action side when you're actually making that big bet that's what I mean making that big bet that's what I mean making that big bet that's what I mean by crazy I was I was trying to refer to by crazy I was I was trying to refer to by crazy I was I was trying to refer to the technical the the technical term of the technical the the technical term of the technical the the technical term of crazy meaning use the big jump in the crazy meaning use the big jump in the crazy meaning use the big jump in the BET to completely throw off the other BET to completely throw off the other BET to completely throw off the other person in terms of person in terms of person in terms of um their ability to reason optimally I um their ability to reason optimally I um their ability to reason optimally I think that's right I think think that's right I think think that's right I think one of the key strategies in poker is to
-
one of the key strategies in poker is to one of the key strategies in poker is to put the other person into an put the other person into an put the other person into an uncomfortable position and if you're uncomfortable position and if you're uncomfortable position and if you're doing that then you're you're playing doing that then you're you're playing doing that then you're you're playing poker well and there's a lot of poker well and there's a lot of poker well and there's a lot of opportunities to do that in no limit opportunities to do that in no limit opportunities to do that in no limit Hold'em you know you can have like 50 in Hold'em you know you can have like 50 in Hold'em you know you can have like 50 in there you throw in a thousand dollar bet there you throw in a thousand dollar bet there you throw in a thousand dollar bet and um you know that's sometimes if you and um you know that's sometimes if you and um you know that's sometimes if you do it right it puts the other person in do it right it puts the other person in do it right it puts the other person in a really tough spot now it's also a really tough spot now it's also a really tough spot now it's also possible that you make huge mistakes possible that you make huge mistakes possible that you make huge mistakes that way and so it's really easy to lose that way and so it's really easy to lose that way and so it's really easy to lose a lot of money and no limit hold them if a lot of money and no limit hold them if a lot of money and no limit hold them if you don't know what you're doing you don't know what you're doing you don't know what you're doing um but there's a lot of upside potential um but there's a lot of upside potential um but there's a lot of upside potential too so when you build systems AI systems too so when you build systems AI systems too so when you build systems AI systems that play these games we'll talk about that play these games we'll talk about that play these games we'll talk about poker we'll talk about diplomacy poker we'll talk about diplomacy poker we'll talk about diplomacy are you um are you drawn in in part by are you um are you drawn in in part by are you um are you drawn in in part by the beauty of the game itself AI aside the beauty of the game itself AI aside the beauty of the game itself AI aside or is it to you primarily a fascinating or is it to you primarily a fascinating or is it to you primarily a fascinating problem set for the eye to solve I'm problem set for the eye to solve I'm problem set for the eye to solve I'm drawn in by the beauty of the game uh drawn in by the beauty of the game uh drawn in by the beauty of the game uh when I I started playing poker when I when I I started playing poker when I when I I started playing poker when I was in high school and the idea to me was in high school and the idea to me was in high school and the idea to me that there is a correct and objectively that there is a correct and objectively that there is a correct and objectively correct way of playing poker and if you correct way of playing poker and if you correct way of playing poker and if you could figure out what that is then could figure out what that is then could figure out what that is then you're you know you're making unlimited you're you know you're making unlimited you're you know you're making unlimited money basically money basically money basically that's like a really fascinating concept that's like a really fascinating concept that's like a really fascinating concept to me to me to me um and so I was fascinated by the um and so I was fascinated by the um and so I was fascinated by the strategy of Poker even when I was like strategy of Poker even when I was like strategy of Poker even when I was like 16 years old it wasn't until like much 16 years old it wasn't until like much 16 years old it wasn't until like much later that I actually worked on poker later that I actually worked on poker later that I actually worked on poker AIS so there was a sense that you can AIS so there was a sense that you can AIS so there was a sense that you can solve poker like uh in the way you can solve poker like uh in the way you can solve poker like uh in the way you can solve chess for example or Checkers I solve chess for example or Checkers I solve chess for example or Checkers I believe Checkers got solved right yeah believe Checkers got solved right yeah believe Checkers got solved right yeah Checkers Checkers are completely solved Checkers Checkers are completely solved Checkers Checkers are completely solved up optimal optimal strategy it's up optimal optimal strategy it's up optimal optimal strategy it's impossible to beat the AI yeah and so in impossible to beat the AI yeah and so in impossible to beat the AI yeah and so in that same way you could technically that same way you could technically that same way you could technically solve chess
-
solve chess solve chess you could solve chess you could solve you could solve chess you could solve you could solve chess you could solve poker you could solve poker so this is poker you could solve poker so this is poker you could solve poker so this is this gets into the concept of an ash this gets into the concept of an ash this gets into the concept of an ash equilibrium yeah so it is a Nash equilibrium yeah so it is a Nash equilibrium yeah so it is a Nash equilibrium Okay so equilibrium Okay so equilibrium Okay so and any finite two-player zero-sum game and any finite two-player zero-sum game and any finite two-player zero-sum game there is an optimal strategy that if you there is an optimal strategy that if you there is an optimal strategy that if you play it you are guaranteed to not lose play it you are guaranteed to not lose play it you are guaranteed to not lose an expectation no matter what your an expectation no matter what your an expectation no matter what your opponent does opponent does opponent does and this is kind of a radical concept to and this is kind of a radical concept to and this is kind of a radical concept to a lot of people a lot of people a lot of people um but it's true in chess it's true in um but it's true in chess it's true in um but it's true in chess it's true in poker it's true in any finite two-player poker it's true in any finite two-player poker it's true in any finite two-player zero-sum game zero-sum game zero-sum game and to give some intuition for this you and to give some intuition for this you and to give some intuition for this you can think of rock paper scissors can think of rock paper scissors can think of rock paper scissors and rock paper scissors if you randomly and rock paper scissors if you randomly and rock paper scissors if you randomly choose between throwing rock paper and choose between throwing rock paper and choose between throwing rock paper and scissors with equal probability then no scissors with equal probability then no scissors with equal probability then no matter what your opponent does you are matter what your opponent does you are matter what your opponent does you are not going to lose an expectation you're not going to lose an expectation you're not going to lose an expectation you're not going to lose an expectation in the not going to lose an expectation in the not going to lose an expectation in the long run long run long run now the same is true for poker there now the same is true for poker there now the same is true for poker there exists some strategy some really exists some strategy some really exists some strategy some really complicated strategy that if you play complicated strategy that if you play complicated strategy that if you play that you are guaranteed to not lose that you are guaranteed to not lose that you are guaranteed to not lose money in the long run and I should say money in the long run and I should say money in the long run and I should say this is for two player poker six player this is for two player poker six player this is for two player poker six player poker is a different story yeah it's a poker is a different story yeah it's a poker is a different story yeah it's a beautiful giant mess when you say in beautiful giant mess when you say in beautiful giant mess when you say in expectation you're guaranteed not to expectation you're guaranteed not to expectation you're guaranteed not to lose in expectation what does in lose in expectation what does in lose in expectation what does in expectation mean Poker is a very high expectation mean Poker is a very high expectation mean Poker is a very high variance game so you're gonna have hands variance game so you're gonna have hands variance game so you're gonna have hands where you win you're gonna have hands where you win you're gonna have hands where you win you're gonna have hands with your lose even if you're playing with your lose even if you're playing with your lose even if you're playing the perfect strategy you can't guarantee the perfect strategy you can't guarantee the perfect strategy you can't guarantee they're going to win every single hand they're going to win every single hand they're going to win every single hand but if you play for long enough then you but if you play for long enough then you but if you play for long enough then you are guaranteed to at least break even are guaranteed to at least break even are guaranteed to at least break even and and practice probably one and and practice probably one and and practice probably one so that's an expectation the size of so that's an expectation the size of so that's an expectation the size of your stack generally speaking now that your stack generally speaking now that your stack generally speaking now that doesn't include anything about the fact doesn't include anything about the fact doesn't include anything about the fact that you can go broke it doesn't include
-
that you can go broke it doesn't include that you can go broke it doesn't include any of those kinds of normal real world any of those kinds of normal real world any of those kinds of normal real world limitations you're talking you know in limitations you're talking you know in limitations you're talking you know in the theoretical world the theoretical world the theoretical world uh what about this the zero-sum aspect uh what about this the zero-sum aspect uh what about this the zero-sum aspect how big of a constraint is that how big how big of a constraint is that how big how big of a constraint is that how big of a constraint is finite of a constraint is finite of a constraint is finite so so so finite's not a huge constraint so I mean finite's not a huge constraint so I mean finite's not a huge constraint so I mean most games that you play are finite in most games that you play are finite in most games that you play are finite in size size size um it's also true actually that there um it's also true actually that there um it's also true actually that there exists this like perfect strategy in exists this like perfect strategy in exists this like perfect strategy in many infinite games as well technically many infinite games as well technically many infinite games as well technically the game has to be compact the game has to be compact the game has to be compact um there are like some edge cases where um there are like some edge cases where um there are like some edge cases where you don't have an ash equilibrium and a you don't have an ash equilibrium and a you don't have an ash equilibrium and a two player zero-sum game so you can two player zero-sum game so you can two player zero-sum game so you can think of a game where you like you know think of a game where you like you know think of a game where you like you know if we're playing a game where whoever if we're playing a game where whoever if we're playing a game where whoever names the bigger number is the winner names the bigger number is the winner names the bigger number is the winner there's no Nash equilibrium to that game there's no Nash equilibrium to that game there's no Nash equilibrium to that game 17 yeah 18. but you beat you win again 17 yeah 18. but you beat you win again 17 yeah 18. but you beat you win again you're good at this I played a lot of you're good at this I played a lot of you're good at this I played a lot of games games games okay uh so that's and then the Zero Sum okay uh so that's and then the Zero Sum okay uh so that's and then the Zero Sum aspect the zero zero sum aspects so aspect the zero zero sum aspects so aspect the zero zero sum aspects so there exists a Nash equilibrium in non there exists a Nash equilibrium in non there exists a Nash equilibrium in non two player zero-sum games as well and by two player zero-sum games as well and by two player zero-sum games as well and by the way just to clarify what I mean by the way just to clarify what I mean by the way just to clarify what I mean by two player Zero Sum I mean there's two two player Zero Sum I mean there's two two player Zero Sum I mean there's two players and whatever one player wins the players and whatever one player wins the players and whatever one player wins the other player loses so if we're applying other player loses so if we're applying other player loses so if we're applying poker and I win 50 that means that poker and I win 50 that means that poker and I win 50 that means that you're losing fifty dollars you're losing fifty dollars you're losing fifty dollars now outside of two player zero-sum games now outside of two player zero-sum games now outside of two player zero-sum games there still exists Nash equilibria but there still exists Nash equilibria but there still exists Nash equilibria but they're not as meaningful because you they're not as meaningful because you they're not as meaningful because you know you can think of a game like Risk know you can think of a game like Risk know you can think of a game like Risk if everybody else at the and on the if everybody else at the and on the if everybody else at the and on the board decides a team up against you and board decides a team up against you and board decides a team up against you and take you out there's no perfect strategy take you out there's no perfect strategy take you out there's no perfect strategy you can play that's gonna guarantee that you can play that's gonna guarantee that you can play that's gonna guarantee that you win there there's just nothing you you win there there's just nothing you you win there there's just nothing you can do so outside of two players zero can do so outside of two players zero can do so outside of two players zero some games there's no guarantee that some games there's no guarantee that some games there's no guarantee that you're going to win by playing a you're going to win by playing a you're going to win by playing a national equilibrium
-
national equilibrium national equilibrium have you ever tried to model in the have you ever tried to model in the have you ever tried to model in the other aspects of the game other aspects of the game other aspects of the game which is like the pleasure you draw from which is like the pleasure you draw from which is like the pleasure you draw from playing the game and then if you're a playing the game and then if you're a playing the game and then if you're a professional poker player if you're professional poker player if you're professional poker player if you're exciting even if you lose exciting even if you lose exciting even if you lose uh the you know the money you would get uh the you know the money you would get uh the you know the money you would get from the attention you get to the from the attention you get to the from the attention you get to the sponsor and all that kind of stuff is sponsor and all that kind of stuff is sponsor and all that kind of stuff is that that would be a fun thing to model that that would be a fun thing to model that that would be a fun thing to model to model in as I'd be make it sort of to model in as I'd be make it sort of to model in as I'd be make it sort of super complex to include the human super complex to include the human super complex to include the human factor in this in this full complexity I factor in this in this full complexity I factor in this in this full complexity I think you bring up a couple good points think you bring up a couple good points think you bring up a couple good points there so I think a lot of professional there so I think a lot of professional there so I think a lot of professional poker players I mean they got a huge poker players I mean they got a huge poker players I mean they got a huge amount of money not from actually amount of money not from actually amount of money not from actually playing poker but from the sponsorships playing poker but from the sponsorships playing poker but from the sponsorships and having a personality that people and having a personality that people and having a personality that people want to tune in and watch that that's a want to tune in and watch that that's a want to tune in and watch that that's a big that's a big way to to make a name big that's a big way to to make a name big that's a big way to to make a name for yourself in poker I just wanted from for yourself in poker I just wanted from for yourself in poker I just wanted from an AI perspective if you create and an AI perspective if you create and an AI perspective if you create and we'll talk about this more maybe a AI we'll talk about this more maybe a AI we'll talk about this more maybe a AI system that also talks trash and all system that also talks trash and all system that also talks trash and all that kind of stuff that that becomes that kind of stuff that that becomes that kind of stuff that that becomes part of the function to maximize so it's part of the function to maximize so it's part of the function to maximize so it's not just optimal poker play maybe not just optimal poker play maybe not just optimal poker play maybe sometimes you want to be chaotic maybe sometimes you want to be chaotic maybe sometimes you want to be chaotic maybe sometimes you want to be suboptimal and sometimes you want to be suboptimal and sometimes you want to be suboptimal and you lose you lose you lose um the the chaos and maybe sometimes you um the the chaos and maybe sometimes you um the the chaos and maybe sometimes you want to be overly aggressive because want to be overly aggressive because want to be overly aggressive because people the audience loves that people the audience loves that people the audience loves that that's fascinating I think I think what that's fascinating I think I think what that's fascinating I think I think what you're getting at here is that there's a you're getting at here is that there's a you're getting at here is that there's a difference between making an AI that difference between making an AI that difference between making an AI that wins a game and an AI That's fun to play wins a game and an AI That's fun to play wins a game and an AI That's fun to play with right yeah yeah and more fun to with right yeah yeah and more fun to with right yeah yeah and more fun to watch so those are all different things watch so those are all different things watch so those are all different things fun to play with and fun to watch yeah fun to play with and fun to watch yeah fun to play with and fun to watch yeah and I think you know I've I've heard uh and I think you know I've I've heard uh and I think you know I've I've heard uh talks from like game designers and and
-
talks from like game designers and and talks from like game designers and and they say like you know people that work they say like you know people that work they say like you know people that work on AI for actual recreational games that on AI for actual recreational games that on AI for actual recreational games that people play and they say yeah there's a people play and they say yeah there's a people play and they say yeah there's a big difference between trying to make an big difference between trying to make an big difference between trying to make an area that actually wins and you know you area that actually wins and you know you area that actually wins and you know you look at a game like Civilization look at a game like Civilization look at a game like Civilization um the way that the AIS play is not um the way that the AIS play is not um the way that the AIS play is not optimal for trying to win they're optimal for trying to win they're optimal for trying to win they're they're playing a different game they're they're playing a different game they're they're playing a different game they're trying to have personalities they're trying to have personalities they're trying to have personalities they're trying to be fun and engaging trying to be fun and engaging trying to be fun and engaging um and that makes for a better game yeah um and that makes for a better game yeah um and that makes for a better game yeah and we also talk about NPCs I just and we also talk about NPCs I just and we also talk about NPCs I just talked to Todd Howard who is the the talked to Todd Howard who is the the talked to Todd Howard who is the the creator of Fallout in the Elder Scrolls creator of Fallout in the Elder Scrolls creator of Fallout in the Elder Scrolls series and series and series and um Starfield the new game coming out and um Starfield the new game coming out and um Starfield the new game coming out and the Creator what I think is the greatest the Creator what I think is the greatest the Creator what I think is the greatest game of all time which is Skyrim and the game of all time which is Skyrim and the game of all time which is Skyrim and the NPCs there the AI that governs that NPCs there the AI that governs that NPCs there the AI that governs that whole game is very interesting but the whole game is very interesting but the whole game is very interesting but the NPCs also are super interesting and NPCs also are super interesting and NPCs also are super interesting and considering what language models might considering what language models might considering what language models might do to NPCs in an open world RPG do to NPCs in an open world RPG do to NPCs in an open world RPG role-playing game role-playing game role-playing game it's super exciting yeah honestly I'm I it's super exciting yeah honestly I'm I it's super exciting yeah honestly I'm I think this is like one of the first think this is like one of the first think this is like one of the first applications where we're going to see applications where we're going to see applications where we're going to see like real consumer interaction with like real consumer interaction with like real consumer interaction with large language models large language models large language models um I guess what sky um Elder Scrolls 6 um I guess what sky um Elder Scrolls 6 um I guess what sky um Elder Scrolls 6 is in development now they're probably is in development now they're probably is in development now they're probably like pretty close to finishing it but I like pretty close to finishing it but I like pretty close to finishing it but I would not be surprised at all if Elder would not be surprised at all if Elder would not be surprised at all if Elder Scrolls 7 was using large language Scrolls 7 was using large language Scrolls 7 was using large language models for their NPCs no they're not models for their NPCs no they're not models for their NPCs no they're not they're I mean I'm not saying anything they're I mean I'm not saying anything they're I mean I'm not saying anything not saying anything this is me not saying anything this is me not saying anything this is me speculating not you no but but there's speculating not you no but but there's speculating not you no but but there's they're just releasing the star feel they're just releasing the star feel they're just releasing the star feel good they do one game at a time yeah and good they do one game at a time yeah and good they do one game at a time yeah and so uh whatever it is whenever the date so uh whatever it is whenever the date so uh whatever it is whenever the date is I don't know what the data is calm is I don't know what the data is calm is I don't know what the data is calm down uh but it would be I don't know
-
down uh but it would be I don't know down uh but it would be I don't know like uh 20 24 25 26 so it's actually like uh 20 24 25 26 so it's actually like uh 20 24 25 26 so it's actually very possible they would include very possible they would include very possible they would include language models I was listening to this language models I was listening to this language models I was listening to this um this talk by a gaming executive uh um this talk by a gaming executive uh um this talk by a gaming executive uh when I was in grad school when I was in grad school when I was in grad school and one of the questions that a person and one of the questions that a person and one of the questions that a person in the audience asked is why are all in the audience asked is why are all in the audience asked is why are all these games so focused on fighting and these games so focused on fighting and these games so focused on fighting and killing and the person responded that killing and the person responded that killing and the person responded that it's just so much harder to make an AI it's just so much harder to make an AI it's just so much harder to make an AI that can talk with you and cooperate that can talk with you and cooperate that can talk with you and cooperate with you than it is to make an AI that with you than it is to make an AI that with you than it is to make an AI that can fight you can fight you can fight you and I think once this technology and I think once this technology and I think once this technology develops further and you can have a you develops further and you can have a you develops further and you can have a you can reach a point where like not every can reach a point where like not every can reach a point where like not every single line of dialogue has to be single line of dialogue has to be single line of dialogue has to be scripted it unlocks a lot of potential scripted it unlocks a lot of potential scripted it unlocks a lot of potential for new kinds of games like much more for new kinds of games like much more for new kinds of games like much more like positive interactions that are not like positive interactions that are not like positive interactions that are not so focused on fighting and I'm really so focused on fighting and I'm really so focused on fighting and I'm really looking forward to that well it might looking forward to that well it might looking forward to that well it might not be positive it might be just drama not be positive it might be just drama not be positive it might be just drama you'll be in like a Call of Duty game you'll be in like a Call of Duty game you'll be in like a Call of Duty game instead of doing the shooting you'll instead of doing the shooting you'll instead of doing the shooting you'll just be hanging out and like arguing just be hanging out and like arguing just be hanging out and like arguing with an AI about like with an AI about like with an AI about like um like passive aggressive and then you um like passive aggressive and then you um like passive aggressive and then you won't be able to sleep that night you won't be able to sleep that night you won't be able to sleep that night you have to return or continue the argument have to return or continue the argument have to return or continue the argument that you were uh emotionally hurt uh that you were uh emotionally hurt uh that you were uh emotionally hurt uh I mean yeah I think that's actually an I mean yeah I think that's actually an I mean yeah I think that's actually an exciting World whatever whatever is the exciting World whatever whatever is the exciting World whatever whatever is the drama the chaos that we love the push drama the chaos that we love the push drama the chaos that we love the push and pull of human connection I think and pull of human connection I think and pull of human connection I think it's possible to do that in the video it's possible to do that in the video it's possible to do that in the video game world and I think you could be game world and I think you could be game world and I think you could be Messier and make more mistakes in the Messier and make more mistakes in the Messier and make more mistakes in the Video Game World which is why it would Video Game World which is why it would Video Game World which is why it would be a nice place and and also it doesn't be a nice place and and also it doesn't be a nice place and and also it doesn't have a deep of a as deep of a real have a deep of a as deep of a real have a deep of a as deep of a real psychological impact because inside psychological impact because inside psychological impact because inside video games it's kind of understood that video games it's kind of understood that video games it's kind of understood that you're in a not a real world so whatever
-
you're in a not a real world so whatever you're in a not a real world so whatever crazy stuff AI does we have some crazy stuff AI does we have some crazy stuff AI does we have some flexibility to play just like with the flexibility to play just like with the flexibility to play just like with the game of diplomacy it's a game this is game of diplomacy it's a game this is game of diplomacy it's a game this is not real geopolitics not real war it's a not real geopolitics not real war it's a not real geopolitics not real war it's a it's a game so you could you can have a it's a game so you could you can have a it's a game so you could you can have a little bit of fun a little bit of chaos little bit of fun a little bit of chaos little bit of fun a little bit of chaos okay back to Natural uh how do we find okay back to Natural uh how do we find okay back to Natural uh how do we find the Nash equilibrium the Nash equilibrium the Nash equilibrium all right so there's different ways to all right so there's different ways to all right so there's different ways to find an ash equilibrium so um find an ash equilibrium so um find an ash equilibrium so um the way that we do it is with this the way that we do it is with this the way that we do it is with this process called self-play process called self-play process called self-play um basically we have this algorithm that um basically we have this algorithm that um basically we have this algorithm that starts by playing totally randomly and starts by playing totally randomly and starts by playing totally randomly and it learns how to play the game by it learns how to play the game by it learns how to play the game by playing against itself playing against itself playing against itself so so so um it will start playing the game um it will start playing the game um it will start playing the game totally randomly and then it you know if totally randomly and then it you know if totally randomly and then it you know if it's playing poker it'll eventually like it's playing poker it'll eventually like it's playing poker it'll eventually like get to the end of the end of the game get to the end of the end of the game get to the end of the end of the game and make fifty dollars and make fifty dollars and make fifty dollars and then it will like review all the and then it will like review all the and then it will like review all the decisions that it made along the way and decisions that it made along the way and decisions that it made along the way and say what would have happened if I had say what would have happened if I had say what would have happened if I had chosen this other action instead you chosen this other action instead you chosen this other action instead you know if I had raised here instead of know if I had raised here instead of know if I had raised here instead of called called called um what would the other player have done um what would the other player have done um what would the other player have done and because it's playing against a copy and because it's playing against a copy and because it's playing against a copy of itself it's able to do that of itself it's able to do that of itself it's able to do that counterfactual reasoning so they can say counterfactual reasoning so they can say counterfactual reasoning so they can say okay well if I took this action and the okay well if I took this action and the okay well if I took this action and the other person takes this action and then other person takes this action and then other person takes this action and then I take this action and eventually I make I take this action and eventually I make I take this action and eventually I make 150 instead of 50.
-
150 instead of 50. 150 instead of 50. and so it updates the regret value for and so it updates the regret value for and so it updates the regret value for that action that action that action regret is basically like how much does regret is basically like how much does regret is basically like how much does it regret having not played that action it regret having not played that action it regret having not played that action in the past in the past in the past and when it encounters that same and when it encounters that same and when it encounters that same situation again it's going to pick situation again it's going to pick situation again it's going to pick actions that have higher regret with actions that have higher regret with actions that have higher regret with higher probability higher probability higher probability now now now it'll just keep simulating the games it'll just keep simulating the games it'll just keep simulating the games this way it'll keep um you know this way it'll keep um you know this way it'll keep um you know accumulating regrets for different accumulating regrets for different accumulating regrets for different situations situations situations um and in the long run if you pick um and in the long run if you pick um and in the long run if you pick actions that have higher regret with actions that have higher regret with actions that have higher regret with higher probability in the correct way higher probability in the correct way higher probability in the correct way it's proven to converge to a Nash it's proven to converge to a Nash it's proven to converge to a Nash equilibrium equilibrium equilibrium even for super complex games even for even for super complex games even for even for super complex games even for imperfect information games it's true imperfect information games it's true imperfect information games it's true for all games it's true for it's true for all games it's true for it's true for all games it's true for it's true for chess it's true for poker it's for chess it's true for poker it's for chess it's true for poker it's particularly useful for poker so this is particularly useful for poker so this is particularly useful for poker so this is the the method of contractual regret the the method of contractual regret the the method of contractual regret minimization this is counter factual minimization this is counter factual minimization this is counter factual regret minimization that doesn't have to regret minimization that doesn't have to regret minimization that doesn't have to do with self-play has to do with just do with self-play has to do with just do with self-play has to do with just any any if you follow this kind of any any if you follow this kind of any any if you follow this kind of process self-play or not you'll be able process self-play or not you'll be able process self-play or not you'll be able to arrive in an optimal set of actions to arrive in an optimal set of actions to arrive in an optimal set of actions so this counterfactual regret so this counterfactual regret so this counterfactual regret minimization is a kind of self-play it's minimization is a kind of self-play it's minimization is a kind of self-play it's a principled kind of self-play that's a principled kind of self-play that's a principled kind of self-play that's proven to converge to Nash equilibria proven to converge to Nash equilibria proven to converge to Nash equilibria even in in private information games now even in in private information games now even in in private information games now you can have other forms of self-play you can have other forms of self-play you can have other forms of self-play and people use other forms of self-play and people use other forms of self-play and people use other forms of self-play for perfect information games for perfect information games for perfect information games um where you have more flexibility the um where you have more flexibility the um where you have more flexibility the algorithm doesn't have to be as algorithm doesn't have to be as algorithm doesn't have to be as theoretically sound in order to converge theoretically sound in order to converge theoretically sound in order to converge to that class of games because there's to that class of games because there's to that class of games because there's uh it's a simpler setting sure so I kind uh it's a simpler setting sure so I kind uh it's a simpler setting sure so I kind of in my brain the word self-play has of in my brain the word self-play has of in my brain the word self-play has mapped in you all networks but we're mapped in you all networks but we're mapped in you all networks but we're speaking something bigger than just speaking something bigger than just speaking something bigger than just neural networks it could be anything
-
neural networks it could be anything neural networks it could be anything the self-play mechanism is just the the self-play mechanism is just the the self-play mechanism is just the mechanism of a system playing itself mechanism of a system playing itself mechanism of a system playing itself exactly yeah self-play is not tied exactly yeah self-play is not tied exactly yeah self-play is not tied specifically to neural Nets it's it's a specifically to neural Nets it's it's a specifically to neural Nets it's it's a kind of reinforcement learning basically kind of reinforcement learning basically kind of reinforcement learning basically okay and I would also say this process okay and I would also say this process okay and I would also say this process of like trying to reason oh what would of like trying to reason oh what would of like trying to reason oh what would the value have been if I had taken this the value have been if I had taken this the value have been if I had taken this other action instead this is very other action instead this is very other action instead this is very similar to how humans learn to play a similar to how humans learn to play a similar to how humans learn to play a game like poker right like you probably game like poker right like you probably game like poker right like you probably played poker before and with your played poker before and with your played poker before and with your friends you probably asked like oh what friends you probably asked like oh what friends you probably asked like oh what do you have called me if I raised there do you have called me if I raised there do you have called me if I raised there you know and that's that's a person you know and that's that's a person you know and that's that's a person trying to do the same kind of like trying to do the same kind of like trying to do the same kind of like learning from a counter factual that the learning from a counter factual that the learning from a counter factual that the AI is doing okay and if you do that at AI is doing okay and if you do that at AI is doing okay and if you do that at scale you're going to be able to learn scale you're going to be able to learn scale you're going to be able to learn an optimal policy yeah now where the an optimal policy yeah now where the an optimal policy yeah now where the neural nets come in I said like okay if neural nets come in I said like okay if neural nets come in I said like okay if it's in that situation again then it it's in that situation again then it it's in that situation again then it will choose the action that has high will choose the action that has high will choose the action that has high regret now the problem is that poker is regret now the problem is that poker is regret now the problem is that poker is such a huge game you know I think no such a huge game you know I think no such a huge game you know I think no limit Texas Hold'em the version that we limit Texas Hold'em the version that we limit Texas Hold'em the version that we were playing has 10 to the 161 different were playing has 10 to the 161 different were playing has 10 to the 161 different decision points which is more than the decision points which is more than the decision points which is more than the number of atoms in the universe squared number of atoms in the universe squared number of atoms in the universe squared that's heads up that's heads up yeah 10 that's heads up that's heads up yeah 10 that's heads up that's heads up yeah 10 to the 161 you said yeah I mean it to the 161 you said yeah I mean it to the 161 you said yeah I mean it depends on the number of chips that you depends on the number of chips that you depends on the number of chips that you have the stacks and everything but like have the stacks and everything but like have the stacks and everything but like the version that we were playing was the version that we were playing was the version that we were playing was tense to the 161. which I assume would tense to the 161. which I assume would tense to the 161. which I assume would be a somewhat simplified version anyway be a somewhat simplified version anyway be a somewhat simplified version anyway because the about there's some like step because the about there's some like step because the about there's some like step function you had for like bets oh no no function you had for like bets oh no no function you had for like bets oh no no that's that's I'm saying like we played that's that's I'm saying like we played that's that's I'm saying like we played the the full game you can bet whatever the the full game you can bet whatever the the full game you can bet whatever amount you want another thought maybe amount you want another thought maybe amount you want another thought maybe was constrained in like what it was constrained in like what it was constrained in like what it considered for bed sizes but the the considered for bed sizes but the the considered for bed sizes but the the person on the other side could bet person on the other side could bet person on the other side could bet whatever they wanted yeah I mean 161 whatever they wanted yeah I mean 161 whatever they wanted yeah I mean 161 plus or minus 10 doesn't matter yeah plus or minus 10 doesn't matter yeah plus or minus 10 doesn't matter yeah um and so the way neural Nets help out
-
um and so the way neural Nets help out um and so the way neural Nets help out here is you know you don't have to run here is you know you don't have to run here is you know you don't have to run into the same exact situation because into the same exact situation because into the same exact situation because that's never going to happen again the that's never going to happen again the that's never going to happen again the odds of you running into the same exact odds of you running into the same exact odds of you running into the same exact situation are pretty slim but if you run situation are pretty slim but if you run situation are pretty slim but if you run into a similar situation then you can into a similar situation then you can into a similar situation then you can generalize from other states that you've generalize from other states that you've generalize from other states that you've been in that kind of look like that one been in that kind of look like that one been in that kind of look like that one and you can say like well these other and you can say like well these other and you can say like well these other situations I had high regret for this situations I had high regret for this situations I had high regret for this action and so maybe I should play that action and so maybe I should play that action and so maybe I should play that action here as well which is the more action here as well which is the more action here as well which is the more complex game chess or poker or go or complex game chess or poker or go or complex game chess or poker or go or poker do you know that is a poker do you know that is a poker do you know that is a controversial question okay um I'm gonna controversial question okay um I'm gonna controversial question okay um I'm gonna it's like somebody screaming on Reddit it's like somebody screaming on Reddit it's like somebody screaming on Reddit right now it depends on which subreddit right now it depends on which subreddit right now it depends on which subreddit you're on is it chess or is it poker I'm you're on is it chess or is it poker I'm you're on is it chess or is it poker I'm sure like David Silver's gonna get sure like David Silver's gonna get sure like David Silver's gonna get really angry at me yeah I'll say I'm really angry at me yeah I'll say I'm really angry at me yeah I'll say I'm gonna say poker actually and I think for gonna say poker actually and I think for gonna say poker actually and I think for a couple reasons a couple reasons a couple reasons um they're not here to defend themselves um they're not here to defend themselves um they're not here to defend themselves so first of all you have the imperfect so first of all you have the imperfect so first of all you have the imperfect information aspect and so it's um it we information aspect and so it's um it we information aspect and so it's um it we can go into that but like once you can go into that but like once you can go into that but like once you introduce imperfect information uh introduce imperfect information uh introduce imperfect information uh things get much more complicated so we things get much more complicated so we things get much more complicated so we should say should say should say maybe you can describe what is seen to maybe you can describe what is seen to maybe you can describe what is seen to the players what is not seen uh in the the players what is not seen uh in the the players what is not seen uh in the game of Texas Hold'em yeah so Texas game of Texas Hold'em yeah so Texas game of Texas Hold'em yeah so Texas Hold'em you get two cards face down that Hold'em you get two cards face down that Hold'em you get two cards face down that only you see only you see only you see um and so that's the hidden information um and so that's the hidden information um and so that's the hidden information of the game the other players also all of the game the other players also all of the game the other players also all get two cards face down that only they get two cards face down that only they get two cards face down that only they see see see um and so you have to kind of as you're um and so you have to kind of as you're um and so you have to kind of as you're playing reason about like okay what do playing reason about like okay what do playing reason about like okay what do they think I have what do they have what they think I have what do they have what they think I have what do they have what do they think I think they have that do they think I think they have that do they think I think they have that kind of stuff and kind of stuff and kind of stuff and um that's that's kind of where bluffing um that's that's kind of where bluffing um that's that's kind of where bluffing comes into play right because the fact comes into play right because the fact comes into play right because the fact that you can Bluff the fact that you can that you can Bluff the fact that you can that you can Bluff the fact that you can bet with a bad hand and still win is
-
bet with a bad hand and still win is bet with a bad hand and still win is because they don't know what your cards because they don't know what your cards because they don't know what your cards are right and that's the that's the key are right and that's the that's the key are right and that's the that's the key difference between a perfect information difference between a perfect information difference between a perfect information game like poker uh sorry like chess and game like poker uh sorry like chess and game like poker uh sorry like chess and go go go um and imprint information games like um and imprint information games like um and imprint information games like poker this is what trash talk looks like poker this is what trash talk looks like poker this is what trash talk looks like the implied statement is the game I the implied statement is the game I the implied statement is the game I solved is much tougher uh but yeah so uh solved is much tougher uh but yeah so uh solved is much tougher uh but yeah so uh when you're playing I'm just gonna do when you're playing I'm just gonna do when you're playing I'm just gonna do random questions here so what when random questions here so what when random questions here so what when you're playing your opponent you're playing your opponent you're playing your opponent under imperfect information under imperfect information under imperfect information is there some degree to which you're is there some degree to which you're is there some degree to which you're trying to estimate the range of hands trying to estimate the range of hands trying to estimate the range of hands that they have that they have that they have or is that not part of the algorithm so or is that not part of the algorithm so or is that not part of the algorithm so how what are the different approaches to how what are the different approaches to how what are the different approaches to the imperfect information game so the the imperfect information game so the the imperfect information game so the key thing to understand about why in key thing to understand about why in key thing to understand about why in perfect information makes things perfect information makes things perfect information makes things difficult is that you have to worry not difficult is that you have to worry not difficult is that you have to worry not just about which actions to play but the just about which actions to play but the just about which actions to play but the probability that you're going to play probability that you're going to play probability that you're going to play those actions those actions those actions so you think about so you think about so you think about um rock paper scissors for example rock um rock paper scissors for example rock um rock paper scissors for example rock paper scissors is an imperfect paper scissors is an imperfect paper scissors is an imperfect information game information game information game um right because you don't know what I'm um right because you don't know what I'm um right because you don't know what I'm about to throw I do but yeah usually not about to throw I do but yeah usually not about to throw I do but yeah usually not yeah yeah and so you can't just say like yeah yeah and so you can't just say like yeah yeah and so you can't just say like I'm just gonna throw a rock every single I'm just gonna throw a rock every single I'm just gonna throw a rock every single time because the other person is going time because the other person is going time because the other person is going to figure that out and notice a pattern to figure that out and notice a pattern to figure that out and notice a pattern and then suddenly you're going to start and then suddenly you're going to start and then suddenly you're going to start losing and so you don't just have to losing and so you don't just have to losing and so you don't just have to figure out like which action to play you figure out like which action to play you figure out like which action to play you have to figure out the probability that have to figure out the probability that have to figure out the probability that you play it and really importantly the you play it and really importantly the you play it and really importantly the value of an action depends on the value of an action depends on the value of an action depends on the probability that you're going to play it probability that you're going to play it probability that you're going to play it so if you're playing Rock every single so if you're playing Rock every single so if you're playing Rock every single time that value is really low but if time that value is really low but if time that value is really low but if you're never playing rock you play Rock you're never playing rock you play Rock you're never playing rock you play Rock like one percent of the time then
-
like one percent of the time then like one percent of the time then suddenly the the other person's probably suddenly the the other person's probably suddenly the the other person's probably gonna be throwing scissors and when you gonna be throwing scissors and when you gonna be throwing scissors and when you throw rock the value of that action is throw rock the value of that action is throw rock the value of that action is going to be really high going to be really high going to be really high now you take that to Poker what that now you take that to Poker what that now you take that to Poker what that means is means is means is the value of bluffing for example if the value of bluffing for example if the value of bluffing for example if you're the kind of person that never you're the kind of person that never you're the kind of person that never Bluffs and you have this reputation as Bluffs and you have this reputation as Bluffs and you have this reputation as somebody that never Bluffs and suddenly somebody that never Bluffs and suddenly somebody that never Bluffs and suddenly you Bluff there's a really good chance you Bluff there's a really good chance you Bluff there's a really good chance that that bluff is going to work and that that bluff is going to work and that that bluff is going to work and you're gonna make a lot of money on the you're gonna make a lot of money on the you're gonna make a lot of money on the other hand if you've got a reputation other hand if you've got a reputation other hand if you've got a reputation like if they seen you play for a long like if they seen you play for a long like if they seen you play for a long time and they see oh you're the kind of time and they see oh you're the kind of time and they see oh you're the kind of person that's bluffing all the time person that's bluffing all the time person that's bluffing all the time when you Bluff they're not going to buy when you Bluff they're not going to buy when you Bluff they're not going to buy it and they're going to call you down it and they're going to call you down it and they're going to call you down you're going to lose a lot of money you're going to lose a lot of money you're going to lose a lot of money and that finding that balance of how and that finding that balance of how and that finding that balance of how often you should be bluffing is uh the often you should be bluffing is uh the often you should be bluffing is uh the key challenge of a game of poker key challenge of a game of poker key challenge of a game of poker and um you contrast that with a game and um you contrast that with a game and um you contrast that with a game like chess like chess like chess it doesn't matter if you're opening with it doesn't matter if you're opening with it doesn't matter if you're opening with the Queen's Gambit 10 of the time or 100 the Queen's Gambit 10 of the time or 100 the Queen's Gambit 10 of the time or 100 of the time the value the expected value of the time the value the expected value of the time the value the expected value is the same is the same is the same so um so that's that's why we need these so um so that's that's why we need these so um so that's that's why we need these algorithms that understand not just we algorithms that understand not just we algorithms that understand not just we have to figure out what actions are good have to figure out what actions are good have to figure out what actions are good but the probabilities we need to get the but the probabilities we need to get the but the probabilities we need to get the exact probabilities correct and that's exact probabilities correct and that's exact probabilities correct and that's actually when we created the bot actually when we created the bot actually when we created the bot labradus libratus means balanced because labradus libratus means balanced because labradus libratus means balanced because the algorithm that we designed was the algorithm that we designed was the algorithm that we designed was designed to find that right balance of designed to find that right balance of designed to find that right balance of how often it should play each action how often it should play each action how often it should play each action the balance of how often in the key sort the balance of how often in the key sort the balance of how often in the key sort of branching is the bluff or not the of branching is the bluff or not the of branching is the bluff or not the bluff bluff bluff is that a is that a good crude is that a is that a good crude is that a is that a good crude simplification of the major decision in simplification of the major decision in simplification of the major decision in poker it's a good simplification I think poker it's a good simplification I think poker it's a good simplification I think that's like the main tension but it's that's like the main tension but it's that's like the main tension but it's it's not just how often the bluff or not
-
it's not just how often the bluff or not it's not just how often the bluff or not to Bluff it's like how often should you to Bluff it's like how often should you to Bluff it's like how often should you bet in general how often should you what bet in general how often should you what bet in general how often should you what what kind of bet should you make what kind of bet should you make what kind of bet should you make um should you bet big or should you bet um should you bet big or should you bet um should you bet big or should you bet small and with which with which hands uh small and with which with which hands uh small and with which with which hands uh and so this is where the idea of a range and so this is where the idea of a range and so this is where the idea of a range comes from because when you are bluffing comes from because when you are bluffing comes from because when you are bluffing with a particular hand in a particular with a particular hand in a particular with a particular hand in a particular spot spot spot you don't want there to be a pattern for you don't want there to be a pattern for you don't want there to be a pattern for the other person to pick up on you don't the other person to pick up on you don't the other person to pick up on you don't want them to figure out oh whenever this want them to figure out oh whenever this want them to figure out oh whenever this person is in this spot they're always person is in this spot they're always person is in this spot they're always bluffing and so you have to reason about bluffing and so you have to reason about bluffing and so you have to reason about okay would I also bet with a good hand okay would I also bet with a good hand okay would I also bet with a good hand in this spot in this spot in this spot you want to be unpredictable so you have you want to be unpredictable so you have you want to be unpredictable so you have to think about what would I do if I had to think about what would I do if I had to think about what would I do if I had this different set of cards is there this different set of cards is there this different set of cards is there explicit estimation of like a theory of explicit estimation of like a theory of explicit estimation of like a theory of mind that the other person has about you mind that the other person has about you mind that the other person has about you or is that just a emergent thing that or is that just a emergent thing that or is that just a emergent thing that happens happens happens the way that the Bots handle it that are the way that the Bots handle it that are the way that the Bots handle it that are really successful they have an explicit really successful they have an explicit really successful they have an explicit theory of mine so they're explicitly theory of mine so they're explicitly theory of mine so they're explicitly reasoning about what are what's the reasoning about what are what's the reasoning about what are what's the common knowledge belief what does what common knowledge belief what does what common knowledge belief what does what do you think I have what do I think you do you think I have what do I think you do you think I have what do I think you have what do you think I think you have have what do you think I think you have have what do you think I think you have um it's explicitly reasoning about that um it's explicitly reasoning about that um it's explicitly reasoning about that is there multiple U's there so is there multiple U's there so is there multiple U's there so maybe that's jumping ahead to six maybe that's jumping ahead to six maybe that's jumping ahead to six players but is there a stickiness to the players but is there a stickiness to the players but is there a stickiness to the person to so it's an iterative game person to so it's an iterative game person to so it's an iterative game you're playing the same person you're playing the same person you're playing the same person there is there's a stickiness to that there is there's a stickiness to that there is there's a stickiness to that right you're gathering information as right you're gathering information as right you're gathering information as you play it's not every every you play it's not every every you play it's not every every um every hand is in your hand is there um every hand is in your hand is there um every hand is in your hand is there um a continuation in terms of estimating um a continuation in terms of estimating um a continuation in terms of estimating what kind of player I'm facing here
-
what kind of player I'm facing here what kind of player I'm facing here that's a good question so that's a good question so that's a good question so you could approach the game that way the you could approach the game that way the you could approach the game that way the way that the Bots do it they don't and way that the Bots do it they don't and way that the Bots do it they don't and the way that humans approach it also the way that humans approach it also the way that humans approach it also expert human players the way they expert human players the way they expert human players the way they approach it is to basically assume that approach it is to basically assume that approach it is to basically assume that you know my strategy so you know my strategy so you know my strategy so I'm going to try to pick a strategy I'm going to try to pick a strategy I'm going to try to pick a strategy where even if I were to play it for 10 where even if I were to play it for 10 where even if I were to play it for 10 000 hands and you could figure out 000 hands and you could figure out 000 hands and you could figure out exactly what it was you still wouldn't exactly what it was you still wouldn't exactly what it was you still wouldn't be able to beat it basically what that be able to beat it basically what that be able to beat it basically what that means is I'm trying to approximate the means is I'm trying to approximate the means is I'm trying to approximate the Nash equilibrium I'm trying to be Nash equilibrium I'm trying to be Nash equilibrium I'm trying to be perfectly balanced because if if I'm perfectly balanced because if if I'm perfectly balanced because if if I'm playing the national equilibrium even if playing the national equilibrium even if playing the national equilibrium even if you know what my strategy is like I said you know what my strategy is like I said you know what my strategy is like I said I'm still unbeatable in expectation so I'm still unbeatable in expectation so I'm still unbeatable in expectation so so that's what that's what the bot aims so that's what that's what the bot aims so that's what that's what the bot aims for and that's actually what a lot of for and that's actually what a lot of for and that's actually what a lot of expert poker players aim for as well to expert poker players aim for as well to expert poker players aim for as well to start by playing the Nash equilibrium start by playing the Nash equilibrium start by playing the Nash equilibrium and then maybe if they spot weaknesses and then maybe if they spot weaknesses and then maybe if they spot weaknesses in the way you're playing then they can in the way you're playing then they can in the way you're playing then they can deviate a little bit to take advantage deviate a little bit to take advantage deviate a little bit to take advantage of that of that of that they aim to be unbeatable in expectation they aim to be unbeatable in expectation they aim to be unbeatable in expectation okay okay okay so who's the greatest poker player of so who's the greatest poker player of so who's the greatest poker player of all time and why is it Phil Hellmuth so all time and why is it Phil Hellmuth so all time and why is it Phil Hellmuth so this is for Phil uh so he's known this is for Phil uh so he's known this is for Phil uh so he's known um um um at least in part for maybe playing at least in part for maybe playing at least in part for maybe playing sub-optimally and he still wins a lot sub-optimally and he still wins a lot sub-optimally and he still wins a lot it's a bit chaotic so maybe it's a bit chaotic so maybe it's a bit chaotic so maybe can you speak from an AI perspective can you speak from an AI perspective can you speak from an AI perspective about the genius of his Madness or The about the genius of his Madness or The about the genius of his Madness or The Madness of his genius Madness of his genius Madness of his genius so playing sub optimally playing so playing sub optimally playing so playing sub optimally playing chaotically um as a way to make it hard to pin down um as a way to make it hard to pin down about what your strategy is so okay the about what your strategy is so okay the about what your strategy is so okay the thing that I should explain first of all thing that I should explain first of all thing that I should explain first of all was like Nash equilibrium it doesn't was like Nash equilibrium it doesn't was like Nash equilibrium it doesn't mean that it's predictable the whole
-
mean that it's predictable the whole mean that it's predictable the whole point of it is that you're trying to be point of it is that you're trying to be point of it is that you're trying to be unpredictable now I think when somebody unpredictable now I think when somebody unpredictable now I think when somebody like Phil Hellmuth might be really like Phil Hellmuth might be really like Phil Hellmuth might be really successful is not in being unpredictable successful is not in being unpredictable successful is not in being unpredictable but in being able to but in being able to but in being able to um take advantage of the other player um take advantage of the other player um take advantage of the other player and figure out where they're being and figure out where they're being and figure out where they're being predictable predictable predictable or guiding the other player into or guiding the other player into or guiding the other player into thinking that you have certain thinking that you have certain thinking that you have certain weaknesses and then and then weaknesses and then and then weaknesses and then and then understanding how they're going to understanding how they're going to understanding how they're going to change their behavior they're going to change their behavior they're going to change their behavior they're going to deviate from a Nash equilibrium style of deviate from a Nash equilibrium style of deviate from a Nash equilibrium style of play to try to take advantage of those play to try to take advantage of those play to try to take advantage of those perceived weaknesses and then counter perceived weaknesses and then counter perceived weaknesses and then counter exploit them so you kind of get into the exploit them so you kind of get into the exploit them so you kind of get into the Mind Games there so you think about Mind Games there so you think about Mind Games there so you think about these heads up poker as a dance between these heads up poker as a dance between these heads up poker as a dance between two agents I guess are you playing the two agents I guess are you playing the two agents I guess are you playing the cards are you playing the the player so cards are you playing the the player so cards are you playing the the player so this this gets down to a big argument in this this gets down to a big argument in this this gets down to a big argument in the poker community and the academic the poker community and the academic the poker community and the academic Community for a long time there was this Community for a long time there was this Community for a long time there was this debate of like what's called GTO Game debate of like what's called GTO Game debate of like what's called GTO Game Theory optimal poker or exploitative Theory optimal poker or exploitative Theory optimal poker or exploitative play play play and um up until about like 2017 when we and um up until about like 2017 when we and um up until about like 2017 when we did the broadest match I think actually did the broadest match I think actually did the broadest match I think actually exploitative play had the advantage a exploitative play had the advantage a exploitative play had the advantage a lot of people were saying like oh this lot of people were saying like oh this lot of people were saying like oh this whole idea of Game Theory it's just whole idea of Game Theory it's just whole idea of Game Theory it's just nonsense and if you really want to make nonsense and if you really want to make nonsense and if you really want to make money you got to like look into the money you got to like look into the money you got to like look into the other person's eyes and read their soul other person's eyes and read their soul other person's eyes and read their soul and figure out what cards they have but and figure out what cards they have but and figure out what cards they have but what happened was people started what happened was people started what happened was people started adopting the game theory optimal adopting the game theory optimal adopting the game theory optimal strategy strategy strategy um and they were making good money and um and they were making good money and um and they were making good money and they weren't trying to adapt so much to they weren't trying to adapt so much to they weren't trying to adapt so much to the other player they were just trying the other player they were just trying the other player they were just trying to play the national equilibrium and to play the national equilibrium and to play the national equilibrium and then what really solidified it I think then what really solidified it I think then what really solidified it I think was the broadest the broadest match was the broadest the broadest match was the broadest the broadest match where we played our bot against four top where we played our bot against four top where we played our bot against four top heads up no limit Hold'em poker players
-
heads up no limit Hold'em poker players heads up no limit Hold'em poker players and the bot wasn't trying to adapt to and the bot wasn't trying to adapt to and the bot wasn't trying to adapt to them it wasn't trying to exploit them it them it wasn't trying to exploit them it them it wasn't trying to exploit them it wasn't trying to do these Mind Games it wasn't trying to do these Mind Games it wasn't trying to do these Mind Games it was just trying to approximate the Nash was just trying to approximate the Nash was just trying to approximate the Nash equilibrium and it crushed them equilibrium and it crushed them equilibrium and it crushed them I think you know I think you know I think you know it we've we're playing for 50 100 blinds it we've we're playing for 50 100 blinds it we've we're playing for 50 100 blinds and over the course of about 120 000 and over the course of about 120 000 and over the course of about 120 000 hands it made close to two million hands it made close to two million hands it made close to two million dollars 120 000 hands 120 000 hands dollars 120 000 hands 120 000 hands dollars 120 000 hands 120 000 hands against humans yeah and this was this against humans yeah and this was this against humans yeah and this was this was fake money to be clear so there was was fake money to be clear so there was was fake money to be clear so there was real money at stake there was 200 000 real money at stake there was 200 000 real money at stake there was 200 000 first of all all money is fake but um first of all all money is fake but um first of all all money is fake but um that's that's that's a different that's that's that's a different that's that's that's a different conversation conversation conversation um we give it meaning uh it's an it's a um we give it meaning uh it's an it's a um we give it meaning uh it's an it's a it's a phenomena that gets meaning from it's a phenomena that gets meaning from it's a phenomena that gets meaning from our uh complex psychology as a human our uh complex psychology as a human our uh complex psychology as a human civilization civilization civilization um it's emerging from the collective um it's emerging from the collective um it's emerging from the collective intelligence of the human species but intelligence of the human species but intelligence of the human species but that's not what you mean you mean like that's not what you mean you mean like that's not what you mean you mean like there's literally you can't you can't there's literally you can't you can't there's literally you can't you can't buy stuff with it okay can you actually buy stuff with it okay can you actually buy stuff with it okay can you actually uh step back and take me through that uh step back and take me through that uh step back and take me through that um competition yeah okay so um competition yeah okay so um competition yeah okay so when I was in grad school when I was in grad school when I was in grad school um there was this thing called the um there was this thing called the um there was this thing called the annual computer poker competition where annual computer poker competition where annual computer poker competition where every year all the different research every year all the different research every year all the different research Labs that were working on AI for poker Labs that were working on AI for poker Labs that were working on AI for poker would get together they would make a bot would get together they would make a bot would get together they would make a bot they would play them against each other they would play them against each other they would play them against each other uh and we made a bot that actually won uh and we made a bot that actually won uh and we made a bot that actually won the um 2014 competition the 2016 the um 2014 competition the 2016 the um 2014 competition the 2016 competition uh and so we decided we're competition uh and so we decided we're competition uh and so we decided we're gonna take this bot build on it and play gonna take this bot build on it and play gonna take this bot build on it and play against Real top professional heads up against Real top professional heads up against Real top professional heads up no limit Texas hold 'em poker players no limit Texas hold 'em poker players no limit Texas hold 'em poker players so we invited four of the world's best so we invited four of the world's best so we invited four of the world's best players in this specialty and we
-
players in this specialty and we players in this specialty and we challenge them to 120 000 hands of poker challenge them to 120 000 hands of poker challenge them to 120 000 hands of poker over the course of 20 days over the course of 20 days over the course of 20 days um and we had 200 000 200 000 in prize um and we had 200 000 200 000 in prize um and we had 200 000 200 000 in prize money at stake where it would basically money at stake where it would basically money at stake where it would basically be divided among them depending on how be divided among them depending on how be divided among them depending on how well they did relative to each other well they did relative to each other well they did relative to each other so we wanted to have some incentive for so we wanted to have some incentive for so we wanted to have some incentive for them to play their best them to play their best them to play their best did you have a confidence did you have a confidence did you have a confidence 2014-16 that this is even possible how 2014-16 that this is even possible how 2014-16 that this is even possible how much doubt was there so and we did a much doubt was there so and we did a much doubt was there so and we did a competition actually in 2015 where we competition actually in 2015 where we competition actually in 2015 where we also played against professional poker also played against professional poker also played against professional poker players and the bot lost by by a pretty players and the bot lost by by a pretty players and the bot lost by by a pretty sizable margin actually now there were sizable margin actually now there were sizable margin actually now there were some big improvements from 2015 to 2017. some big improvements from 2015 to 2017. some big improvements from 2015 to 2017. and so can you speak to the improvements and so can you speak to the improvements and so can you speak to the improvements is it computational nature is it the is it computational nature is it the is it computational nature is it the algorithm the the methods it was it was algorithm the the methods it was it was algorithm the the methods it was it was really an algorithmic approach that was really an algorithmic approach that was really an algorithmic approach that was the difference so 2015 it was much more the difference so 2015 it was much more the difference so 2015 it was much more focused on trying to come up with a focused on trying to come up with a focused on trying to come up with a strategy up front like trying to solve strategy up front like trying to solve strategy up front like trying to solve the entire game of poker like and then the entire game of poker like and then the entire game of poker like and then just have a lookup table where you're just have a lookup table where you're just have a lookup table where you're saying like oh I'm in this situation saying like oh I'm in this situation saying like oh I'm in this situation what's the strategy what's the strategy what's the strategy um the approach that we took in 2017 was um the approach that we took in 2017 was um the approach that we took in 2017 was much more search based it was trying to much more search based it was trying to much more search based it was trying to say okay well let me in real time try to say okay well let me in real time try to say okay well let me in real time try to compute a much better strategy than what compute a much better strategy than what compute a much better strategy than what I had pre-computed by playing against I had pre-computed by playing against I had pre-computed by playing against myself during self-play what is the myself during self-play what is the myself during self-play what is the search space for search space for search space for poker what are you searching over poker what are you searching over poker what are you searching over what's that look like there's different what's that look like there's different what's that look like there's different actions like raising calling yeah what actions like raising calling yeah what actions like raising calling yeah what are the actions are the actions are the actions um is it just a search over actions so um is it just a search over actions so um is it just a search over actions so in a game like chess the the search is
-
in a game like chess the the search is in a game like chess the the search is like okay I'm in this chess position and like okay I'm in this chess position and like okay I'm in this chess position and I can like you know move these different I can like you know move these different I can like you know move these different pieces and see where things end up in pieces and see where things end up in pieces and see where things end up in poker what you're searching over is the poker what you're searching over is the poker what you're searching over is the actions you can take for your hand the actions you can take for your hand the actions you can take for your hand the probabilities that you take those probabilities that you take those probabilities that you take those actions and then also the probabilities actions and then also the probabilities actions and then also the probabilities that you take other actions with other that you take other actions with other that you take other actions with other hands that you might have hands that you might have hands that you might have um and and that's kind of like a hard to um and and that's kind of like a hard to um and and that's kind of like a hard to wrap your head around like why are you wrap your head around like why are you wrap your head around like why are you searching over these like other hands searching over these like other hands searching over these like other hands that you might have and like trying to that you might have and like trying to that you might have and like trying to figure out what you would do with those figure out what you would do with those figure out what you would do with those hands hands hands um and the idea is is again you you um and the idea is is again you you um and the idea is is again you you wanna wanna wanna you wanna always be balanced and you wanna always be balanced and you wanna always be balanced and unpredictable and so if you're a search unpredictable and so if you're a search unpredictable and so if you're a search algorithm that's saying like oh I want algorithm that's saying like oh I want algorithm that's saying like oh I want to raise with this hand well in order to to raise with this hand well in order to to raise with this hand well in order to know whether that's a good action like know whether that's a good action like know whether that's a good action like let's say it's a bluff you know let's let's say it's a bluff you know let's let's say it's a bluff you know let's say you have a bad hand and you're say you have a bad hand and you're say you have a bad hand and you're saying like oh I I think I should be saying like oh I I think I should be saying like oh I I think I should be betting here with this really bad hand betting here with this really bad hand betting here with this really bad hand and bluffing well that all that's only a and bluffing well that all that's only a and bluffing well that all that's only a good action if you're also good action if you're also good action if you're also betting with a strong hand otherwise betting with a strong hand otherwise betting with a strong hand otherwise it's an obvious Bluff so if your action it's an obvious Bluff so if your action it's an obvious Bluff so if your action in some sense maximizes your in some sense maximizes your in some sense maximizes your unpredictability so that action could be unpredictability so that action could be unpredictability so that action could be mapped by your opponent to a lot of mapped by your opponent to a lot of mapped by your opponent to a lot of different hands then that's a good different hands then that's a good different hands then that's a good action basically what you want to do is action basically what you want to do is action basically what you want to do is put your opponent into a tough spot so put your opponent into a tough spot so put your opponent into a tough spot so you want them to always have some doubt you want them to always have some doubt you want them to always have some doubt like should I call here should I fold like should I call here should I fold like should I call here should I fold here and if you are raising in the here and if you are raising in the here and if you are raising in the appropriate balance between Bluffs and appropriate balance between Bluffs and appropriate balance between Bluffs and good hands then you're putting them into good hands then you're putting them into good hands then you're putting them into that tough spot and so that's what we're that tough spot and so that's what we're that tough spot and so that's what we're trying to do we're always trying to trying to do we're always trying to trying to do we're always trying to search for a strategy that would put the search for a strategy that would put the search for a strategy that would put the opponent into a difficult position can opponent into a difficult position can opponent into a difficult position can you give a metric that you're trying to you give a metric that you're trying to you give a metric that you're trying to maximize or minimize does this have to maximize or minimize does this have to maximize or minimize does this have to do with the regret thing what we're do with the regret thing what we're do with the regret thing what we're talking about in terms of putting your
-
talking about in terms of putting your talking about in terms of putting your opponent in a maximally tough spot yeah opponent in a maximally tough spot yeah opponent in a maximally tough spot yeah ultimately what you're trying to ultimately what you're trying to ultimately what you're trying to maximize is your expected winnings like maximize is your expected winnings like maximize is your expected winnings like your expected value the amount of money your expected value the amount of money your expected value the amount of money that you're going to walk away from that you're going to walk away from that you're going to walk away from assuming that your opponent was playing assuming that your opponent was playing assuming that your opponent was playing optimally in response so you're going to optimally in response so you're going to optimally in response so you're going to assume that your opponent is is also assume that your opponent is is also assume that your opponent is is also playing um like as as well as possible playing um like as as well as possible playing um like as as well as possible Nash equilibrium approach because if Nash equilibrium approach because if Nash equilibrium approach because if they're not then you're just going to they're not then you're just going to they're not then you're just going to make more money right like anything that make more money right like anything that make more money right like anything that deviates like by definition the national deviates like by definition the national deviates like by definition the national equilibrium is the strategy that does equilibrium is the strategy that does equilibrium is the strategy that does the best in expectation and so if you're the best in expectation and so if you're the best in expectation and so if you're deviating from that then you're just deviating from that then you're just deviating from that then you're just they're going to lose money and since they're going to lose money and since they're going to lose money and since it's a two player zero-sum game that it's a two player zero-sum game that it's a two player zero-sum game that means you're gonna make money so there's means you're gonna make money so there's means you're gonna make money so there's not an explicit like objective function not an explicit like objective function not an explicit like objective function that maximizes the toughness of the spot that maximizes the toughness of the spot that maximizes the toughness of the spot they're put in you're always they're put in you're always they're put in you're always this is not from like a self-play this is not from like a self-play this is not from like a self-play reinforcement learning perspective reinforcement learning perspective reinforcement learning perspective you're just trying to maximize winnings you're just trying to maximize winnings you're just trying to maximize winnings and the rest is implicit that's right and the rest is implicit that's right and the rest is implicit that's right yeah so we're what we're actually trying yeah so we're what we're actually trying yeah so we're what we're actually trying to maximize is the expected value given to maximize is the expected value given to maximize is the expected value given that the opponent is playing optimally that the opponent is playing optimally that the opponent is playing optimally in response to us now in practice what in response to us now in practice what in response to us now in practice what that ends up looking like is it's that ends up looking like is it's that ends up looking like is it's putting the opponent into difficult putting the opponent into difficult putting the opponent into difficult situations where there's no obvious situations where there's no obvious situations where there's no obvious decision to be made so the the system decision to be made so the the system decision to be made so the the system doesn't know anything about the doesn't know anything about the doesn't know anything about the difficulty of the situation not at all difficulty of the situation not at all difficulty of the situation not at all it doesn't care okay yeah all right my it doesn't care okay yeah all right my it doesn't care okay yeah all right my head was getting excited whenever I was head was getting excited whenever I was head was getting excited whenever I was making the other the opponent's sweat making the other the opponent's sweat making the other the opponent's sweat okay so you're in 2015 you didn't do as okay so you're in 2015 you didn't do as okay so you're in 2015 you didn't do as well so what's the journey from that to well so what's the journey from that to well so what's the journey from that to a system that in your mind could have a a system that in your mind could have a a system that in your mind could have a chance so 2015 we we got we got beat chance so 2015 we we got we got beat chance so 2015 we we got we got beat pretty badly and we actually learned a
-
pretty badly and we actually learned a pretty badly and we actually learned a lot from that competition and in lot from that competition and in lot from that competition and in particular you know what became clear to particular you know what became clear to particular you know what became clear to me is that the way the humans were me is that the way the humans were me is that the way the humans were approaching the game was very different approaching the game was very different approaching the game was very different from how the bot was approaching the from how the bot was approaching the from how the bot was approaching the game the bot would not be doing search game the bot would not be doing search game the bot would not be doing search it would just be trying to compute you it would just be trying to compute you it would just be trying to compute you know it would do like months of know it would do like months of know it would do like months of self-play it would just be playing self-play it would just be playing self-play it would just be playing against itself for months but then when against itself for months but then when against itself for months but then when it's actually playing the game it would it's actually playing the game it would it's actually playing the game it would just act instantly just act instantly just act instantly um and the humans when they're in a um and the humans when they're in a um and the humans when they're in a tough spot they would sit there and tough spot they would sit there and tough spot they would sit there and think for sometimes even like five think for sometimes even like five think for sometimes even like five minutes about whether they're going to minutes about whether they're going to minutes about whether they're going to call or fold a hand call or fold a hand call or fold a hand um and it became clear to me that that's um and it became clear to me that that's um and it became clear to me that that's there's a good chance that that's what there's a good chance that that's what there's a good chance that that's what that's what's missing from our bot so I that's what's missing from our bot so I that's what's missing from our bot so I actually did some actually did some actually did some um initial experiments to try to figure um initial experiments to try to figure um initial experiments to try to figure out how much of a difference this is out how much of a difference this is out how much of a difference this is actually make and the difference was actually make and the difference was actually make and the difference was huge as a signal to the human player how huge as a signal to the human player how huge as a signal to the human player how long you took to think no no I'm not long you took to think no no I'm not long you took to think no no I'm not saying that there were any timing tells saying that there were any timing tells saying that there were any timing tells I was saying when the human like the bot I was saying when the human like the bot I was saying when the human like the bot would always act instantly it wouldn't would always act instantly it wouldn't would always act instantly it wouldn't try to come up with a better strategy in try to come up with a better strategy in try to come up with a better strategy in real time real time real time um over what it had precomputed during um over what it had precomputed during um over what it had precomputed during training whereas the human like they training whereas the human like they training whereas the human like they have all this intuition about how to have all this intuition about how to have all this intuition about how to play but they're also in real time play but they're also in real time play but they're also in real time leveraging their ability to think just leveraging their ability to think just leveraging their ability to think just to search to plan to search to plan to search to plan um and coming up with an even better um and coming up with an even better um and coming up with an even better strategy than what their intuition would strategy than what their intuition would strategy than what their intuition would say so you're saying that there's you're say so you're saying that there's you're say so you're saying that there's you're doing that's what you mean by you're doing that's what you mean by you're doing that's what you mean by you're doing search also you have an you have a doing search also you have an you have a doing search also you have an you have a intuition and searched on top of that intuition and searched on top of that intuition and searched on top of that looking for a better solution yeah looking for a better solution yeah looking for a better solution yeah that's that's what I mean by search that that's that's what I mean by search that that's that's what I mean by search that um instead of acting instantly you know um instead of acting instantly you know um instead of acting instantly you know a neural net usually gives you a
-
a neural net usually gives you a a neural net usually gives you a response in like 100 milliseconds or response in like 100 milliseconds or response in like 100 milliseconds or something it depends on the size of the something it depends on the size of the something it depends on the size of the of the net but if you can leverage extra of the net but if you can leverage extra of the net but if you can leverage extra computational resources computational resources computational resources you can't possibly get a much better you can't possibly get a much better you can't possibly get a much better outcome and we did some experiments in outcome and we did some experiments in outcome and we did some experiments in small scale versions of Poker and what small scale versions of Poker and what small scale versions of Poker and what we what we found was that if you we what we found was that if you we what we found was that if you do a little bit of search even just a do a little bit of search even just a do a little bit of search even just a little bit it was the equivalent of little bit it was the equivalent of little bit it was the equivalent of making your you know your pre-computed making your you know your pre-computed making your you know your pre-computed strategy like you could kind of think it strategy like you could kind of think it strategy like you could kind of think it as your neural net a thousand times as your neural net a thousand times as your neural net a thousand times bigger bigger bigger with just a little bit of search and it with just a little bit of search and it with just a little bit of search and it just like blew away all of the research just like blew away all of the research just like blew away all of the research that we had been working on and trying that we had been working on and trying that we had been working on and trying to like scale up this like pre-computed to like scale up this like pre-computed to like scale up this like pre-computed solution it was dwarfed by the benefit solution it was dwarfed by the benefit solution it was dwarfed by the benefit that we got from search that we got from search that we got from search can you just Linger on what you mean by can you just Linger on what you mean by can you just Linger on what you mean by search here you're searching over a search here you're searching over a search here you're searching over a space of actions space of actions space of actions for your hand and for other hands how for your hand and for other hands how for your hand and for other hands how are you selecting the other hands to are you selecting the other hands to are you selecting the other hands to search over search over search over and so yeah randomly no it's all the and so yeah randomly no it's all the and so yeah randomly no it's all the other hands that you could have so when other hands that you could have so when other hands that you could have so when you're playing No Limit taxes hold on you're playing No Limit taxes hold on you're playing No Limit taxes hold on you've got two face down cards and so you've got two face down cards and so you've got two face down cards and so that's 52 choose two one thousand three that's 52 choose two one thousand three that's 52 choose two one thousand three hundred twenty six different hundred twenty six different hundred twenty six different combinations now that's actually a combinations now that's actually a combinations now that's actually a little bit lower because there's little bit lower because there's little bit lower because there's Facebook cards in the middle and so you Facebook cards in the middle and so you Facebook cards in the middle and so you can eliminate those as well but you're can eliminate those as well but you're can eliminate those as well but you're looking at like around a thousand looking at like around a thousand looking at like around a thousand different possible hands that you can different possible hands that you can different possible hands that you can have and so when we're doing when the have and so when we're doing when the have and so when we're doing when the bot's doing search It's thinking bot's doing search It's thinking bot's doing search It's thinking explicitly there are these thousand explicitly there are these thousand explicitly there are these thousand different hands that I could have there different hands that I could have there different hands that I could have there are these thousand different hands that are these thousand different hands that are these thousand different hands that you could have you could have you could have let me try to figure out what would it let me try to figure out what would it let me try to figure out what would it be a better strategy than what I've be a better strategy than what I've be a better strategy than what I've pre-computed for these hands and your
-
pre-computed for these hands and your pre-computed for these hands and your hands hands hands Okay so Okay so Okay so that search how do you fuse that with that search how do you fuse that with that search how do you fuse that with what the neural net is telling you or what the neural net is telling you or what the neural net is telling you or what the the the train system is telling what the the the train system is telling what the the the train system is telling you yeah so you yeah so you yeah so you kind of like where the train system you kind of like where the train system you kind of like where the train system comes in is is the value comes in is is the value comes in is is the value um at the end so there's um at the end so there's um at the end so there's um you only look so far ahead you look um you only look so far ahead you look um you only look so far ahead you look like maybe you know one round ahead so like maybe you know one round ahead so like maybe you know one round ahead so if you're on the Flop you're looking to if you're on the Flop you're looking to if you're on the Flop you're looking to the start of the turn the start of the turn the start of the turn um um um and at that point you can use the and at that point you can use the and at that point you can use the pre-computed solution to figure out what pre-computed solution to figure out what pre-computed solution to figure out what are what's the value here of like of are what's the value here of like of are what's the value here of like of this strategy this strategy this strategy is it of a single action essentially in is it of a single action essentially in is it of a single action essentially in that spot you're getting a value or is that spot you're getting a value or is that spot you're getting a value or is it the value of the entire series of it the value of the entire series of it the value of the entire series of actions well it's kind of both actions well it's kind of both actions well it's kind of both um because you're trying to maximize the um because you're trying to maximize the um because you're trying to maximize the value for value for value for the hand that you have but in the the hand that you have but in the the hand that you have but in the process in order to maximize the value process in order to maximize the value process in order to maximize the value of the hand that you have you have to of the hand that you have you have to of the hand that you have you have to figure out what would I be doing with figure out what would I be doing with figure out what would I be doing with all these other hands as well okay but all these other hands as well okay but all these other hands as well okay but you are you in the search always going you are you in the search always going you are you in the search always going to the end of the game in liberatis we to the end of the game in liberatis we to the end of the game in liberatis we did uh so we only use search starting on did uh so we only use search starting on did uh so we only use search starting on the turn and then we searched all the the turn and then we searched all the the turn and then we searched all the way to the end of the game the turn the way to the end of the game the turn the way to the end of the game the turn the river river river uh can we take it through the uh can we take it through the uh can we take it through the terminology yeah there's four rounds of terminology yeah there's four rounds of terminology yeah there's four rounds of Poker so there's the pre-flop the Flop Poker so there's the pre-flop the Flop Poker so there's the pre-flop the Flop the turn and the river uh and so we the turn and the river uh and so we the turn and the river uh and so we would start doing search halfway through would start doing search halfway through would start doing search halfway through the game now the first half of the game the game now the first half of the game the game now the first half of the game that was all pre-computed it would just that was all pre-computed it would just that was all pre-computed it would just act instantly and then when it got to at act instantly and then when it got to at act instantly and then when it got to at the halfway point then it would always
-
the halfway point then it would always the halfway point then it would always search to the end of the game now we search to the end of the game now we search to the end of the game now we later improved this so wouldn't have to later improved this so wouldn't have to later improved this so wouldn't have to search all the way to the end of the search all the way to the end of the search all the way to the end of the game it would actually search game it would actually search game it would actually search um just a few moves ahead um just a few moves ahead um just a few moves ahead um but that that came later and that um but that that came later and that um but that that came later and that drastically reduced the num the amount drastically reduced the num the amount drastically reduced the num the amount of computational resources that we of computational resources that we of computational resources that we needed but the moves because you can needed but the moves because you can needed but the moves because you can keep betting on top of each other that's keep betting on top of each other that's keep betting on top of each other that's what you mean by moves so like that's what you mean by moves so like that's what you mean by moves so like that's where you don't just get one bet where you don't just get one bet where you don't just get one bet per Turner poker you can have multiple per Turner poker you can have multiple per Turner poker you can have multiple arbitrary number of bets right right I'm arbitrary number of bets right right I'm arbitrary number of bets right right I'm trying to think like I'm gonna bet and trying to think like I'm gonna bet and trying to think like I'm gonna bet and then what are you gonna do in response then what are you gonna do in response then what are you gonna do in response are you gonna raise me are you going to are you gonna raise me are you going to are you gonna raise me are you going to call and then if you raise what should I call and then if you raise what should I call and then if you raise what should I do so it's reasoning about that whole do so it's reasoning about that whole do so it's reasoning about that whole process up until the end of the game in process up until the end of the game in process up until the end of the game in the case of liberatis so for liberatis the case of liberatis so for liberatis the case of liberatis so for liberatis what's the the most number of re-racists what's the the most number of re-racists what's the the most number of re-racists have you ever seen have you ever seen have you ever seen uh you probably cap out at like five or uh you probably cap out at like five or uh you probably cap out at like five or something because at that point you're something because at that point you're something because at that point you're basically all in you know I mean is basically all in you know I mean is basically all in you know I mean is there like uh interesting patterns like there like uh interesting patterns like there like uh interesting patterns like that that you've seen that the game does that that you've seen that the game does that that you've seen that the game does like you you'll have like Alpha zero like you you'll have like Alpha zero like you you'll have like Alpha zero doing way more sacrifices than humans doing way more sacrifices than humans doing way more sacrifices than humans usually do is there something like the usually do is there something like the usually do is there something like the bratis was constantly re-raising or bratis was constantly re-raising or bratis was constantly re-raising or something like that even noticed there something like that even noticed there something like that even noticed there was there was something really was there was something really was there was something really interesting that we observed with the interesting that we observed with the interesting that we observed with the broadest broadest broadest um so um so um so humans when they're playing poker they humans when they're playing poker they humans when they're playing poker they usually size their bets relative to the usually size their bets relative to the usually size their bets relative to the size of the pot so you know if the pot size of the pot so you know if the pot size of the pot so you know if the pot has a hundred dollars in there maybe you has a hundred dollars in there maybe you has a hundred dollars in there maybe you bet like 75 or somewhere around there bet like 75 or somewhere around there bet like 75 or somewhere around there somewhere between like 50 and 100 somewhere between like 50 and 100 somewhere between like 50 and 100 um and with libratus we gave it the um and with libratus we gave it the um and with libratus we gave it the option to basically bet whatever it option to basically bet whatever it option to basically bet whatever it wanted it was actually really easy for wanted it was actually really easy for wanted it was actually really easy for us to say like oh if you want you can
-
us to say like oh if you want you can us to say like oh if you want you can bet like 10 times the pot and we didn't bet like 10 times the pot and we didn't bet like 10 times the pot and we didn't think it would actually do that it was think it would actually do that it was think it would actually do that it was just like why not give it the option and just like why not give it the option and just like why not give it the option and then during the competition it actually then during the competition it actually then during the competition it actually started doing this and by the way this started doing this and by the way this started doing this and by the way this is like a very last minute decision on is like a very last minute decision on is like a very last minute decision on our part to add this option and so we our part to add this option and so we our part to add this option and so we did not we did not think the bot would did not we did not think the bot would did not we did not think the bot would would do this and uh I was actually kind would do this and uh I was actually kind would do this and uh I was actually kind of worried when it did start to do this of worried when it did start to do this of worried when it did start to do this like oh is this is a problem like humans like oh is this is a problem like humans like oh is this is a problem like humans don't do this like is it screwing up don't do this like is it screwing up don't do this like is it screwing up um but it would put the humans into um but it would put the humans into um but it would put the humans into really difficult spots when it would do really difficult spots when it would do really difficult spots when it would do that that that because you know you can imagine like because you know you can imagine like because you know you can imagine like you have the second best hand that's you have the second best hand that's you have the second best hand that's possible given the board and you're possible given the board and you're possible given the board and you're thinking like oh you're in a really thinking like oh you're in a really thinking like oh you're in a really great spot here and suddenly the bot great spot here and suddenly the bot great spot here and suddenly the bot bets twenty thousand dollars into a you bets twenty thousand dollars into a you bets twenty thousand dollars into a you know a thousand dollar pot and and it's know a thousand dollar pot and and it's know a thousand dollar pot and and it's basically saying like I have the best basically saying like I have the best basically saying like I have the best hand or I'm bluffing and you having the hand or I'm bluffing and you having the hand or I'm bluffing and you having the second best hand like now you get a second best hand like now you get a second best hand like now you get a really tough choice to make and so the really tough choice to make and so the really tough choice to make and so the humans would sometimes think like five humans would sometimes think like five humans would sometimes think like five or ten minutes about like what do you do or ten minutes about like what do you do or ten minutes about like what do you do should I call should I fold and um and should I call should I fold and um and should I call should I fold and um and when I saw the humans like really when I saw the humans like really when I saw the humans like really struggling with that decision like struggling with that decision like struggling with that decision like that's when I realized like oh actually that's when I realized like oh actually that's when I realized like oh actually this is maybe a good thing to do after this is maybe a good thing to do after this is maybe a good thing to do after all and of course the system doesn't all and of course the system doesn't all and of course the system doesn't know that it's making again like we said know that it's making again like we said know that it's making again like we said that it's putting them in a tough spot that it's putting them in a tough spot that it's putting them in a tough spot it's it's it's just that's part of the it's it's it's just that's part of the it's it's it's just that's part of the optimal the game theory optimal right optimal the game theory optimal right optimal the game theory optimal right from the Bots perspective it's just it's from the Bots perspective it's just it's from the Bots perspective it's just it's just doing the thing that's going to just doing the thing that's going to just doing the thing that's going to make it the most money make it the most money make it the most money um and the fact that it's putting the um and the fact that it's putting the um and the fact that it's putting the humans in a difficult spot like that's humans in a difficult spot like that's humans in a difficult spot like that's just um you know a side effect of that just um you know a side effect of that just um you know a side effect of that and this was I think the the one thing I and this was I think the the one thing I and this was I think the the one thing I mean there were a few things that the mean there were a few things that the mean there were a few things that the humans walked away from but this was the humans walked away from but this was the humans walked away from but this was the the number one thing that the humans
-
the number one thing that the humans the number one thing that the humans walked away from the competition saying walked away from the competition saying walked away from the competition saying like we need to start doing this like we need to start doing this like we need to start doing this um and now these over bats what are um and now these over bats what are um and now these over bats what are called over bets have become really called over bets have become really called over bets have become really common in high level poker play have you common in high level poker play have you common in high level poker play have you ever talked to like somebody like Danny ever talked to like somebody like Danny ever talked to like somebody like Danny on the ground about this he seems to be on the ground about this he seems to be on the ground about this he seems to be a student of the game I did actually a student of the game I did actually a student of the game I did actually have a conversation with Daniel degrania have a conversation with Daniel degrania have a conversation with Daniel degrania once yeah I was uh I was visiting the once yeah I was uh I was visiting the once yeah I was uh I was visiting the Isle of Man to talk to Poker Stars about Isle of Man to talk to Poker Stars about Isle of Man to talk to Poker Stars about AI AI AI um and Daniel legrandi was there when we um and Daniel legrandi was there when we um and Daniel legrandi was there when we had dinner together with uh some other had dinner together with uh some other had dinner together with uh some other people and um yeah he was really people and um yeah he was really people and um yeah he was really interested in it he mentioned that he interested in it he mentioned that he interested in it he mentioned that he was like you know excited about like was like you know excited about like was like you know excited about like learning from these AIS learning from these AIS learning from these AIS um so he wasn't scared he was excited he um so he wasn't scared he was excited he um so he wasn't scared he was excited he was excited and uh and he all he was excited and uh and he all he was excited and uh and he all he honestly he wanted to play against the honestly he wanted to play against the honestly he wanted to play against the bot he thought he thought he had a bot he thought he thought he had a bot he thought he thought he had a decent chance of beating it decent chance of beating it decent chance of beating it um I I think he's you know um I I think he's you know um I I think he's you know this was like several years ago and I this was like several years ago and I this was like several years ago and I think it was like not as clear to think it was like not as clear to think it was like not as clear to everybody that you know the AIS were everybody that you know the AIS were everybody that you know the AIS were taking over I think now people recognize taking over I think now people recognize taking over I think now people recognize that like if you're playing against uh a that like if you're playing against uh a that like if you're playing against uh a bot there's like no chance that you have bot there's like no chance that you have bot there's like no chance that you have in a game like Pokemon so consistently in a game like Pokemon so consistently in a game like Pokemon so consistently the Bots will win the Bots have heads up the Bots will win the Bots have heads up the Bots will win the Bots have heads up and in in other variants too so multi and in in other variants too so multi and in in other variants too so multi multi six player Texas Hold'em No Limit multi six player Texas Hold'em No Limit multi six player Texas Hold'em No Limit taxes hold them as the Bots win yeah taxes hold them as the Bots win yeah taxes hold them as the Bots win yeah that's the case so I think there's some that's the case so I think there's some that's the case so I think there's some debate about like is it true for every debate about like is it true for every debate about like is it true for every single variant of Poker I think I think single variant of Poker I think I think single variant of Poker I think I think for every single variant of Poker if for every single variant of Poker if for every single variant of Poker if somebody really put in the effort they somebody really put in the effort they somebody really put in the effort they can make an AI that would beat all can make an AI that would beat all can make an AI that would beat all humans at it humans at it humans at it um we've focused on the most popular um we've focused on the most popular um we've focused on the most popular variants so heads up no limit Texas variants so heads up no limit Texas variants so heads up no limit Texas Hold'em and then we followed it up with
-
Hold'em and then we followed it up with Hold'em and then we followed it up with um with uh six player poker as well um with uh six player poker as well um with uh six player poker as well where we managed to uh make a bot that where we managed to uh make a bot that where we managed to uh make a bot that beat expert human players and I think beat expert human players and I think beat expert human players and I think even there now uh it's pretty clear that even there now uh it's pretty clear that even there now uh it's pretty clear that humans don't stand a chance see I would humans don't stand a chance see I would humans don't stand a chance see I would love to hook up an AI system that looks love to hook up an AI system that looks love to hook up an AI system that looks at EEG at EEG at EEG like how like actually tries to optimize like how like actually tries to optimize like how like actually tries to optimize the toughness of the spot it puts a the toughness of the spot it puts a the toughness of the spot it puts a human in and I I would I would love to human in and I I would I would love to human in and I I would I would love to see how different is that from the game see how different is that from the game see how different is that from the game theory optimal so you try to maximize theory optimal so you try to maximize theory optimal so you try to maximize the heart rate of the human player like the heart rate of the human player like the heart rate of the human player like the freaking out over a long period of the freaking out over a long period of the freaking out over a long period of time I wonder if there's going to be time I wonder if there's going to be time I wonder if there's going to be different strategies that emerge uh that different strategies that emerge uh that different strategies that emerge uh that are close in terms of Effectiveness are close in terms of Effectiveness are close in terms of Effectiveness because something tells me you could because something tells me you could because something tells me you could still be still be still be um achieved superhuman level performance um achieved superhuman level performance um achieved superhuman level performance by just making people sweat by just making people sweat by just making people sweat I feel like that there's a good chance I feel like that there's a good chance I feel like that there's a good chance that that is the case yeah if you're that that is the case yeah if you're that that is the case yeah if you're able to see like that it's like it's able to see like that it's like it's able to see like that it's like it's like a decent proxy for score right like a decent proxy for score right like a decent proxy for score right right um and this is actually like the right um and this is actually like the right um and this is actually like the the common poker wisdom when they're the common poker wisdom when they're the common poker wisdom when they're telling where they're teaching players telling where they're teaching players telling where they're teaching players before the robots and they were trying before the robots and they were trying before the robots and they were trying to teach people how to play poker they to teach people how to play poker they to teach people how to play poker they would say like the key to the game is to would say like the key to the game is to would say like the key to the game is to put your opponent into difficult spots put your opponent into difficult spots put your opponent into difficult spots it's a good um a good estimate for if it's a good um a good estimate for if it's a good um a good estimate for if you're making the right decision so what you're making the right decision so what you're making the right decision so what else can you say about the fundamental else can you say about the fundamental else can you say about the fundamental role of search in poker and maybe if you role of search in poker and maybe if you role of search in poker and maybe if you can also relate it to chess and go in can also relate it to chess and go in can also relate it to chess and go in these games these games these games um um um what's the role of search to solve in what's the role of search to solve in what's the role of search to solve in these games these games these games yeah I think a lot of people under this yeah I think a lot of people under this yeah I think a lot of people under this is true for the general public and I
-
is true for the general public and I is true for the general public and I think it's true for the AI Community a think it's true for the AI Community a think it's true for the AI Community a lot of people underestimate the lot of people underestimate the lot of people underestimate the importance of search for these kinds of importance of search for these kinds of importance of search for these kinds of game AI results game AI results game AI results um an example of this is uh TD Gammon um an example of this is uh TD Gammon um an example of this is uh TD Gammon that came out in 1992 this was the the that came out in 1992 this was the the that came out in 1992 this was the the first real instance of a neural net first real instance of a neural net first real instance of a neural net being used in a game AI it's a landmark being used in a game AI it's a landmark being used in a game AI it's a landmark achievement it was actually the achievement it was actually the achievement it was actually the inspiration for Alpha zero and it used inspiration for Alpha zero and it used inspiration for Alpha zero and it used search it used two-ply search to figure search it used two-ply search to figure search it used two-ply search to figure out its next move out its next move out its next move you got deep blue there he was very you got deep blue there he was very you got deep blue there he was very heavily focused on search heavily focused on search heavily focused on search um looking many many moves ahead farther um looking many many moves ahead farther um looking many many moves ahead farther than any human could and that was key than any human could and that was key than any human could and that was key for why it won and then even with for why it won and then even with for why it won and then even with something like alphago I mean alphago is something like alphago I mean alphago is something like alphago I mean alphago is commonly hailed as a landmark commonly hailed as a landmark commonly hailed as a landmark achievement for neural Nets and it is achievement for neural Nets and it is achievement for neural Nets and it is but there's also this huge component of but there's also this huge component of but there's also this huge component of search Monte Carlo tree search to search Monte Carlo tree search to search Monte Carlo tree search to alphago that was key absolutely alphago that was key absolutely alphago that was key absolutely essential for the AI to be able to beat essential for the AI to be able to beat essential for the AI to be able to beat top humans top humans top humans um I think a good example of this is you um I think a good example of this is you um I think a good example of this is you look at the latest versions of alpha of look at the latest versions of alpha of look at the latest versions of alpha of alphago like it was called Alpha zero alphago like it was called Alpha zero alphago like it was called Alpha zero um and there's this metric called ELO um and there's this metric called ELO um and there's this metric called ELO rating where you can compare different rating where you can compare different rating where you can compare different humans and you can compare Bots to humans and you can compare Bots to humans and you can compare Bots to humans now a top human player is around humans now a top human player is around humans now a top human player is around 3600 ELO maybe a little bit higher now 3600 ELO maybe a little bit higher now 3600 ELO maybe a little bit higher now um Alpha zero the strongest version is um Alpha zero the strongest version is um Alpha zero the strongest version is around 5200 ELO around 5200 ELO around 5200 ELO but if you take out the search that's but if you take out the search that's but if you take out the search that's being done at test time and by the way being done at test time and by the way being done at test time and by the way what I mean by search is the planning what I mean by search is the planning what I mean by search is the planning ahead the thinking of like oh if I move ahead the thinking of like oh if I move ahead the thinking of like oh if I move my if I place the stone here and then he my if I place the stone here and then he my if I place the stone here and then he does this and then you look like five
-
does this and then you look like five does this and then you look like five moves ahead and you see like what the moves ahead and you see like what the moves ahead and you see like what the board state looks like board state looks like board state looks like um that's what I mean by search if you um that's what I mean by search if you um that's what I mean by search if you take out the search that's done during take out the search that's done during take out the search that's done during the game the ELO rating drops to around the game the ELO rating drops to around the game the ELO rating drops to around three thousand three thousand three thousand so even today so even today so even today what seven years after alphago what seven years after alphago what seven years after alphago if you take out the Monte Carlo research if you take out the Monte Carlo research if you take out the Monte Carlo research that's being done at one playing against that's being done at one playing against that's being done at one playing against the human the human the human the Bots are not superhuman nobody has the Bots are not superhuman nobody has the Bots are not superhuman nobody has made a raw neural net that is superhuman made a raw neural net that is superhuman made a raw neural net that is superhuman and go and go and go that's worth lingering on that's that's that's worth lingering on that's that's that's worth lingering on that's that's quite profound quite profound quite profound so without search that just means so without search that just means so without search that just means looking at the next move looking at the next move looking at the next move and saying this is the best move so and saying this is the best move so and saying this is the best move so having a function that estimates having a function that estimates having a function that estimates accurately what the best move is that's accurately what the best move is that's accurately what the best move is that's right without search yeah and all these right without search yeah and all these right without search yeah and all these Bots they have the what's called a Bots they have the what's called a Bots they have the what's called a policy Network where it will tell you policy Network where it will tell you policy Network where it will tell you this is what the neural net thinks is this is what the neural net thinks is this is what the neural net thinks is the next best move the next best move the next best move um um um and it's kind of like a the intuition and it's kind of like a the intuition and it's kind of like a the intuition that a human has you know the human that a human has you know the human that a human has you know the human looks at the board and and any uh go or looks at the board and and any uh go or looks at the board and and any uh go or chess master will be able to tell you chess master will be able to tell you chess master will be able to tell you like oh instantly here's what I think like oh instantly here's what I think like oh instantly here's what I think the right move is the right move is the right move is um and the bot is able to do the same um and the bot is able to do the same um and the bot is able to do the same thing but just like how a human thing but just like how a human thing but just like how a human Grandmaster can make a better decision Grandmaster can make a better decision Grandmaster can make a better decision if they have more time to think when you if they have more time to think when you if they have more time to think when you add on this Monte Carlo tree search the add on this Monte Carlo tree search the add on this Monte Carlo tree search the bot is able to make a better decision bot is able to make a better decision bot is able to make a better decision yeah I mean of course a human is doing yeah I mean of course a human is doing yeah I mean of course a human is doing something like searching their brain but something like searching their brain but something like searching their brain but it's not it's not it's not I hesitate to draw a hard line but it's I hesitate to draw a hard line but it's I hesitate to draw a hard line but it's not like uh Monte Carlo tree search it's not like uh Monte Carlo tree search it's not like uh Monte Carlo tree search it's more like
-
more like more like sequential language model generation so sequential language model generation so sequential language model generation so it's like a different it's a the neural it's like a different it's a the neural it's like a different it's a the neural network is doing the searching and I network is doing the searching and I network is doing the searching and I wonder what the human brain is doing in wonder what the human brain is doing in wonder what the human brain is doing in terms of searching because you're doing terms of searching because you're doing terms of searching because you're doing that like computation the human is that like computation the human is that like computation the human is Computing they have intuition they've Computing they have intuition they've Computing they have intuition they've got got got they have a really strong ability to they have a really strong ability to they have a really strong ability to estimate you know amongst the top estimate you know amongst the top estimate you know amongst the top players of what is a good and not players of what is a good and not players of what is a good and not position without calculating all the position without calculating all the position without calculating all the details details details but they're still doing search in their but they're still doing search in their but they're still doing search in their head but it's a different kind of search head but it's a different kind of search head but it's a different kind of search have you ever thought about like what is have you ever thought about like what is have you ever thought about like what is the difference between the human the difference between the human the difference between the human the search that the human is performing the search that the human is performing the search that the human is performing versus what versus what versus what computers are doing I have thought a lot computers are doing I have thought a lot computers are doing I have thought a lot about that and I think it's a really about that and I think it's a really about that and I think it's a really important question so the AI in Alpha important question so the AI in Alpha important question so the AI in Alpha and Alphas in alphago or any of these go and Alphas in alphago or any of these go and Alphas in alphago or any of these go AIS they're all doing Monte College AIS they're all doing Monte College AIS they're all doing Monte College research which is a particular kind of research which is a particular kind of research which is a particular kind of search and it it's actually a symbolic search and it it's actually a symbolic search and it it's actually a symbolic tabular search it uses the neural net to tabular search it uses the neural net to tabular search it uses the neural net to guide its search but it isn't actually guide its search but it isn't actually guide its search but it isn't actually like full full-on neural net like full full-on neural net like full full-on neural net now that kind of search is very now that kind of search is very now that kind of search is very successful in these kinds of like successful in these kinds of like successful in these kinds of like perfect information board games like perfect information board games like perfect information board games like chess and go but if you take it to a chess and go but if you take it to a chess and go but if you take it to a game like poker for example it doesn't game like poker for example it doesn't game like poker for example it doesn't work it can't it can't understand the work it can't it can't understand the work it can't it can't understand the concept of hidden information it doesn't concept of hidden information it doesn't concept of hidden information it doesn't understand the balance that you have to understand the balance that you have to understand the balance that you have to strike between like the amount that strike between like the amount that strike between like the amount that you're raising versus the amount that you're raising versus the amount that you're raising versus the amount that you're calling and in every one of these you're calling and in every one of these you're calling and in every one of these games you see a different kind of search games you see a different kind of search games you see a different kind of search and the human brain is able to plan for and the human brain is able to plan for and the human brain is able to plan for all these different games in a very all these different games in a very all these different games in a very general way now I think that's one thing general way now I think that's one thing general way now I think that's one thing that we're missing from AI today and I
-
that we're missing from AI today and I that we're missing from AI today and I think it's a really important missing think it's a really important missing think it's a really important missing piece the ability to plan and reason piece the ability to plan and reason piece the ability to plan and reason more generally across a wide variety of more generally across a wide variety of more generally across a wide variety of different settings different settings different settings in a way where the general reasoning in a way where the general reasoning in a way where the general reasoning makes you better at each one of the makes you better at each one of the makes you better at each one of the games not worse yeah so you can kind of games not worse yeah so you can kind of games not worse yeah so you can kind of think of it as like neural Nets today think of it as like neural Nets today think of it as like neural Nets today they'll give you like Transformers for they'll give you like Transformers for they'll give you like Transformers for example or super General but you know example or super General but you know example or super General but you know they'll give you it'll output an answer they'll give you it'll output an answer they'll give you it'll output an answer in like 100 milliseconds and if you tell in like 100 milliseconds and if you tell in like 100 milliseconds and if you tell it like oh you've got five minutes to it like oh you've got five minutes to it like oh you've got five minutes to give you a decision you know feel free give you a decision you know feel free give you a decision you know feel free to take more time to make a better to take more time to make a better to take more time to make a better decision it's not gonna know what to do decision it's not gonna know what to do decision it's not gonna know what to do with that with that with that um but a human if you're playing a game um but a human if you're playing a game um but a human if you're playing a game like chess they're going to give you a like chess they're going to give you a like chess they're going to give you a very different answer depending on if very different answer depending on if very different answer depending on if you say oh you've got 100 milliseconds you say oh you've got 100 milliseconds you say oh you've got 100 milliseconds or you've got five minutes or you've got five minutes or you've got five minutes yeah there I mean that people have yeah there I mean that people have yeah there I mean that people have started using right Transformers the started using right Transformers the started using right Transformers the language models like the in an iterative language models like the in an iterative language models like the in an iterative way that does improve the answer or like way that does improve the answer or like way that does improve the answer or like showing the work the kind of kind of showing the work the kind of kind of showing the work the kind of kind of idea yeah they got this thing called idea yeah they got this thing called idea yeah they got this thing called Chain of Thought reasoning and that's I Chain of Thought reasoning and that's I Chain of Thought reasoning and that's I think um super promising right yeah I think um super promising right yeah I think um super promising right yeah I think and I think it's a good step in think and I think it's a good step in think and I think it's a good step in the right direction the right direction the right direction um I I would kind of like say it's um I I would kind of like say it's um I I would kind of like say it's similar to Monte Carlo rollouts in in a similar to Monte Carlo rollouts in in a similar to Monte Carlo rollouts in in a game like chess there's a kind of search game like chess there's a kind of search game like chess there's a kind of search that you can do where you're saying like that you can do where you're saying like that you can do where you're saying like I'm Gonna Roll Out My intuition and see I'm Gonna Roll Out My intuition and see I'm Gonna Roll Out My intuition and see like without really thinking you know like without really thinking you know like without really thinking you know what are the better decisions I can make what are the better decisions I can make what are the better decisions I can make farther down the path farther down the path farther down the path um what would I do if I just acted um what would I do if I just acted um what would I do if I just acted according to intuition for the next 10 according to intuition for the next 10 according to intuition for the next 10 moves moves moves um and that gets you an improvement but um and that gets you an improvement but um and that gets you an improvement but I think that there's much uh much richer I think that there's much uh much richer I think that there's much uh much richer kinds of of planning that we could do kinds of of planning that we could do kinds of of planning that we could do so when the broadest actually beat the so when the broadest actually beat the so when the broadest actually beat the poker plays what did I feel like what
-
poker plays what did I feel like what poker plays what did I feel like what was that I mean actually on that day was that I mean actually on that day was that I mean actually on that day what were you feeling like were you were what were you feeling like were you were what were you feeling like were you were you nervous you nervous you nervous I mean Poco was one of the games that he I mean Poco was one of the games that he I mean Poco was one of the games that he thought like is not going to be solvable thought like is not going to be solvable thought like is not going to be solvable because it's the human factor so at because it's the human factor so at because it's the human factor so at least in the narratives we tell least in the narratives we tell least in the narratives we tell ourselves the human factor so ourselves the human factor so ourselves the human factor so fundamental to the game of poker fundamental to the game of poker fundamental to the game of poker yeah the liberatis competition was super yeah the liberatis competition was super yeah the liberatis competition was super stressful for me stressful for me stressful for me um also I mean I was working on this um also I mean I was working on this um also I mean I was working on this like basically continuously for a year like basically continuously for a year like basically continuously for a year leading up to the competition I mean for leading up to the competition I mean for leading up to the competition I mean for me it became like very clear like okay me it became like very clear like okay me it became like very clear like okay this is the search technique this is the this is the search technique this is the this is the search technique this is the approach that we need and then I spent a approach that we need and then I spent a approach that we need and then I spent a year working on this pretty much like year working on this pretty much like year working on this pretty much like non-stop oh can we actually get into non-stop oh can we actually get into non-stop oh can we actually get into details like what programming language details like what programming language details like what programming language is it written in what's some interesting is it written in what's some interesting is it written in what's some interesting uh implementation details that are like uh implementation details that are like uh implementation details that are like fun slash painful yeah so one of the fun slash painful yeah so one of the fun slash painful yeah so one of the interesting things about liberatis is interesting things about liberatis is interesting things about liberatis is that we had no idea what the bar was to that we had no idea what the bar was to that we had no idea what the bar was to actually beat top humans yeah we could actually beat top humans yeah we could actually beat top humans yeah we could play against like our prior Bots and play against like our prior Bots and play against like our prior Bots and that kind of gives us some sense of like that kind of gives us some sense of like that kind of gives us some sense of like are we making progress are we going in are we making progress are we going in are we making progress are we going in the right direction uh but we had no the right direction uh but we had no the right direction uh but we had no idea like what the bar actually was and idea like what the bar actually was and idea like what the bar actually was and so we threw a huge amount of resources so we threw a huge amount of resources so we threw a huge amount of resources at trying to make the strongest bot at trying to make the strongest bot at trying to make the strongest bot possible so we use C plus plus it was possible so we use C plus plus it was possible so we use C plus plus it was parallelized we were using I think like parallelized we were using I think like parallelized we were using I think like a thousand CPUs uh maybe maybe more a thousand CPUs uh maybe maybe more a thousand CPUs uh maybe maybe more actually actually actually um and you know today that sounds like um and you know today that sounds like um and you know today that sounds like nothing but for a grad student back in nothing but for a grad student back in nothing but for a grad student back in 2016 that was a huge amount of resources 2016 that was a huge amount of resources 2016 that was a huge amount of resources but still a lot for even any gratitude but still a lot for even any gratitude but still a lot for even any gratitude today it's still tough to to get today it's still tough to to get today it's still tough to to get or even to allow yourself to think in or even to allow yourself to think in or even to allow yourself to think in that in terms of scale at CMU at MIT
-
that in terms of scale at CMU at MIT that in terms of scale at CMU at MIT anything like that yeah and you know anything like that yeah and you know anything like that yeah and you know talking about terabytes of memory talking about terabytes of memory talking about terabytes of memory um so it's a very paralyzed um so it's a very paralyzed um so it's a very paralyzed um and it had to be very fast too um and it had to be very fast too um and it had to be very fast too because the more games that you could because the more games that you could because the more games that you could simulate uh the stronger the bot would simulate uh the stronger the bot would simulate uh the stronger the bot would be so is there some like John Carmack be so is there some like John Carmack be so is there some like John Carmack Style Style Style like efficiencies you have to come up like efficiencies you have to come up like efficiencies you have to come up with like an efficient way to represent with like an efficient way to represent with like an efficient way to represent a hand all that kind of stuff there were a hand all that kind of stuff there were a hand all that kind of stuff there were all sorts of optimizations that I had to all sorts of optimizations that I had to all sorts of optimizations that I had to make to try to get this thing to run as make to try to get this thing to run as make to try to get this thing to run as fast as possible they were like how do fast as possible they were like how do fast as possible they were like how do you minimize the latency how do you like you minimize the latency how do you like you minimize the latency how do you like you know package things together so that you know package things together so that you know package things together so that like you minimize the amount of like you minimize the amount of like you minimize the amount of communication between the different communication between the different communication between the different nodes nodes nodes um how do you like optimize the um how do you like optimize the um how do you like optimize the algorithm so that you can you know try algorithm so that you can you know try algorithm so that you can you know try to squeeze out more and more from the to squeeze out more and more from the to squeeze out more and more from the game that you're actually playing all game that you're actually playing all game that you're actually playing all these kinds of different decisions that these kinds of different decisions that these kinds of different decisions that that I you know had to make uh just a that I you know had to make uh just a that I you know had to make uh just a fun question what what id did you use fun question what what id did you use fun question what what id did you use what uh for for C plus plus what uh for for C plus plus what uh for for C plus plus I think I used a visual studio actually I think I used a visual studio actually I think I used a visual studio actually yeah okay yeah is that still carried yeah okay yeah is that still carried yeah okay yeah is that still carried through to today vs code is is what I through to today vs code is is what I through to today vs code is is what I use today it seems like it's pretty the use today it seems like it's pretty the use today it seems like it's pretty the community basically converged on Okay community basically converged on Okay community basically converged on Okay cool so you got you got this cool so you got you got this cool so you got you got this super optimized C plus plus system super optimized C plus plus system super optimized C plus plus system and then you show up to the day of and then you show up to the day of and then you show up to the day of competition competition competition yeah yeah yeah humans versus machine humans versus machine humans versus machine um how did it feel throughout the day um how did it feel throughout the day um how did it feel throughout the day super stressful super stressful super stressful um I mean I thought going into it that um I mean I thought going into it that um I mean I thought going into it that we had like a 50 50 chance because we had like a 50 50 chance because we had like a 50 50 chance because basically I thought if if they play in a basically I thought if if they play in a basically I thought if if they play in a totally normal style I think we'll totally normal style I think we'll totally normal style I think we'll squeak out a win but there's always a squeak out a win but there's always a squeak out a win but there's always a chance that they can find some weakness
-
chance that they can find some weakness chance that they can find some weakness in the bot and if they do and we're in the bot and if they do and we're in the bot and if they do and we're playing like for 20 days 120 000 hands playing like for 20 days 120 000 hands playing like for 20 days 120 000 hands of poker they have a lot of time to find of poker they have a lot of time to find of poker they have a lot of time to find weaknesses in the system and if they do weaknesses in the system and if they do weaknesses in the system and if they do we're gonna get crushed and that's we're gonna get crushed and that's we're gonna get crushed and that's actually what happened in the previous actually what happened in the previous actually what happened in the previous competition competition competition um the humans you know they started out um the humans you know they started out um the humans you know they started out it wasn't like a they were winning from it wasn't like a they were winning from it wasn't like a they were winning from the start but then they found these the start but then they found these the start but then they found these weaknesses that they could take weaknesses that they could take weaknesses that they could take advantage of and for the next they know advantage of and for the next they know advantage of and for the next they know like 10 days they were just just like 10 days they were just just like 10 days they were just just crushing the bot stealing money from it crushing the bot stealing money from it crushing the bot stealing money from it what were the weaknesses they found like what were the weaknesses they found like what were the weaknesses they found like maybe over betting was effective that maybe over betting was effective that maybe over betting was effective that kind of stuff so certain betting kind of stuff so certain betting kind of stuff so certain betting strategies worked what they found is strategies worked what they found is strategies worked what they found is yeah over betting like betting certain yeah over betting like betting certain yeah over betting like betting certain amounts the bot would have a lot of amounts the bot would have a lot of amounts the bot would have a lot of trouble dealing with those sizes and trouble dealing with those sizes and trouble dealing with those sizes and then also then also then also um when it the bot got into really um when it the bot got into really um when it the bot got into really difficult all-in situations it wasn't difficult all-in situations it wasn't difficult all-in situations it wasn't able to because it wasn't doing search able to because it wasn't doing search able to because it wasn't doing search it had to Clump different hands together it had to Clump different hands together it had to Clump different hands together and it wouldn't it would treat them and it wouldn't it would treat them and it wouldn't it would treat them identically yeah um and so it wouldn't identically yeah um and so it wouldn't identically yeah um and so it wouldn't be able to distinguish you know like be able to distinguish you know like be able to distinguish you know like having a king High flush versus an ace having a king High flush versus an ace having a king High flush versus an ace high flush and in some situations that high flush and in some situations that high flush and in some situations that really matters a lot and so they could really matters a lot and so they could really matters a lot and so they could put the bot into those situations and put the bot into those situations and put the bot into those situations and then the bot would would just bleed then the bot would would just bleed then the bot would would just bleed money clever humans yeah okay so I money clever humans yeah okay so I money clever humans yeah okay so I didn't realize it was over 20 days so didn't realize it was over 20 days so didn't realize it was over 20 days so um um um what were the humans like over those 20 what were the humans like over those 20 what were the humans like over those 20 days and what was the bot like so we had days and what was the bot like so we had days and what was the bot like so we had set up the competition you know like I set up the competition you know like I set up the competition you know like I said there was two hundred thousand said there was two hundred thousand said there was two hundred thousand dollars in prize money and they would dollars in prize money and they would dollars in prize money and they would get paid a fraction of that depending on get paid a fraction of that depending on get paid a fraction of that depending on how well they did relative to each other how well they did relative to each other how well they did relative to each other yeah so I was kind of hoping that they yeah so I was kind of hoping that they yeah so I was kind of hoping that they wouldn't work together to try to find wouldn't work together to try to find wouldn't work together to try to find weaknesses in the bot but they entered
-
weaknesses in the bot but they entered weaknesses in the bot but they entered the competition with their like number the competition with their like number the competition with their like number one objective being to beat the bot and one objective being to beat the bot and one objective being to beat the bot and they didn't care about like individual they didn't care about like individual they didn't care about like individual Glory they were like we're all going to Glory they were like we're all going to Glory they were like we're all going to work as a team to try to take down the work as a team to try to take down the work as a team to try to take down the spot yeah and so they immediately spot yeah and so they immediately spot yeah and so they immediately started comparing notes what they would started comparing notes what they would started comparing notes what they would do is they would coordinate looking at do is they would coordinate looking at do is they would coordinate looking at different parts of the strategy to try different parts of the strategy to try different parts of the strategy to try to try to you know find out weaknesses to try to you know find out weaknesses to try to you know find out weaknesses um and then at the end of the day we um and then at the end of the day we um and then at the end of the day we actually sent them a log of all the actually sent them a log of all the actually sent them a log of all the hands that were played and what cards hands that were played and what cards hands that were played and what cards the bot had on each of those hands oh the bot had on each of those hands oh the bot had on each of those hands oh wow yeah that's that's gutsy yeah it was wow yeah that's that's gutsy yeah it was wow yeah that's that's gutsy yeah it was honestly and I'm not sure why we did honestly and I'm not sure why we did honestly and I'm not sure why we did that in retrospect but um I mean I'm that in retrospect but um I mean I'm that in retrospect but um I mean I'm glad we did it because we ended up glad we did it because we ended up glad we did it because we ended up winning anyway but that if if you've winning anyway but that if if you've winning anyway but that if if you've ever played poker before like that is ever played poker before like that is ever played poker before like that is golden information I mean to know golden information I mean to know golden information I mean to know usually when you play poker you see usually when you play poker you see usually when you play poker you see about a third of the hands to Showdown about a third of the hands to Showdown about a third of the hands to Showdown um and to just hand them all the cards um and to just hand them all the cards um and to just hand them all the cards that the bot had on every single hand that the bot had on every single hand that the bot had on every single hand that was that was that was um just just a gold mine for them yeah um just just a gold mine for them yeah um just just a gold mine for them yeah and so then they would review the hands and so then they would review the hands and so then they would review the hands and try to see like okay could they find and try to see like okay could they find and try to see like okay could they find patterns in the bot the weaknesses and patterns in the bot the weaknesses and patterns in the bot the weaknesses and could they then then they would could they then then they would could they then then they would coordinate and study together and try to coordinate and study together and try to coordinate and study together and try to figure out okay now this person's gonna figure out okay now this person's gonna figure out okay now this person's gonna explore this part of the strategy for explore this part of the strategy for explore this part of the strategy for weaknesses this person's gonna explore weaknesses this person's gonna explore weaknesses this person's gonna explore this part of the strategy for weaknesses this part of the strategy for weaknesses this part of the strategy for weaknesses it's a kind of psychological warfare it's a kind of psychological warfare it's a kind of psychological warfare showing in the hands yeah showing in the hands yeah showing in the hands yeah um I mean I'm sure you didn't think of um I mean I'm sure you didn't think of um I mean I'm sure you didn't think of it that way but like doing that means it that way but like doing that means it that way but like doing that means you're confident in the possibility to you're confident in the possibility to you're confident in the possibility to win well that's that's one way of win well that's that's one way of win well that's that's one way of putting it I wasn't uh super confident putting it I wasn't uh super confident putting it I wasn't uh super confident yeah so yeah so yeah so you know going in like I said I think I you know going in like I said I think I you know going in like I said I think I had like 50 50 odds on us winning the had like 50 50 odds on us winning the had like 50 50 odds on us winning the when we actually when we announced the when we actually when we announced the when we actually when we announced the competition the poker Community decided
-
competition the poker Community decided competition the poker Community decided to gamble on who would win and their to gamble on who would win and their to gamble on who would win and their initial odds against us were like four initial odds against us were like four initial odds against us were like four to one they were really convinced that to one they were really convinced that to one they were really convinced that the humans were gonna the humans were gonna the humans were gonna pull out a win pull out a win pull out a win um the bot ended up winning for three um the bot ended up winning for three um the bot ended up winning for three days straight and even then after three days straight and even then after three days straight and even then after three days the betting odds were still just 50 days the betting odds were still just 50 days the betting odds were still just 50 50. 50. 50. um um um and then at that point it started to and then at that point it started to and then at that point it started to look like the humans were coming back look like the humans were coming back look like the humans were coming back um they started to like you know but but um they started to like you know but but um they started to like you know but but poker is a very high variance game poker is a very high variance game poker is a very high variance game um and I think what happened is like um and I think what happened is like um and I think what happened is like they thought that they spotted some they thought that they spotted some they thought that they spotted some weaknesses that weren't actually there weaknesses that weren't actually there weaknesses that weren't actually there and then around day eight it was just and then around day eight it was just and then around day eight it was just very clear that they were getting very clear that they were getting very clear that they were getting absolutely crushed absolutely crushed absolutely crushed um and and from that point I mean for um and and from that point I mean for um and and from that point I mean for for a while there I was super stressed for a while there I was super stressed for a while there I was super stressed out thinking like oh my God the humans out thinking like oh my God the humans out thinking like oh my God the humans are coming back and we're just they've are coming back and we're just they've are coming back and we're just they've found weaknesses and now we're just found weaknesses and now we're just found weaknesses and now we're just gonna lose the whole thing but no it gonna lose the whole thing but no it gonna lose the whole thing but no it ended up going in the other direction ended up going in the other direction ended up going in the other direction and the bot ended up like crushing them and the bot ended up like crushing them and the bot ended up like crushing them in the long run in the long run in the long run how did it uh feel at the end like as a how did it uh feel at the end like as a how did it uh feel at the end like as a human being what it as a person who human being what it as a person who human being what it as a person who loves appreciates the beauty of the Game loves appreciates the beauty of the Game loves appreciates the beauty of the Game of Poker and the person who appreciates of Poker and the person who appreciates of Poker and the person who appreciates the beauty of AI is there did you feel a the beauty of AI is there did you feel a the beauty of AI is there did you feel a certain kind of way about it certain kind of way about it certain kind of way about it uh I felt a lot of a lot of things man uh I felt a lot of a lot of things man uh I felt a lot of a lot of things man um I mean at that point in my life I had um I mean at that point in my life I had um I mean at that point in my life I had spent five years working on this project spent five years working on this project spent five years working on this project and um it was a huge sense of and um it was a huge sense of and um it was a huge sense of accomplishment I mean to spend five accomplishment I mean to spend five accomplishment I mean to spend five years working on something and finally years working on something and finally years working on something and finally see it succeed see it succeed see it succeed um yeah I wouldn't trade that for um yeah I wouldn't trade that for um yeah I wouldn't trade that for anything in the world yeah because it's anything in the world yeah because it's anything in the world yeah because it's uh that's a real Benchmark it's not like uh that's a real Benchmark it's not like uh that's a real Benchmark it's not like uh getting us some percent accuracy and
-
uh getting us some percent accuracy and uh getting us some percent accuracy and a data set this is like real this is a data set this is like real this is a data set this is like real this is real world it's it's just a game but real world it's it's just a game but real world it's it's just a game but it's also a game it means a lot to a lot it's also a game it means a lot to a lot it's also a game it means a lot to a lot of people and this is humans doing their of people and this is humans doing their of people and this is humans doing their best to beat the machine so this is a best to beat the machine so this is a best to beat the machine so this is a real Benchmark unlike anything else yeah real Benchmark unlike anything else yeah real Benchmark unlike anything else yeah and I mean this is this is what I have and I mean this is this is what I have and I mean this is this is what I have been dreaming about since I was like 16 been dreaming about since I was like 16 been dreaming about since I was like 16 playing poker you know with my friends playing poker you know with my friends playing poker you know with my friends in high school the idea that you could in high school the idea that you could in high school the idea that you could find a strategy find a strategy find a strategy um you know approximate the national um you know approximate the national um you know approximate the national equilibrium be able to beat all the equilibrium be able to beat all the equilibrium be able to beat all the poker players in the world with it you poker players in the world with it you poker players in the world with it you know so to actually see that come to know so to actually see that come to know so to actually see that come to fruition and be realized uh that was fruition and be realized uh that was fruition and be realized uh that was is kind of magical is kind of magical is kind of magical yeah especially money is on the line too yeah especially money is on the line too yeah especially money is on the line too it's a different it's different than it's a different it's different than it's a different it's different than chess chess chess and that aspect like people get that's and that aspect like people get that's and that aspect like people get that's why you want to look at Betty Marcus if why you want to look at Betty Marcus if why you want to look at Betty Marcus if you want to actually understand what you want to actually understand what you want to actually understand what people really think in the same sense people really think in the same sense people really think in the same sense poker it's really high stakes because poker it's really high stakes because poker it's really high stakes because it's money and to solve that game that's it's money and to solve that game that's it's money and to solve that game that's that's an amazing accomplishment so the that's an amazing accomplishment so the that's an amazing accomplishment so the leap from that to leap from that to leap from that to multi-way six player poker what's how multi-way six player poker what's how multi-way six player poker what's how difficult does that jump difficult does that jump difficult does that jump and what are some interesting and what are some interesting and what are some interesting differences between heads up poker and differences between heads up poker and differences between heads up poker and and multi-way poker yeah so I mentioned and multi-way poker yeah so I mentioned and multi-way poker yeah so I mentioned you know Nash equilibrium and two player you know Nash equilibrium and two player you know Nash equilibrium and two player zero-sum games zero-sum games zero-sum games if you play that strategy you are if you play that strategy you are if you play that strategy you are guaranteed to not lose an expectation no guaranteed to not lose an expectation no guaranteed to not lose an expectation no matter what your opponent does now once matter what your opponent does now once matter what your opponent does now once you go to six player poker you're no you go to six player poker you're no you go to six player poker you're no longer playing a two player zero-sum longer playing a two player zero-sum longer playing a two player zero-sum game and so there was a lot of debate game and so there was a lot of debate game and so there was a lot of debate among the academic community and among among the academic community and among among the academic community and among the poker Community about how well these the poker Community about how well these the poker Community about how well these techniques would extend beyond just techniques would extend beyond just techniques would extend beyond just two-player heads-up poker now
-
two-player heads-up poker now two-player heads-up poker now what I have come to realize is that what I have come to realize is that what I have come to realize is that um the techniques actually I thought um the techniques actually I thought um the techniques actually I thought really would extend to six player poker really would extend to six player poker really would extend to six player poker because even though in theory they don't because even though in theory they don't because even though in theory they don't give you these guarantees outside of two give you these guarantees outside of two give you these guarantees outside of two player zero some games in practice it player zero some games in practice it player zero some games in practice it still gives you a really strong strategy still gives you a really strong strategy still gives you a really strong strategy now there were a lot of complications now there were a lot of complications now there were a lot of complications that would come up with six player poker that would come up with six player poker that would come up with six player poker besides like the game theoretic aspect I besides like the game theoretic aspect I besides like the game theoretic aspect I mean for one the game is Just mean for one the game is Just mean for one the game is Just exponentially larger exponentially larger exponentially larger um so the main thing that allowed us to um so the main thing that allowed us to um so the main thing that allowed us to go from two player to six player was the go from two player to six player was the go from two player to six player was the idea of depth limited search idea of depth limited search idea of depth limited search so I said before like you know we would so I said before like you know we would so I said before like you know we would do search we would plan out the bot do search we would plan out the bot do search we would plan out the bot would plan out like what what it's going would plan out like what what it's going would plan out like what what it's going to do next and for the next several to do next and for the next several to do next and for the next several moves and in liberatis that search was moves and in liberatis that search was moves and in liberatis that search was done extending all the way to the end of done extending all the way to the end of done extending all the way to the end of the game so it would have to start the game so it would have to start the game so it would have to start um it from from the turn onwards like um it from from the turn onwards like um it from from the turn onwards like looking maybe 10 moves ahead looking maybe 10 moves ahead looking maybe 10 moves ahead um it would have to figure out what it um it would have to figure out what it um it would have to figure out what it was doing for all those moves was doing for all those moves was doing for all those moves now when you get to six player poker it now when you get to six player poker it now when you get to six player poker it can't do that exhaustive search anymore can't do that exhaustive search anymore can't do that exhaustive search anymore because the game is just way too large because the game is just way too large because the game is just way too large um but by only having to look a few um but by only having to look a few um but by only having to look a few moves ahead and then stopping there and moves ahead and then stopping there and moves ahead and then stopping there and substituting a Value Estimate of like substituting a Value Estimate of like substituting a Value Estimate of like how good is that strategy at that point how good is that strategy at that point how good is that strategy at that point then we're able to do a much more then we're able to do a much more then we're able to do a much more scalable form of search is there something cool looking at the is there something cool looking at the paper right now is there something cool paper right now is there something cool paper right now is there something cool in the paper in terms of Graphics a game in the paper in terms of Graphics a game in the paper in terms of Graphics a game tree Traverse of via Monte Carlo I think tree Traverse of via Monte Carlo I think tree Traverse of via Monte Carlo I think if you go down a bit uh if you go down a bit uh if you go down a bit uh uh figure one an example of equilibrium
-
uh figure one an example of equilibrium uh figure one an example of equilibrium selection problem ooh so yeah uh what do selection problem ooh so yeah uh what do selection problem ooh so yeah uh what do we know about equilibria one is there's we know about equilibria one is there's we know about equilibria one is there's multiple players so when you go outside multiple players so when you go outside multiple players so when you go outside of two players you're a sum so a Nash of two players you're a sum so a Nash of two players you're a sum so a Nash equilibrium is a set of strategies like equilibrium is a set of strategies like equilibrium is a set of strategies like one strategy for each player where no one strategy for each player where no one strategy for each player where no player has an incentive to switch to a player has an incentive to switch to a player has an incentive to switch to a different strategy different strategy different strategy um and so you can kind of think of it as um and so you can kind of think of it as um and so you can kind of think of it as like imagine you have a game where like imagine you have a game where like imagine you have a game where there's a ring that's actually the there's a ring that's actually the there's a ring that's actually the visual here you got a ring and the visual here you got a ring and the visual here you got a ring and the object of the game is to be as far away object of the game is to be as far away object of the game is to be as far away from the other players as possible from the other players as possible from the other players as possible there's an ash equilibrium is for all there's an ash equilibrium is for all there's an ash equilibrium is for all the players to be spaced equally apart the players to be spaced equally apart the players to be spaced equally apart around this ring around this ring around this ring but there's infinitely many different but there's infinitely many different but there's infinitely many different Nash equilibria right there's infinitely Nash equilibria right there's infinitely Nash equilibria right there's infinitely many ways to space four dots along a many ways to space four dots along a many ways to space four dots along a ring ring ring And if every single player independently And if every single player independently And if every single player independently computes a Nash equilibrium computes a Nash equilibrium computes a Nash equilibrium then there's no guarantee that the joint then there's no guarantee that the joint then there's no guarantee that the joint strategy that they're all playing is strategy that they're all playing is strategy that they're all playing is going to result is going to be in Ash going to result is going to be in Ash going to result is going to be in Ash equilibrium there they're just going to equilibrium there they're just going to equilibrium there they're just going to be like random dots scattered along this be like random dots scattered along this be like random dots scattered along this ring rather than four coordinated dots ring rather than four coordinated dots ring rather than four coordinated dots being equally spaced apart is it being equally spaced apart is it being equally spaced apart is it possible to sort of optimally do this possible to sort of optimally do this possible to sort of optimally do this kind of selection kind of selection kind of selection to do the selection about to do the selection about to do the selection about um of the equilibrium you're chasing so um of the equilibrium you're chasing so um of the equilibrium you're chasing so is there like a meta problem to be is there like a meta problem to be is there like a meta problem to be solved here so the meta problem is in solved here so the meta problem is in solved here so the meta problem is in some sense some sense some sense um how do you how do you understand the um how do you how do you understand the um how do you how do you understand the national equilibria that the other national equilibria that the other national equilibria that the other players are going to play players are going to play players are going to play um and and even if you do that again um and and even if you do that again um and and even if you do that again there's no guarantee that you're going there's no guarantee that you're going there's no guarantee that you're going to win so to win so to win so you know if you're playing you know if you're playing you know if you're playing uh if you're playing risk like I said
-
uh if you're playing risk like I said uh if you're playing risk like I said and and all the other players decide to and and all the other players decide to and and all the other players decide to team up against you You're Gonna Lose team up against you You're Gonna Lose team up against you You're Gonna Lose Nash equilibrium doesn't help you there Nash equilibrium doesn't help you there Nash equilibrium doesn't help you there and so there was this big debate about and so there was this big debate about and so there was this big debate about whether Nash equilibrium and all these whether Nash equilibrium and all these whether Nash equilibrium and all these techniques that compute it are even techniques that compute it are even techniques that compute it are even useful once you go outside of two player useful once you go outside of two player useful once you go outside of two player zero some games now I think for many zero some games now I think for many zero some games now I think for many games there is a valid criticism here games there is a valid criticism here games there is a valid criticism here and I think when we talk about when we and I think when we talk about when we and I think when we talk about when we go to something like diplomacy we run go to something like diplomacy we run go to something like diplomacy we run into this issue that the approach of into this issue that the approach of into this issue that the approach of trying to approximate a Nash equilibrium trying to approximate a Nash equilibrium trying to approximate a Nash equilibrium doesn't really work anymore but it turns doesn't really work anymore but it turns doesn't really work anymore but it turns out that in six player poker out that in six player poker out that in six player poker um because six player poker is such an um because six player poker is such an um because six player poker is such an adversarial game adversarial game adversarial game um where none of the players really try um where none of the players really try um where none of the players really try to work with each other to work with each other to work with each other the techniques that were used in the techniques that were used in the techniques that were used in two-player poker to try to approximate two-player poker to try to approximate two-player poker to try to approximate an equilibrium those still end up an equilibrium those still end up an equilibrium those still end up working in practice in in six player working in practice in in six player working in practice in in six player poker there's some poker there's some poker there's some deep way in which six player poker is deep way in which six player poker is deep way in which six player poker is just a bunch of heads up poker like just a bunch of heads up poker like just a bunch of heads up poker like games in one it's like uh it's like games in one it's like uh it's like games in one it's like uh it's like embedded in it so the competitiveness embedded in it so the competitiveness embedded in it so the competitiveness um is more fundamental to Poker than the um is more fundamental to Poker than the um is more fundamental to Poker than the cooperation right yeah poker is just cooperation right yeah poker is just cooperation right yeah poker is just such an adversarial game there's no real such an adversarial game there's no real such an adversarial game there's no real cooperation in fact you're not even cooperation in fact you're not even cooperation in fact you're not even allowed to cooperate in poker it's allowed to cooperate in poker it's allowed to cooperate in poker it's considered collusion it's against the considered collusion it's against the considered collusion it's against the rules rules rules um um um and so for that reason the techniques and so for that reason the techniques and so for that reason the techniques end up working really well and I think end up working really well and I think end up working really well and I think that's true more more broadly in that's true more more broadly in that's true more more broadly in extremely adverse serial games in extremely adverse serial games in extremely adverse serial games in general but that's sort of in practice general but that's sort of in practice general but that's sort of in practice versus being able to prove something versus being able to prove something versus being able to prove something that's right nobody has a proof that that's right nobody has a proof that that's right nobody has a proof that that's the case and it could be that that's the case and it could be that that's the case and it could be that that six player poker belongs to some that six player poker belongs to some that six player poker belongs to some class of games where a pro approximating
-
class of games where a pro approximating class of games where a pro approximating an Azure equilibrium through self-play an Azure equilibrium through self-play an Azure equilibrium through self-play provably works well provably works well provably works well um and you know there are other classes um and you know there are other classes um and you know there are other classes of games Beyond just two player Zero Sum of games Beyond just two player Zero Sum of games Beyond just two player Zero Sum where this is proven to work well so where this is proven to work well so where this is proven to work well so there are these you know kinds of games there are these you know kinds of games there are these you know kinds of games called potential games which I won't go called potential games which I won't go called potential games which I won't go into it's kind of like a complicated into it's kind of like a complicated into it's kind of like a complicated concept but concept but concept but um there are classes of games where uh um there are classes of games where uh um there are classes of games where uh this approach to approximating an ash this approach to approximating an ash this approach to approximating an ash equilibrium is proven to work well now equilibrium is proven to work well now equilibrium is proven to work well now six player poker is not known to belong six player poker is not known to belong six player poker is not known to belong to one of those classes but it is to one of those classes but it is to one of those classes but it is possible that there is some classic possible that there is some classic possible that there is some classic games where it either provably performs games where it either provably performs games where it either provably performs well or provably performs not that badly well or provably performs not that badly well or provably performs not that badly so what are some interesting things so what are some interesting things so what are some interesting things about uh pluribus that was able to about uh pluribus that was able to about uh pluribus that was able to achieve human level performance on this achieve human level performance on this achieve human level performance on this or superhuman level performance on the or superhuman level performance on the or superhuman level performance on the six player version of Poker I personally six player version of Poker I personally six player version of Poker I personally I think the most interesting interesting I think the most interesting interesting I think the most interesting interesting thing about pluribus is that it was so thing about pluribus is that it was so thing about pluribus is that it was so much cheaper than libratus I mean much cheaper than libratus I mean much cheaper than libratus I mean libratus if you had to put a price tag libratus if you had to put a price tag libratus if you had to put a price tag on on the computational resources that on on the computational resources that on on the computational resources that went into it I would say the final went into it I would say the final went into it I would say the final training run took about a hundred training run took about a hundred training run took about a hundred thousand dollars thousand dollars thousand dollars you go to pluribus the final training you go to pluribus the final training you go to pluribus the final training run would cost like less than 150 on AWS run would cost like less than 150 on AWS run would cost like less than 150 on AWS is this normalized to computational is this normalized to computational is this normalized to computational inflation so meaning uh this is is this inflation so meaning uh this is is this inflation so meaning uh this is is this just does this just have to do with the just does this just have to do with the just does this just have to do with the fact that pluribus was trained like a fact that pluribus was trained like a fact that pluribus was trained like a year later year later year later no no it's not it's I mean first of all no no it's not it's I mean first of all no no it's not it's I mean first of all like yeah Computing resources are are like yeah Computing resources are are like yeah Computing resources are are getting cheaper every day and like but getting cheaper every day and like but getting cheaper every day and like but you're not going to see a thousand-fold you're not going to see a thousand-fold you're not going to see a thousand-fold decrease in the computational resources decrease in the computational resources decrease in the computational resources over two years over two years over two years um or even anywhere close to that the
-
um or even anywhere close to that the um or even anywhere close to that the the real Improvement was algorithmic the real Improvement was algorithmic the real Improvement was algorithmic improvements and in particular the improvements and in particular the improvements and in particular the ability to do depth limited search ability to do depth limited search ability to do depth limited search so it does depth limited search also so it does depth limited search also so it does depth limited search also work for libratus yeah yes so where this work for libratus yeah yes so where this work for libratus yeah yes so where this deploymented search came from is you deploymented search came from is you deploymented search came from is you know I I developed this technique and know I I developed this technique and know I I developed this technique and um ran it on two-player poker first and um ran it on two-player poker first and um ran it on two-player poker first and that reduced the computational resources that reduced the computational resources that reduced the computational resources needed to make an AI that was superhuman needed to make an AI that was superhuman needed to make an AI that was superhuman from you know a hundred thousand dollars from you know a hundred thousand dollars from you know a hundred thousand dollars for the broadest to something you could for the broadest to something you could for the broadest to something you could train on your laptop what do you learn train on your laptop what do you learn train on your laptop what do you learn from that um from that discovery um from that discovery what I would take away from that is that what I would take away from that is that what I would take away from that is that algorithmic improvements really do algorithmic improvements really do algorithmic improvements really do matter how would you describe the more matter how would you describe the more matter how would you describe the more General case of limited Dev search General case of limited Dev search General case of limited Dev search so it's basically constraining the scale so it's basically constraining the scale so it's basically constraining the scale a temporal or in some other way of the a temporal or in some other way of the a temporal or in some other way of the computation you're doing in some clever computation you're doing in some clever computation you're doing in some clever way way way so like with like how else can you so like with like how else can you so like with like how else can you significantly constrain computation significantly constrain computation significantly constrain computation right right right well I think the idea is that we want to well I think the idea is that we want to well I think the idea is that we want to be able to leverage search as much as be able to leverage search as much as be able to leverage search as much as possible and the way that we were doing possible and the way that we were doing possible and the way that we were doing it in liberatis required us to search it in liberatis required us to search it in liberatis required us to search all the way to the end of the game now all the way to the end of the game now all the way to the end of the game now if you're playing a game like chess the if you're playing a game like chess the if you're playing a game like chess the idea that you're going to search always idea that you're going to search always idea that you're going to search always to the end of the game is kind of to the end of the game is kind of to the end of the game is kind of unimaginable right like there's just so unimaginable right like there's just so unimaginable right like there's just so many situations where you just won't be many situations where you just won't be many situations where you just won't be able to use search in that case or the able to use search in that case or the able to use search in that case or the cost would be cost would be cost would be um you know prohibitive um you know prohibitive um you know prohibitive and this technique allowed us to and this technique allowed us to and this technique allowed us to leverage search and without having to leverage search and without having to leverage search and without having to pay such a huge computational cost for pay such a huge computational cost for pay such a huge computational cost for it and be able to apply it more broadly it and be able to apply it more broadly it and be able to apply it more broadly so to what degree did you use neural
-
so to what degree did you use neural so to what degree did you use neural nets for uh libratus and pluribus and nets for uh libratus and pluribus and nets for uh libratus and pluribus and more generally what role do neural Nets more generally what role do neural Nets more generally what role do neural Nets have to play have to play have to play in um in super human level performance in um in super human level performance in um in super human level performance in poker so we actually did not use in poker so we actually did not use in poker so we actually did not use neural Nets at all for libratus or neural Nets at all for libratus or neural Nets at all for libratus or pluribus and a lot of people found this pluribus and a lot of people found this pluribus and a lot of people found this surprising back in 2017 I think they surprising back in 2017 I think they surprising back in 2017 I think they found it surprising today found it surprising today found it surprising today um that we were able to do this without um that we were able to do this without um that we were able to do this without using any neural Nets um using any neural Nets um using any neural Nets um and I think the reason for that I mean I and I think the reason for that I mean I and I think the reason for that I mean I think neural Nets are think neural Nets are think neural Nets are um incredibly powerful and the um incredibly powerful and the um incredibly powerful and the techniques that are used today even for techniques that are used today even for techniques that are used today even for poker AIS do rely uh quite heavily on poker AIS do rely uh quite heavily on poker AIS do rely uh quite heavily on neural Nets neural Nets neural Nets um but it wasn't the main challenge for um but it wasn't the main challenge for um but it wasn't the main challenge for poker like I think what neural Nets are poker like I think what neural Nets are poker like I think what neural Nets are really good for if you're in a situation really good for if you're in a situation really good for if you're in a situation where finding features for a value where finding features for a value where finding features for a value function is really difficult then neural function is really difficult then neural function is really difficult then neural Nets are really powerful and this was Nets are really powerful and this was Nets are really powerful and this was the problem in go right like the problem the problem in go right like the problem the problem in go right like the problem and go and go and go was that or the final problem in go at was that or the final problem in go at was that or the final problem in go at least was that nobody had a good way of least was that nobody had a good way of least was that nobody had a good way of looking at a board and figuring out who looking at a board and figuring out who looking at a board and figuring out who was winning or describing was winning or describing was winning or describing um through a simple algorithm who was um through a simple algorithm who was um through a simple algorithm who was winning or losing winning or losing winning or losing and so there neural Nets were super and so there neural Nets were super and so there neural Nets were super helpful because you could just feed in a helpful because you could just feed in a helpful because you could just feed in a ton of different board positions into ton of different board positions into ton of different board positions into this neural net and it would be able to this neural net and it would be able to this neural net and it would be able to predict then who was winning or losing predict then who was winning or losing predict then who was winning or losing but in poker the features weren't the but in poker the features weren't the but in poker the features weren't the challenge the the challenge was how do challenge the the challenge was how do challenge the the challenge was how do you design a scalable algorithm that you design a scalable algorithm that you design a scalable algorithm that would allow you to find this balance would allow you to find this balance would allow you to find this balance strategy that would understand that you strategy that would understand that you strategy that would understand that you have to Bluff with the right probability
-
have to Bluff with the right probability have to Bluff with the right probability so can that be somehow incorporated into so can that be somehow incorporated into so can that be somehow incorporated into the value function this the value function this the value function this the complexity of polka that you've the complexity of polka that you've the complexity of polka that you've described yeah so the way the value described yeah so the way the value described yeah so the way the value functions work in like the latest and functions work in like the latest and functions work in like the latest and greatest poker AIS they do use neural greatest poker AIS they do use neural greatest poker AIS they do use neural nets for the value function the way it's nets for the value function the way it's nets for the value function the way it's done is is very different from how it's done is is very different from how it's done is is very different from how it's done in a game like chess or go because done in a game like chess or go because done in a game like chess or go because in poker you have to reason about in poker you have to reason about in poker you have to reason about beliefs and so the value of a state beliefs and so the value of a state beliefs and so the value of a state depends on the beliefs that players have depends on the beliefs that players have depends on the beliefs that players have about what the different cards are like about what the different cards are like about what the different cards are like if you have pocket aces then whether if you have pocket aces then whether if you have pocket aces then whether that's a really really good hand or just that's a really really good hand or just that's a really really good hand or just an okay hand depends on whether you know an okay hand depends on whether you know an okay hand depends on whether you know I have pocket aces where like if you I have pocket aces where like if you I have pocket aces where like if you know that I have pocket aces then if I know that I have pocket aces then if I know that I have pocket aces then if I bet you're going to fold immediately but bet you're going to fold immediately but bet you're going to fold immediately but if you think that I have a really bad if you think that I have a really bad if you think that I have a really bad hand then I could bet with pocket aces hand then I could bet with pocket aces hand then I could bet with pocket aces and make a ton of money so and make a ton of money so and make a ton of money so the value function in poker these days the value function in poker these days the value function in poker these days takes the beliefs as an input which is takes the beliefs as an input which is takes the beliefs as an input which is very different from like how how chess very different from like how how chess very different from like how how chess and go AIS work and go AIS work and go AIS work so as a person who appreciates the game so as a person who appreciates the game so as a person who appreciates the game uh uh uh who do you think is the greatest poker who do you think is the greatest poker who do you think is the greatest poker player of all time player of all time player of all time that's a that's a tough question that's a that's a tough question that's a that's a tough question um Can an AI help answer that question um Can an AI help answer that question um Can an AI help answer that question can you can actually add analyze the can you can actually add analyze the can you can actually add analyze the quality of play right so the AHS engines quality of play right so the AHS engines quality of play right so the AHS engines can can can can give estimates of the quality of can give estimates of the quality of can give estimates of the quality of play right play right play right um um um I wonder if there's a is there an ELO I wonder if there's a is there an ELO I wonder if there's a is there an ELO rating type of system for poker
-
rating type of system for poker rating type of system for poker I suppose you could but there's just not I suppose you could but there's just not I suppose you could but there's just not enough enough enough you would have to play a lot of games you would have to play a lot of games you would have to play a lot of games right a very large number of games like right a very large number of games like right a very large number of games like more than you would in chess the more than you would in chess the more than you would in chess the deterministic game makes it easier to deterministic game makes it easier to deterministic game makes it easier to estimate yellow estimate yellow estimate yellow I think I think it is much harder to I think I think it is much harder to I think I think it is much harder to estimate something like ELO rating in estimate something like ELO rating in estimate something like ELO rating in poker I think it's doable the problem is poker I think it's doable the problem is poker I think it's doable the problem is that the game is very high variants so that the game is very high variants so that the game is very high variants so you could play you could be profitable you could play you could be profitable you could play you could be profitable in poker for a year and you could in poker for a year and you could in poker for a year and you could actually be a bad player just because actually be a bad player just because actually be a bad player just because the variance is so high I mean you've the variance is so high I mean you've the variance is so high I mean you've got top professional poker players that got top professional poker players that got top professional poker players that would lose for a year just because would lose for a year just because would lose for a year just because they're on a really bad they're on a really bad they're on a really bad um bad streak so yeah so for ELO you um bad streak so yeah so for ELO you um bad streak so yeah so for ELO you have to have a nice clean way of saying have to have a nice clean way of saying have to have a nice clean way of saying if player a played player B if player a played player B if player a played player B and a B's B that says something that's a and a B's B that says something that's a and a B's B that says something that's a signal in poker it's a very noisy signal signal in poker it's a very noisy signal signal in poker it's a very noisy signal it's a very noisy signal now there is a it's a very noisy signal now there is a it's a very noisy signal now there is a signal there and so you could do this signal there and so you could do this signal there and so you could do this this calculation it would just be much this calculation it would just be much this calculation it would just be much harder harder harder um but the same way that AIS have now um but the same way that AIS have now um but the same way that AIS have now taken over chess and you know all the taken over chess and you know all the taken over chess and you know all the top professional chess players train top professional chess players train top professional chess players train with with AIS the same is true for poker with with AIS the same is true for poker with with AIS the same is true for poker the game has become a very computational the game has become a very computational the game has become a very computational um people trained with AIS to try to um people trained with AIS to try to um people trained with AIS to try to find out where they're making mistakes find out where they're making mistakes find out where they're making mistakes try to learn from the AIS to improve try to learn from the AIS to improve try to learn from the AIS to improve their strategy so their strategy so their strategy so now yeah so the game has been now yeah so the game has been now yeah so the game has been revolutionized in the past five years by revolutionized in the past five years by revolutionized in the past five years by by the development of AI in this sport by the development of AI in this sport by the development of AI in this sport the skill with which you avoided the the skill with which you avoided the the skill with which you avoided the question of the greatest of all time was question of the greatest of all time was question of the greatest of all time was impressive so my feeling is that it's a impressive so my feeling is that it's a impressive so my feeling is that it's a difficult it's a difficult question
-
difficult it's a difficult question difficult it's a difficult question because just like in chess where you because just like in chess where you because just like in chess where you can't really compare Magnus Carlson can't really compare Magnus Carlson can't really compare Magnus Carlson today to Gary Kasparov today to Gary Kasparov today to Gary Kasparov um because the game has evolved so much um because the game has evolved so much um because the game has evolved so much um the poker players today are so far um the poker players today are so far um the poker players today are so far beyond the the skills of like people beyond the the skills of like people beyond the the skills of like people that were playing even 10 or 20 years that were playing even 10 or 20 years that were playing even 10 or 20 years ago ago ago um so you look at the kinds of like um so you look at the kinds of like um so you look at the kinds of like All-Stars that were on ESPN at like the All-Stars that were on ESPN at like the All-Stars that were on ESPN at like the height of the poker boom height of the poker boom height of the poker boom pretty much all those players are pretty much all those players are pretty much all those players are actually not that good at the game today actually not that good at the game today actually not that good at the game today at least at least the the strategy at least at least the the strategy at least at least the the strategy aspect I mean there might be still be aspect I mean there might be still be aspect I mean there might be still be good at like reading the player at the good at like reading the player at the good at like reading the player at the other side of the table and trying to other side of the table and trying to other side of the table and trying to figure out like are they bluffing or not figure out like are they bluffing or not figure out like are they bluffing or not but in terms of the actual like but in terms of the actual like but in terms of the actual like computational strategy of the game computational strategy of the game computational strategy of the game um a lot of them have really struggled um a lot of them have really struggled um a lot of them have really struggled to keep up with that development now to keep up with that development now to keep up with that development now so for that reason I'll give an answer so for that reason I'll give an answer so for that reason I'll give an answer and I'm gonna say Daniel negranio who and I'm gonna say Daniel negranio who and I'm gonna say Daniel negranio who you actually had on the podcast recently you actually had on the podcast recently you actually had on the podcast recently I saw was a great episode and I love I saw was a great episode and I love I saw was a great episode and I love this so much and Phil's gonna hate this this so much and Phil's gonna hate this this so much and Phil's gonna hate this so much and I'm gonna give him I'm gonna so much and I'm gonna give him I'm gonna so much and I'm gonna give him I'm gonna give him credit because he is one of the give him credit because he is one of the give him credit because he is one of the few like old school really strong few like old school really strong few like old school really strong players that have kept up with the players that have kept up with the players that have kept up with the development of AI so he is trying to development of AI so he is trying to development of AI so he is trying to he's constantly studying the the game he's constantly studying the the game he's constantly studying the the game theory optimal way of playing exactly theory optimal way of playing exactly theory optimal way of playing exactly yeah and I think a lot of a lot of the yeah and I think a lot of a lot of the yeah and I think a lot of a lot of the old school poker players are just kind old school poker players are just kind old school poker players are just kind of given up on that aspect and and I got of given up on that aspect and and I got of given up on that aspect and and I got to give them the ground you credit for to give them the ground you credit for to give them the ground you credit for for keeping up with all the developments for keeping up with all the developments for keeping up with all the developments that are happening in the sport yeah that are happening in the sport yeah that are happening in the sport yeah it's fascinating to watch it's it's fascinating to watch it's it's fascinating to watch it's fascinating to watch where it's headed fascinating to watch where it's headed fascinating to watch where it's headed um yeah so there you go some love for
-
um yeah so there you go some love for um yeah so there you go some love for Daniel Daniel Daniel quick pause bathroom break yeah let's do quick pause bathroom break yeah let's do quick pause bathroom break yeah let's do it it it let's go from poker to diplomacy let's go from poker to diplomacy let's go from poker to diplomacy what is at a high level the game of what is at a high level the game of what is at a high level the game of diplomacy diplomacy diplomacy yeah so I talked a lot about two player yeah so I talked a lot about two player yeah so I talked a lot about two player zero some games and what's interesting zero some games and what's interesting zero some games and what's interesting about diplomacy is that it's very about diplomacy is that it's very about diplomacy is that it's very different from these like adversarial uh different from these like adversarial uh different from these like adversarial uh games like chess go poker even Starcraft games like chess go poker even Starcraft games like chess go poker even Starcraft and DOTA diplomacy has a much bigger and DOTA diplomacy has a much bigger and DOTA diplomacy has a much bigger Cooperative element to it it's a seven Cooperative element to it it's a seven Cooperative element to it it's a seven player game it was actually created in player game it was actually created in player game it was actually created in the 50s the 50s the 50s um and it takes place um and it takes place um and it takes place uh before World War one it's like a map uh before World War one it's like a map uh before World War one it's like a map of Europe with seven Great Powers of Europe with seven Great Powers of Europe with seven Great Powers um and they're all trying to form um and they're all trying to form um and they're all trying to form alliances with each other there's a lot alliances with each other there's a lot alliances with each other there's a lot of negotiation going on of negotiation going on of negotiation going on um and so the whole focus of the game is um and so the whole focus of the game is um and so the whole focus of the game is on on on forming alliances with the other players forming alliances with the other players forming alliances with the other players to take on the other players England to take on the other players England to take on the other players England Germany Russia turkey Austria Hungary Germany Russia turkey Austria Hungary Germany Russia turkey Austria Hungary Italy and France that's right yeah Italy and France that's right yeah Italy and France that's right yeah so the way the game works is so the way the game works is so the way the game works is on each turn you spend about you know on each turn you spend about you know on each turn you spend about you know five to fifteen minutes talking to the five to fifteen minutes talking to the five to fifteen minutes talking to the other players in privates and you make other players in privates and you make other players in privates and you make all sorts of deals with them you say all sorts of deals with them you say all sorts of deals with them you say like hey let's work together like hey let's work together like hey let's work together um you know let's team up against this um you know let's team up against this um you know let's team up against this other player because the only way that other player because the only way that other player because the only way that you can make progress is by working with you can make progress is by working with you can make progress is by working with somebody else against the others somebody else against the others somebody else against the others um and then after that negotiation um and then after that negotiation um and then after that negotiation period is done all the players period is done all the players period is done all the players simultaneously submit their moves and simultaneously submit their moves and simultaneously submit their moves and they're all executed at the same time they're all executed at the same time they're all executed at the same time and so you can tell people like hey I'm
-
and so you can tell people like hey I'm and so you can tell people like hey I'm going to support you this turn going to support you this turn going to support you this turn um but then you don't follow through um but then you don't follow through um but then you don't follow through with it and they're only going to figure with it and they're only going to figure with it and they're only going to figure that out once they see the moves being that out once they see the moves being that out once they see the moves being read off how much of it is natural read off how much of it is natural read off how much of it is natural language like written actual text how language like written actual text how language like written actual text how much is like uh you're actually saying much is like uh you're actually saying much is like uh you're actually saying phrases that are structured so there's phrases that are structured so there's phrases that are structured so there's different ways to play the game you know different ways to play the game you know different ways to play the game you know you can play it in person and in that you can play it in person and in that you can play it in person and in that case it's all natural language freeform case it's all natural language freeform case it's all natural language freeform communication there's no constraints on communication there's no constraints on communication there's no constraints on the kinds of deals that you can make the the kinds of deals that you can make the the kinds of deals that you can make the kinds of things that you can discuss kinds of things that you can discuss kinds of things that you can discuss um it can also play it online so you can um it can also play it online so you can um it can also play it online so you can you know send along emails back and you know send along emails back and you know send along emails back and forth you can play it like live online forth you can play it like live online forth you can play it like live online or over voice chat but the the focus the or over voice chat but the the focus the or over voice chat but the the focus the important thing to understand is that important thing to understand is that important thing to understand is that this is unstructured communication you this is unstructured communication you this is unstructured communication you can say whatever you want can say whatever you want can say whatever you want um you can make any sorts of deals that um you can make any sorts of deals that um you can make any sorts of deals that you want and everything is done you want and everything is done you want and everything is done privately so it's not like you're all privately so it's not like you're all privately so it's not like you're all around the board together having a around the board together having a around the board together having a conversation you're grabbing somebody conversation you're grabbing somebody conversation you're grabbing somebody going off into a corner and conspiring going off into a corner and conspiring going off into a corner and conspiring behind everybody else's back about what behind everybody else's back about what behind everybody else's back about what you're planning and uh there's no limit you're planning and uh there's no limit you're planning and uh there's no limit in theory to the conversation you can in theory to the conversation you can in theory to the conversation you can have directly with one person that's have directly with one person that's have directly with one person that's right you can make all sorts of you can right you can make all sorts of you can right you can make all sorts of you can talk about anything you could say like talk about anything you could say like talk about anything you could say like hey let's have a long-term alliance hey let's have a long-term alliance hey let's have a long-term alliance against this guy you can say like hey against this guy you can say like hey against this guy you can say like hey can you support me this turn and in can you support me this turn and in can you support me this turn and in return I'll do this other thing for you return I'll do this other thing for you return I'll do this other thing for you next turn or um you know yeah just you next turn or um you know yeah just you next turn or um you know yeah just you can talk about like what you talked can talk about like what you talked can talk about like what you talked about with somebody else and gossip about with somebody else and gossip about with somebody else and gossip about like what they're planning about like what they're planning about like what they're planning um the way that I would describe the um the way that I would describe the um the way that I would describe the game is that it's kind of like a mix game is that it's kind of like a mix game is that it's kind of like a mix between risk poker and the TV show between risk poker and the TV show between risk poker and the TV show Survivor there's like this big element Survivor there's like this big element Survivor there's like this big element of like trying to
-
of like trying to of like trying to um yeah there's a big social element and um yeah there's a big social element and um yeah there's a big social element and the best way that I would describe the the best way that I would describe the the best way that I would describe the game is that it's really a game about game is that it's really a game about game is that it's really a game about people rather than the pieces people rather than the pieces people rather than the pieces so risk because it is a map it's kind of so risk because it is a map it's kind of so risk because it is a map it's kind of war game like war game like war game like uh poker because there's a game theory uh poker because there's a game theory uh poker because there's a game theory component that's very kind of strategic component that's very kind of strategic component that's very kind of strategic so you could convert it into an so you could convert it into an so you could convert it into an artificial intelligence problem and then artificial intelligence problem and then artificial intelligence problem and then survive it because of the social survive it because of the social survive it because of the social component that's a strong social component that's a strong social component that's a strong social component I saw that somebody said component I saw that somebody said component I saw that somebody said online that the internet version of the online that the internet version of the online that the internet version of the game has this quality of that is easier game has this quality of that is easier game has this quality of that is easier to almost to do like role playing to almost to do like role playing to almost to do like role playing as opposed to being yourself you can as opposed to being yourself you can as opposed to being yourself you can actually like be the like really imagine actually like be the like really imagine actually like be the like really imagine yourself as the leader of France or yourself as the leader of France or yourself as the leader of France or Russia and so on like really pretend to Russia and so on like really pretend to Russia and so on like really pretend to be that person it's actually fun to be that person it's actually fun to be that person it's actually fun to really lean into being that that leader really lean into being that that leader really lean into being that that leader yeah so some some players do go this yeah so some some players do go this yeah so some some players do go this route where they just like kind of view route where they just like kind of view route where they just like kind of view it as a strategy game but also a it as a strategy game but also a it as a strategy game but also a role-playing game where they can like role-playing game where they can like role-playing game where they can like act out like what would I be like if I act out like what would I be like if I act out like what would I be like if I was you know a leader of France in 1900 was you know a leader of France in 1900 was you know a leader of France in 1900 a forfeit right away no I'm just kidding a forfeit right away no I'm just kidding a forfeit right away no I'm just kidding um um um and they sometimes use like the and they sometimes use like the and they sometimes use like the old-timey language to like old-timey language to like old-timey language to like um or how they imagine the elites would um or how they imagine the elites would um or how they imagine the elites would talk at that time anyway so the what are talk at that time anyway so the what are talk at that time anyway so the what are the different turns of the game like the different turns of the game like the different turns of the game like what are the rounds yeah so on on every what are the rounds yeah so on on every what are the rounds yeah so on on every turn you got like a bunch of different turn you got like a bunch of different turn you got like a bunch of different units that you start out with so you units that you start out with so you units that you start out with so you start out um controlling like just a few start out um controlling like just a few start out um controlling like just a few units and the object of the game is to units and the object of the game is to units and the object of the game is to gain control of a majority of the map if gain control of a majority of the map if gain control of a majority of the map if you able if you're able to do that then
-
you able if you're able to do that then you able if you're able to do that then you've won the game but like I said the you've won the game but like I said the you've won the game but like I said the only way that you're able to do that is only way that you're able to do that is only way that you're able to do that is by working with other players so on by working with other players so on by working with other players so on every turn you can issue a move order so every turn you can issue a move order so every turn you can issue a move order so for each of your units you can move them for each of your units you can move them for each of your units you can move them to an adjacent territory to an adjacent territory to an adjacent territory or you can keep them where they are or or you can keep them where they are or or you can keep them where they are or you can support a move or a hold of a you can support a move or a hold of a you can support a move or a hold of a different units so what are the different units so what are the different units so what are the territories how how is the map divided territories how how is the map divided territories how how is the map divided up it's kind of like Risk where the the up it's kind of like Risk where the the up it's kind of like Risk where the the map is divided up into like 50 different map is divided up into like 50 different map is divided up into like 50 different territories territories territories um now you can enter a territory if um now you can enter a territory if um now you can enter a territory if you're moving into that territory with you're moving into that territory with you're moving into that territory with more supports than the person that's in more supports than the person that's in more supports than the person that's in there or the person that's trying to there or the person that's trying to there or the person that's trying to move in there so if you're moving in and move in there so if you're moving in and move in there so if you're moving in and there's somebody already there there's somebody already there there's somebody already there um then if neither of you have support um then if neither of you have support um then if neither of you have support it's a one versus one and you'll bounce it's a one versus one and you'll bounce it's a one versus one and you'll bounce back another if you'll make progress if back another if you'll make progress if back another if you'll make progress if you have a unit that's supporting that you have a unit that's supporting that you have a unit that's supporting that move into the territory then it's a two move into the territory then it's a two move into the territory then it's a two versus one and you'll kick them out and versus one and you'll kick them out and versus one and you'll kick them out and they'll have to retreat somewhere what they'll have to retreat somewhere what they'll have to retreat somewhere what does support mean support is like it's does support mean support is like it's does support mean support is like it's it's an action that you can issue in the it's an action that you can issue in the it's an action that you can issue in the game so you can say this unit you write game so you can say this unit you write game so you can say this unit you write down this unit is supporting this other down this unit is supporting this other down this unit is supporting this other unit into this territory are these units unit into this territory are these units unit into this territory are these units from opposing forces they could be they from opposing forces they could be they from opposing forces they could be they could be and this is this is where the could be and this is this is where the could be and this is this is where the interesting aspect of the game comes in interesting aspect of the game comes in interesting aspect of the game comes in because you can support your own units because you can support your own units because you can support your own units into territory but you can also support into territory but you can also support into territory but you can also support other people's units into territories other people's units into territories other people's units into territories and so that's what the negotiations and so that's what the negotiations and so that's what the negotiations really revolve around but you don't have really revolve around but you don't have really revolve around but you don't have to do the thing you say you're going to to do the thing you say you're going to to do the thing you say you're going to do and this yeah and so you can say I'm do and this yeah and so you can say I'm do and this yeah and so you can say I'm going to support you but then backstab going to support you but then backstab going to support you but then backstab the person yeah that's absolutely right the person yeah that's absolutely right the person yeah that's absolutely right and that tension is core to the game the and that tension is core to the game the and that tension is core to the game the attention is absolutely core to the game
-
attention is absolutely core to the game attention is absolutely core to the game the the fact that you can make all sorts the the fact that you can make all sorts the the fact that you can make all sorts of promises but you have to reason about of promises but you have to reason about of promises but you have to reason about the fact that like hey they might not the fact that like hey they might not the fact that like hey they might not trust you if you say you're going to do trust you if you say you're going to do trust you if you say you're going to do something or they might be lying to you something or they might be lying to you something or they might be lying to you when they say they're going to support when they say they're going to support when they say they're going to support you you you so maybe just just to jump back what's so maybe just just to jump back what's so maybe just just to jump back what's what's the history of the game in what's the history of the game in what's the history of the game in general is it true that Henry Kissinger general is it true that Henry Kissinger general is it true that Henry Kissinger loved the game and JFK and all those loved the game and JFK and all those loved the game and JFK and all those I've heard like a bunch of different I've heard like a bunch of different I've heard like a bunch of different people that or is that just one of those people that or is that just one of those people that or is that just one of those things that the cool kids say they do things that the cool kids say they do things that the cool kids say they do but they don't actually play so the game but they don't actually play so the game but they don't actually play so the game was created in the 50s yeah was created in the 50s yeah was created in the 50s yeah um and from what I understand it was um um and from what I understand it was um um and from what I understand it was um JFK's it was played in like the JFK JFK's it was played in like the JFK JFK's it was played in like the JFK White House Henry Kissinger's favorite White House Henry Kissinger's favorite White House Henry Kissinger's favorite game I don't know if it's true but um game I don't know if it's true but um game I don't know if it's true but um that's definitely what I've heard it's that's definitely what I've heard it's that's definitely what I've heard it's interesting that they went with World interesting that they went with World interesting that they went with World War One War One War One when it was created after World War II when it was created after World War II when it was created after World War II so the story that I've heard for the so the story that I've heard for the so the story that I've heard for the creation of the game is it was created creation of the game is it was created creation of the game is it was created by by by um somebody that had looked at the um somebody that had looked at the um somebody that had looked at the history of the 20th century and they saw history of the 20th century and they saw history of the 20th century and they saw World War one as a failure of diplomacy World War one as a failure of diplomacy World War one as a failure of diplomacy so sure you know they saw the fact that so sure you know they saw the fact that so sure you know they saw the fact that this war broke out as like the the this war broke out as like the the this war broke out as like the the diplomats of all these countries like diplomats of all these countries like diplomats of all these countries like really failed to prevent a war and he really failed to prevent a war and he really failed to prevent a war and he wanted to create a game that would wanted to create a game that would wanted to create a game that would basically teach people about diplomacy basically teach people about diplomacy basically teach people about diplomacy um and it's really fascinating that like um and it's really fascinating that like um and it's really fascinating that like in his ideal version of the game of in his ideal version of the game of in his ideal version of the game of diplomacy nobody actually wins the game diplomacy nobody actually wins the game diplomacy nobody actually wins the game because the whole point is that if because the whole point is that if because the whole point is that if somebody is about to win then the other somebody is about to win then the other somebody is about to win then the other players should be able to work together players should be able to work together players should be able to work together to stop that person from winning and so to stop that person from winning and so to stop that person from winning and so the ideal version of the game is just the ideal version of the game is just the ideal version of the game is just one where nobody actually wins and you
-
one where nobody actually wins and you one where nobody actually wins and you know it kind of has a nice like know it kind of has a nice like know it kind of has a nice like wholesome take-home message then that wholesome take-home message then that wholesome take-home message then that you know war war is ultimately futile you know war war is ultimately futile you know war war is ultimately futile and uh and uh and uh and that optimal and that optimal and that optimal that feudal optimal could be achieved that feudal optimal could be achieved that feudal optimal could be achieved through great diplomacy yeah so uh is through great diplomacy yeah so uh is through great diplomacy yeah so uh is there some asymmetry in in terms of there some asymmetry in in terms of there some asymmetry in in terms of which is more powerful Russia versus which is more powerful Russia versus which is more powerful Russia versus Germany versus Germany versus Germany versus France and so on so I think the general France and so on so I think the general France and so on so I think the general consensus is that France is the consensus is that France is the consensus is that France is the strongest power in the game but the strongest power in the game but the strongest power in the game but the beautiful thing about diplomacy is that beautiful thing about diplomacy is that beautiful thing about diplomacy is that it's it's self-balancing right so it's it's it's self-balancing right so it's it's it's self-balancing right so it's the fact that France has an inherent the fact that France has an inherent the fact that France has an inherent Advantage from the beginning means that Advantage from the beginning means that Advantage from the beginning means that the other players are less likely to the other players are less likely to the other players are less likely to work with it I saw that Russia has four work with it I saw that Russia has four work with it I saw that Russia has four units or four of something that the units or four of something that the units or four of something that the others have three of something that's others have three of something that's others have three of something that's true yeah so Russia starts off with four true yeah so Russia starts off with four true yeah so Russia starts off with four units while all the other players start units while all the other players start units while all the other players start with three but Russia is also in a much with three but Russia is also in a much with three but Russia is also in a much more vulnerable position because they more vulnerable position because they more vulnerable position because they have to like have to like have to like um they have a lot more neighbors as um they have a lot more neighbors as um they have a lot more neighbors as well got it larger territory more uh well got it larger territory more uh well got it larger territory more uh yeah right more border to defend okay uh yeah right more border to defend okay uh yeah right more border to defend okay uh what else is what else is important to what else is what else is important to what else is what else is important to know about the rules so there how many know about the rules so there how many know about the rules so there how many rounds are there like is this iterative rounds are there like is this iterative rounds are there like is this iterative game is there is it is it finite you game is there is it is it finite you game is there is it is it finite you just keep going indefinitely usually the just keep going indefinitely usually the just keep going indefinitely usually the game lasts uh I would say about game lasts uh I would say about game lasts uh I would say about 15 or 20 turns 15 or 20 turns 15 or 20 turns um there's in theory No Limit it could um there's in theory No Limit it could um there's in theory No Limit it could last longer but at some point I mean if last longer but at some point I mean if last longer but at some point I mean if you're playing a house game with friends you're playing a house game with friends you're playing a house game with friends at some point you just get tired and you at some point you just get tired and you at some point you just get tired and you all agree like okay we're gonna end the all agree like okay we're gonna end the all agree like okay we're gonna end the game here and call it a draw game here and call it a draw game here and call it a draw um if you're playing online there's um if you're playing online there's um if you're playing online there's usually like set limits on when the game usually like set limits on when the game usually like set limits on when the game will actually end and what's the end
-
will actually end and what's the end will actually end and what's the end what's the termination condition like what's the termination condition like what's the termination condition like this this one country have to conquer this this one country have to conquer this this one country have to conquer everything else so if somebody is able everything else so if somebody is able everything else so if somebody is able to actually gain control of a majority to actually gain control of a majority to actually gain control of a majority of the map then then they've won the of the map then then they've won the of the map then then they've won the game and that is a solo Victory as it's game and that is a solo Victory as it's game and that is a solo Victory as it's called now that pretty rarely happens called now that pretty rarely happens called now that pretty rarely happens especially with strong players because especially with strong players because especially with strong players because like I said the game is designed to like I said the game is designed to like I said the game is designed to incentivize the other players to put a incentivize the other players to put a incentivize the other players to put a stop to that and all work together to stop to that and all work together to stop to that and all work together to stop the superpower stop the superpower stop the superpower um usually what ends up happening is um usually what ends up happening is um usually what ends up happening is that you know all the players agree to a that you know all the players agree to a that you know all the players agree to a draw and then the the score the the win draw and then the the score the the win draw and then the the score the the win is divided among the remaining players is divided among the remaining players is divided among the remaining players um there's a lot of different scoring um there's a lot of different scoring um there's a lot of different scoring systems the one that we used in our systems the one that we used in our systems the one that we used in our research research research um basically um basically um basically um gives a score relative to how much um gives a score relative to how much um gives a score relative to how much control you have of the map so the more control you have of the map so the more control you have of the map so the more that you control the higher you score that you control the higher you score that you control the higher you score what's the history of using this game as what's the history of using this game as what's the history of using this game as a benchmark for AI research do people a benchmark for AI research do people a benchmark for AI research do people use it yeah so people have been working use it yeah so people have been working use it yeah so people have been working on AI for diplomacy since about the 80s on AI for diplomacy since about the 80s on AI for diplomacy since about the 80s um there was some really exciting um there was some really exciting um there was some really exciting research back then but the approach that research back then but the approach that research back then but the approach that was taken was very different from what was taken was very different from what was taken was very different from what we see today I mean the research in the we see today I mean the research in the we see today I mean the research in the 80s was a very rule-based approach kind 80s was a very rule-based approach kind 80s was a very rule-based approach kind of kind of a heuristic approach it was of kind of a heuristic approach it was of kind of a heuristic approach it was very in line with the kind of research very in line with the kind of research very in line with the kind of research that was being done in the 80s you know that was being done in the 80s you know that was being done in the 80s you know basically trying to encode human basically trying to encode human basically trying to encode human knowledge into the strategy of the AI knowledge into the strategy of the AI knowledge into the strategy of the AI sure sure sure um and you know it's understandable I um and you know it's understandable I um and you know it's understandable I mean the game is so incredibly different mean the game is so incredibly different mean the game is so incredibly different and so so much more complicated than the and so so much more complicated than the and so so much more complicated than the kinds of games that people were working kinds of games that people were working kinds of games that people were working on like chess and go uh and poker that
-
on like chess and go uh and poker that on like chess and go uh and poker that it was honestly even hard to like start it was honestly even hard to like start it was honestly even hard to like start getting making any progress in in getting making any progress in in getting making any progress in in diplomacy can you just formulate what is diplomacy can you just formulate what is diplomacy can you just formulate what is the problem from an AI perspective and the problem from an AI perspective and the problem from an AI perspective and why is it hard why is it a challenging why is it hard why is it a challenging why is it hard why is it a challenging game to solve so there's a lot of game to solve so there's a lot of game to solve so there's a lot of aspects in diplomacy that make it a huge aspects in diplomacy that make it a huge aspects in diplomacy that make it a huge challenge first of all you have the challenge first of all you have the challenge first of all you have the natural language components and I think natural language components and I think natural language components and I think this really is what makes it are really this really is what makes it are really this really is what makes it are really the most difficult the most difficult the most difficult game among like the major benchmarks the game among like the major benchmarks the game among like the major benchmarks the fact that you have to it's not about fact that you have to it's not about fact that you have to it's not about moving pieces on the board moving pieces on the board moving pieces on the board your action space is basically all the your action space is basically all the your action space is basically all the different sentences that you could different sentences that you could different sentences that you could communicate to somebody else in this communicate to somebody else in this communicate to somebody else in this game and um is there can we just like game and um is there can we just like game and um is there can we just like Linger on that so Linger on that so Linger on that so is part of it like the ambiguity in the is part of it like the ambiguity in the is part of it like the ambiguity in the language language language if it was like very strict if it was like very strict if it was like very strict if you narrowed the set of possible if you narrowed the set of possible if you narrowed the set of possible sentences you could do it would that sentences you could do it would that sentences you could do it would that simplify the game significantly the the simplify the game significantly the the simplify the game significantly the the real real real difficulty is the breadth of things that difficulty is the breadth of things that difficulty is the breadth of things that you can talk about you can talk about you can talk about um you can have natural language and um you can have natural language and um you can have natural language and other games and like Sellers of Catan other games and like Sellers of Catan other games and like Sellers of Catan for example like you could have a for example like you could have a for example like you could have a natural language Settlers of Catan AI natural language Settlers of Catan AI natural language Settlers of Catan AI but the things that you're going to talk but the things that you're going to talk but the things that you're going to talk about are basically like am I trading about are basically like am I trading about are basically like am I trading you two sheep for a wood or three sheep you two sheep for a wood or three sheep you two sheep for a wood or three sheep for a wood for a wood for a wood um whereas in a game like diplomacy the um whereas in a game like diplomacy the um whereas in a game like diplomacy the breadth of conversations that you're breadth of conversations that you're breadth of conversations that you're going to have are like you know am I going to have are like you know am I going to have are like you know am I going to support you are you going to going to support you are you going to going to support you are you going to support me in return which units are support me in return which units are support me in return which units are going to do what uh what did this other going to do what uh what did this other going to do what uh what did this other person say promise you uh they're lying person say promise you uh they're lying person say promise you uh they're lying because they told this other person that
-
because they told this other person that because they told this other person that they're going to do this instead they're going to do this instead they're going to do this instead um if you help me out this turn then in um if you help me out this turn then in um if you help me out this turn then in the future I'll do these things that the future I'll do these things that the future I'll do these things that will help you out will help you out will help you out um the the depth and breadth of these um the the depth and breadth of these um the the depth and breadth of these conversations is is really complicated conversations is is really complicated conversations is is really complicated and it's all being done in natural and it's all being done in natural and it's all being done in natural language language language um now you could approach it and we um now you could approach it and we um now you could approach it and we actually consider doing this like you actually consider doing this like you actually consider doing this like you you know having a simplified language to you know having a simplified language to you know having a simplified language to make this complexity uh smaller But make this complexity uh smaller But make this complexity uh smaller But ultimately we thought the most impactful ultimately we thought the most impactful ultimately we thought the most impactful way of doing this research would be to way of doing this research would be to way of doing this research would be to address the natural language component address the natural language component address the natural language component head-on and just try to go for the full head-on and just try to go for the full head-on and just try to go for the full game up front game up front game up front just looking at sample games and what just looking at sample games and what just looking at sample games and what the conversations look like greetings the conversations look like greetings the conversations look like greetings England this should prove to be a fun England this should prove to be a fun England this should prove to be a fun game since all the private press is game since all the private press is game since all the private press is going to be made public at the end going to be made public at the end going to be made public at the end at the least it will be interesting to at the least it will be interesting to at the least it will be interesting to see if the Press changes because of that see if the Press changes because of that see if the Press changes because of that anyway good okay so there's like uh yeah anyway good okay so there's like uh yeah anyway good okay so there's like uh yeah that's just kind of like the generic that's just kind of like the generic that's just kind of like the generic readings at the beginning of the game I readings at the beginning of the game I readings at the beginning of the game I think that the meat comes a little bit think that the meat comes a little bit think that the meat comes a little bit later when you're starting to talk about later when you're starting to talk about later when you're starting to talk about like specific strategy and stuff like specific strategy and stuff like specific strategy and stuff I agree there are a lot of advantages to I agree there are a lot of advantages to I agree there are a lot of advantages to the two of us keeping in touch in our the two of us keeping in touch in our the two of us keeping in touch in our Nations makes strong natural allies in Nations makes strong natural allies in Nations makes strong natural allies in the middle game so that kind of stuff uh the middle game so that kind of stuff uh the middle game so that kind of stuff uh making friends making enemies yeah or making friends making enemies yeah or making friends making enemies yeah or like if you look at the next line so the like if you look at the next line so the like if you look at the next line so the person's saying like I've heard uh bits person's saying like I've heard uh bits person's saying like I've heard uh bits about a Lepanto and an octopus opening about a Lepanto and an octopus opening about a Lepanto and an octopus opening and basically telling Austria like hey and basically telling Austria like hey and basically telling Austria like hey just a heads up you know I've heard just a heads up you know I've heard just a heads up you know I've heard these whispers about like what might be these whispers about like what might be these whispers about like what might be going on behind your back yeah but so going on behind your back yeah but so going on behind your back yeah but so there's all kinds of complexities in
-
there's all kinds of complexities in there's all kinds of complexities in that that that in the in the language of that right in the in the language of that right in the in the language of that right like to interpret what that what the like to interpret what that what the like to interpret what that what the heck that means it's hard for us humans heck that means it's hard for us humans heck that means it's hard for us humans but for yeah it's even harder because but for yeah it's even harder because but for yeah it's even harder because you have to understand like at every you have to understand like at every you have to understand like at every level the the semantics of that right I level the the semantics of that right I level the the semantics of that right I mean there's there's a complexity and mean there's there's a complexity and mean there's there's a complexity and understanding when somebody is saying understanding when somebody is saying understanding when somebody is saying this to me what does that mean and then this to me what does that mean and then this to me what does that mean and then there's also the complexity of like there's also the complexity of like there's also the complexity of like should I be telling this person this should I be telling this person this should I be telling this person this like I've overheard these these Whispers like I've overheard these these Whispers like I've overheard these these Whispers should I be telling this person that should I be telling this person that should I be telling this person that like hey you might be getting attacked like hey you might be getting attacked like hey you might be getting attacked by by this other power Okay so by by this other power Okay so by by this other power Okay so what how we're supposed to think about what how we're supposed to think about what how we're supposed to think about okay so that's the natural language how okay so that's the natural language how okay so that's the natural language how do you even begin trying to solve this do you even begin trying to solve this do you even begin trying to solve this game it seems like this seems like the game it seems like this seems like the game it seems like this seems like the touring test on steroids yeah and I mean touring test on steroids yeah and I mean touring test on steroids yeah and I mean there's there's the natural language there's there's the natural language there's there's the natural language aspect and then even besides the natural aspect and then even besides the natural aspect and then even besides the natural language aspect you also have the The language aspect you also have the The language aspect you also have the The Cooperative elements of the game and I Cooperative elements of the game and I Cooperative elements of the game and I think this is actually think this is actually think this is actually um something that I find really um something that I find really um something that I find really interesting if you look at all the interesting if you look at all the interesting if you look at all the previous game AI uh breakthroughs previous game AI uh breakthroughs previous game AI uh breakthroughs they've all happened in these purely they've all happened in these purely they've all happened in these purely adversarial games where you don't adversarial games where you don't adversarial games where you don't actually need to understand how humans actually need to understand how humans actually need to understand how humans play the game it's all just AI versus AI play the game it's all just AI versus AI play the game it's all just AI versus AI right like you look at uh Checkers chess right like you look at uh Checkers chess right like you look at uh Checkers chess go poker Starcraft Dota 2 like in some go poker Starcraft Dota 2 like in some go poker Starcraft Dota 2 like in some of those cases they leveraged human data of those cases they leveraged human data of those cases they leveraged human data but they never needed to they were but they never needed to they were but they never needed to they were always just trying to have a scalable always just trying to have a scalable always just trying to have a scalable algorithm that then they could throw a algorithm that then they could throw a algorithm that then they could throw a lot of computational resources out a lot lot of computational resources out a lot lot of computational resources out a lot of memory at and then eventually it of memory at and then eventually it of memory at and then eventually it would converge to an approximation of a would converge to an approximation of a would converge to an approximation of a Nash equilibrium this Nash equilibrium this Nash equilibrium this perfect strategy that in the two player
-
perfect strategy that in the two player perfect strategy that in the two player zero some game guarantees that they're zero some game guarantees that they're zero some game guarantees that they're going to be able to not lose to any going to be able to not lose to any going to be able to not lose to any opponent so you can't leverage self-play opponent so you can't leverage self-play opponent so you can't leverage self-play to solve this game you you can leverage to solve this game you you can leverage to solve this game you you can leverage self-play but it's no longer sufficient self-play but it's no longer sufficient self-play but it's no longer sufficient to beat humans so how do you integrate to beat humans so how do you integrate to beat humans so how do you integrate the human into the loop of this so what the human into the loop of this so what the human into the loop of this so what you have to do is incorporate human data you have to do is incorporate human data you have to do is incorporate human data and to kind of give you some intuition and to kind of give you some intuition and to kind of give you some intuition for why this is the case like imagine for why this is the case like imagine for why this is the case like imagine you're playing a negotiation game like you're playing a negotiation game like you're playing a negotiation game like like diplomacy like diplomacy like diplomacy um but you're training completely from um but you're training completely from um but you're training completely from scratch scratch scratch without any human data the AI is not without any human data the AI is not without any human data the AI is not going to suddenly like figure out how to going to suddenly like figure out how to going to suddenly like figure out how to communicate in English it's going to communicate in English it's going to communicate in English it's going to figure out some weird robot language figure out some weird robot language figure out some weird robot language that only it will understand yeah and that only it will understand yeah and that only it will understand yeah and then when you stick that in a game with then when you stick that in a game with then when you stick that in a game with six other humans they're gonna think six other humans they're gonna think six other humans they're gonna think this person's talking gibberish and this person's talking gibberish and this person's talking gibberish and they're just going to Ally with each they're just going to Ally with each they're just going to Ally with each other and team up against the bot other and team up against the bot other and team up against the bot or not even team up against the ball but or not even team up against the ball but or not even team up against the ball but just not work with the bot and so in just not work with the bot and so in just not work with the bot and so in order to be able to play this game with order to be able to play this game with order to be able to play this game with humans it has to understand the human humans it has to understand the human humans it has to understand the human way of playing the game not this machine way of playing the game not this machine way of playing the game not this machine way of playing the game yeah yeah that's way of playing the game yeah yeah that's way of playing the game yeah yeah that's fascinating so right the the there's a fascinating so right the the there's a fascinating so right the the there's a nuanced thing to understand because the nuanced thing to understand because the nuanced thing to understand because the a chess playing program doesn't need to a chess playing program doesn't need to a chess playing program doesn't need to play like a human to beat a human play like a human to beat a human play like a human to beat a human exactly but here you have to play like a exactly but here you have to play like a exactly but here you have to play like a human in order to beat them or at least human in order to beat them or at least human in order to beat them or at least you have to understand how humans play you have to understand how humans play you have to understand how humans play the game so that you can understand how the game so that you can understand how the game so that you can understand how to work with them if they have certain to work with them if they have certain to work with them if they have certain expectations about what does it mean to expectations about what does it mean to expectations about what does it mean to be a good Ally what does it mean to have be a good Ally what does it mean to have be a good Ally what does it mean to have like a reciprocal relationship where like a reciprocal relationship where like a reciprocal relationship where we're working together you have to abide we're working together you have to abide we're working together you have to abide by those conventions and if you don't by those conventions and if you don't by those conventions and if you don't they're just going to work with somebody
-
they're just going to work with somebody they're just going to work with somebody else instead do you think of this as a else instead do you think of this as a else instead do you think of this as a clean in some deep sense of the spirit clean in some deep sense of the spirit clean in some deep sense of the spirit of the touring test is formulated by of the touring test is formulated by of the touring test is formulated by Alan Turing is is it in some sense this Alan Turing is is it in some sense this Alan Turing is is it in some sense this is what the Turing test actually looks is what the Turing test actually looks is what the Turing test actually looks like like like so because of open-ended natural so because of open-ended natural so because of open-ended natural language conversation seems like language conversation seems like language conversation seems like very difficult to evaluate like here at very difficult to evaluate like here at very difficult to evaluate like here at a high stakes where humans are trying to a high stakes where humans are trying to a high stakes where humans are trying to win a game that seems like how you win a game that seems like how you win a game that seems like how you actually actually actually perform the Turing test I think it's perform the Turing test I think it's perform the Turing test I think it's different from the touring test like the different from the touring test like the different from the touring test like the way that the touring test is formulated way that the touring test is formulated way that the touring test is formulated it's about trying to distinguish a human it's about trying to distinguish a human it's about trying to distinguish a human from a machine and seeing Oh could the from a machine and seeing Oh could the from a machine and seeing Oh could the machine uh successfully pass as a human machine uh successfully pass as a human machine uh successfully pass as a human in this adversarial setting where the in this adversarial setting where the in this adversarial setting where the eight where the player is trying to eight where the player is trying to eight where the player is trying to figure out whether it's a machine or a figure out whether it's a machine or a figure out whether it's a machine or a human whereas in diplomacy it's not human whereas in diplomacy it's not human whereas in diplomacy it's not about trying to figure out whether this about trying to figure out whether this about trying to figure out whether this player is a human or a machine it's player is a human or a machine it's player is a human or a machine it's ultimately about whether I can work with ultimately about whether I can work with ultimately about whether I can work with this player regardless of whether they this player regardless of whether they this player regardless of whether they are a human or machine and can the are a human or machine and can the are a human or machine and can the machine do that better than a human can machine do that better than a human can machine do that better than a human can yeah I'm going to think about that but yeah I'm going to think about that but yeah I'm going to think about that but that just feels like that just feels like that just feels like the implied requirement for that is for the implied requirement for that is for the implied requirement for that is for the machine to be human-like the machine to be human-like the machine to be human-like I think that's I think that's true that I think that's I think that's true that I think that's I think that's true that if you're going to play in this human if you're going to play in this human if you're going to play in this human game game game you have to somehow adapt to the to the you have to somehow adapt to the to the you have to somehow adapt to the to the human surroundings and the human human surroundings and the human human surroundings and the human playstyle and to win you have to adapt playstyle and to win you have to adapt playstyle and to win you have to adapt so you can't if you're the outsider so you can't if you're the outsider so you can't if you're the outsider if you're not human-like I feel like
-
if you're not human-like I feel like if you're not human-like I feel like that's a losing strategy I think that's that's a losing strategy I think that's that's a losing strategy I think that's I think that's correct yeah yeah so okay I think that's correct yeah yeah so okay I think that's correct yeah yeah so okay uh uh uh what what are the complexities here what what what are the complexities here what what what are the complexities here what was your approach to it before I get to was your approach to it before I get to was your approach to it before I get to that one thing I should explain like why that one thing I should explain like why that one thing I should explain like why we decided to work on diplomacy so we decided to work on diplomacy so we decided to work on diplomacy so basically what happened is in 2019 basically what happened is in 2019 basically what happened is in 2019 um I was wrapping up the work on six um I was wrapping up the work on six um I was wrapping up the work on six player poker on pluribus and was trying player poker on pluribus and was trying player poker on pluribus and was trying to think about what to work on next and to think about what to work on next and to think about what to work on next and I had been seeing like all these other I had been seeing like all these other I had been seeing like all these other breakthroughs happening in AI I mean breakthroughs happening in AI I mean breakthroughs happening in AI I mean like 2019 you have Starcraft you have like 2019 you have Starcraft you have like 2019 you have Starcraft you have Alpha star beating humans and Starcraft Alpha star beating humans and Starcraft Alpha star beating humans and Starcraft you've got the Dota 2 stuff happening at you've got the Dota 2 stuff happening at you've got the Dota 2 stuff happening at open AI you have GPT 2 or GPD 3 coming I open AI you have GPT 2 or GPD 3 coming I open AI you have GPT 2 or GPD 3 coming I think it was gpd2 at the time and it think it was gpd2 at the time and it think it was gpd2 at the time and it became clear that AI was progressing became clear that AI was progressing became clear that AI was progressing really really rapidly really really rapidly really really rapidly and people were throwing out these like and people were throwing out these like and people were throwing out these like other games about you know what should other games about you know what should other games about you know what should be the next challenge for for be the next challenge for for be the next challenge for for multi-agent AI and I just felt like we multi-agent AI and I just felt like we multi-agent AI and I just felt like we had to aim bigger had to aim bigger had to aim bigger um um um if you look at a game like chess or a if you look at a game like chess or a if you look at a game like chess or a game like go they took decades for game like go they took decades for game like go they took decades for researchers to to ultimately reach researchers to to ultimately reach researchers to to ultimately reach superhuman performance at I mean like superhuman performance at I mean like superhuman performance at I mean like chess took 40 Years of AI research go chess took 40 Years of AI research go chess took 40 Years of AI research go took another 20 years took another 20 years took another 20 years um and um and um and we we thought that diplomacy would be we we thought that diplomacy would be we we thought that diplomacy would be this incredibly difficult challenge that this incredibly difficult challenge that this incredibly difficult challenge that could easily take a decade to make an AI could easily take a decade to make an AI could easily take a decade to make an AI that could play competently that could play competently that could play competently um but we felt like that was that was a um but we felt like that was that was a um but we felt like that was that was a goal worth aiming for goal worth aiming for goal worth aiming for um um um and so honestly I was kind of reluctant and so honestly I was kind of reluctant and so honestly I was kind of reluctant to work on it at first because I thought
-
to work on it at first because I thought to work on it at first because I thought it was like too far out of the realm of it was like too far out of the realm of it was like too far out of the realm of possibility but you know I was talking possibility but you know I was talking possibility but you know I was talking to a co-worker of mine Adam Lear and he to a co-worker of mine Adam Lear and he to a co-worker of mine Adam Lear and he was basically saying like yeah why not was basically saying like yeah why not was basically saying like yeah why not aim for it you know we'll learn some aim for it you know we'll learn some aim for it you know we'll learn some interesting things along the way and interesting things along the way and interesting things along the way and maybe it'll be possible maybe it'll be possible maybe it'll be possible um and so so we decided to go for it and um and so so we decided to go for it and um and so so we decided to go for it and I think I think it was the right choice I think I think it was the right choice I think I think it was the right choice considering just how much progress there considering just how much progress there considering just how much progress there there was in Ai and that that progress there was in Ai and that that progress there was in Ai and that that progress has continued in the years since so has continued in the years since so has continued in the years since so winning in diplomacy what does that winning in diplomacy what does that winning in diplomacy what does that really look like it means talking to six really look like it means talking to six really look like it means talking to six other players six other entities agents other players six other entities agents other players six other entities agents and convincing and convincing and convincing and convincing them of stuff that you and convincing them of stuff that you and convincing them of stuff that you want them to be convinced of like what want them to be convinced of like what want them to be convinced of like what what exactly I'm trying to get like to what exactly I'm trying to get like to what exactly I'm trying to get like to deeply understand what the problem is deeply understand what the problem is deeply understand what the problem is ultimately ultimately ultimately the problem is it's simple to to the problem is it's simple to to the problem is it's simple to to quantify right like you're going to play quantify right like you're going to play quantify right like you're going to play this game with humans and you want your this game with humans and you want your this game with humans and you want your score on average to be score on average to be score on average to be um as high as possible you know if you um as high as possible you know if you um as high as possible you know if you can say like I am winning more than any can say like I am winning more than any can say like I am winning more than any any human alive any human alive any human alive um then you're a champion diplomacy um then you're a champion diplomacy um then you're a champion diplomacy player player player um now ultimately we haven't we didn't um now ultimately we haven't we didn't um now ultimately we haven't we didn't reach that we got to human level reach that we got to human level reach that we got to human level performance we actually so we played performance we actually so we played performance we actually so we played about 40 games with with real humans about 40 games with with real humans about 40 games with with real humans online uh the bot came in second out of online uh the bot came in second out of online uh the bot came in second out of all players that played five or more all players that played five or more all players that played five or more games and um so not like number one but games and um so not like number one but games and um so not like number one but way way higher than well what was the way way higher than well what was the way way higher than well what was the expertise level are the beginners are expertise level are the beginners are expertise level are the beginners are they intermediate players Advanced they intermediate players Advanced they intermediate players Advanced players so no sense that's a great
-
players so no sense that's a great players so no sense that's a great question and so I think question and so I think question and so I think this kind of goes into how do you this kind of goes into how do you this kind of goes into how do you measure the performance in diplomacy and measure the performance in diplomacy and measure the performance in diplomacy and I would argue that when you're measuring I would argue that when you're measuring I would argue that when you're measuring performance in a game like this you performance in a game like this you performance in a game like this you don't actually want to measure it in don't actually want to measure it in don't actually want to measure it in games with all expert players uh it's games with all expert players uh it's games with all expert players uh it's kind of like if you're developing a kind of like if you're developing a kind of like if you're developing a self-driving car you don't want to self-driving car you don't want to self-driving car you don't want to measure that car on the road with a measure that car on the road with a measure that car on the road with a bunch of expert stunt drivers you want bunch of expert stunt drivers you want bunch of expert stunt drivers you want to put it on a road of like an actual to put it on a road of like an actual to put it on a road of like an actual American city and see is this car American city and see is this car American city and see is this car crashing less often than an expert crashing less often than an expert crashing less often than an expert driver would driver would driver would so so that's the metric that we've used so so that's the metric that we've used so so that's the metric that we've used we we're saying like we're going to we we're saying like we're going to we we're saying like we're going to stick this game we're gonna stick this stick this game we're gonna stick this stick this game we're gonna stick this bot in games with a wide variety of bot in games with a wide variety of bot in games with a wide variety of skill levels and then are we doing skill levels and then are we doing skill levels and then are we doing better than a strong or expert human better than a strong or expert human better than a strong or expert human player would in the same situation player would in the same situation player would in the same situation that's quite brilliant because I played that's quite brilliant because I played that's quite brilliant because I played a lot of sports in my life like as a a lot of sports in my life like as a a lot of sports in my life like as a tennis Judo whatever tennis Judo whatever tennis Judo whatever and it's it's somehow almost easier to and it's it's somehow almost easier to and it's it's somehow almost easier to go against experts almost always I don't go against experts almost always I don't go against experts almost always I don't I think they're more predictable in the I think they're more predictable in the I think they're more predictable in the quality of play the the space of quality of play the the space of quality of play the the space of strategies you're operating under is strategies you're operating under is strategies you're operating under is narrower against experts it's more fun narrower against experts it's more fun narrower against experts it's more fun it's really frustrating to go against it's really frustrating to go against it's really frustrating to go against beginners also because beginners talk beginners also because beginners talk beginners also because beginners talk trash to you when they somehow do beat trash to you when they somehow do beat trash to you when they somehow do beat you so that's a human thing that AI you so that's a human thing that AI you so that's a human thing that AI doesn't have to be worry about that but doesn't have to be worry about that but doesn't have to be worry about that but yeah the variants and strategies right yeah the variants and strategies right yeah the variants and strategies right is greater especially with natural is greater especially with natural is greater especially with natural language it's just all over the place language it's just all over the place language it's just all over the place then true yeah and honestly when you then true yeah and honestly when you then true yeah and honestly when you look at what makes a good human look at what makes a good human look at what makes a good human diplomacy player diplomacy player diplomacy player um obviously they're able to handle um obviously they're able to handle um obviously they're able to handle themselves in games with other expert themselves in games with other expert themselves in games with other expert humans but where they really shine is
-
humans but where they really shine is humans but where they really shine is when they're playing with these weak when they're playing with these weak when they're playing with these weak players and they know how to take players and they know how to take players and they know how to take advantage of the fact that they're a advantage of the fact that they're a advantage of the fact that they're a weak player that they won't be able to weak player that they won't be able to weak player that they won't be able to like pull off a stab as well or that like pull off a stab as well or that like pull off a stab as well or that they have certain Tendencies and they they have certain Tendencies and they they have certain Tendencies and they can take them under their wing and can take them under their wing and can take them under their wing and persuade them to do things that might persuade them to do things that might persuade them to do things that might not even be in their interest not even be in their interest not even be in their interest um the really good diplomacy players are um the really good diplomacy players are um the really good diplomacy players are able to to take advantage of the fact able to to take advantage of the fact able to to take advantage of the fact that there is that there are some weak that there is that there are some weak that there is that there are some weak players in the game okay so if you have players in the game okay so if you have players in the game okay so if you have to incorporate human play data how do to incorporate human play data how do to incorporate human play data how do you do that how do you do that in order you do that how do you do that in order you do that how do you do that in order to train an AI system to play diplomacy to train an AI system to play diplomacy to train an AI system to play diplomacy yeah so that's that's really the Crux of yeah so that's that's really the Crux of yeah so that's that's really the Crux of the problem how do we the problem how do we the problem how do we um leverage the benefits of self-play um leverage the benefits of self-play um leverage the benefits of self-play that have been so successful in all that have been so successful in all that have been so successful in all these other previous games while keeping these other previous games while keeping these other previous games while keeping the strategy as uh as human compatible the strategy as uh as human compatible the strategy as uh as human compatible as possible as possible as possible and so what we did is we first trained a and so what we did is we first trained a and so what we did is we first trained a language model language model language model um and then we made that language model um and then we made that language model um and then we made that language model controllable on a set of in a set of controllable on a set of in a set of controllable on a set of in a set of intents what we call intense which are intents what we call intense which are intents what we call intense which are basically like an action that we want to basically like an action that we want to basically like an action that we want to play and an action that we would like play and an action that we would like play and an action that we would like the other player to play and so this the other player to play and so this the other player to play and so this gives us a way to generate dialogue gives us a way to generate dialogue gives us a way to generate dialogue that's not just trying to imitate the that's not just trying to imitate the that's not just trying to imitate the human style human style human style um whatever a human would say in the um whatever a human would say in the um whatever a human would say in the situation but to actually give it a a an situation but to actually give it a a an situation but to actually give it a a an intent of purpose in its communication intent of purpose in its communication intent of purpose in its communication we can talk about a specific move or we we can talk about a specific move or we we can talk about a specific move or we can make a specific request and the can make a specific request and the can make a specific request and the determination of what that move is that determination of what that move is that determination of what that move is that we're discussing comes from we're discussing comes from we're discussing comes from um strategic strategic reasoning model um strategic strategic reasoning model um strategic strategic reasoning model that uses reinforcement learning and that uses reinforcement learning and that uses reinforcement learning and planning so the Computing the intents
-
planning so the Computing the intents planning so the Computing the intents for all the players for all the players for all the players how's that done just so as a starting how's that done just so as a starting how's that done just so as a starting point is that with reinforcement point is that with reinforcement point is that with reinforcement learning or is that just optimal learning or is that just optimal learning or is that just optimal determining what the optimal is for determining what the optimal is for determining what the optimal is for intents It's a combination of intents It's a combination of intents It's a combination of reinforcement learning and planning reinforcement learning and planning reinforcement learning and planning um actually very similar to how you um actually very similar to how you um actually very similar to how you approach how we approached poker and how approach how we approached poker and how approach how we approached poker and how people approached like chess and go as people approached like chess and go as people approached like chess and go as well we're using self-play and and well we're using self-play and and well we're using self-play and and search to try to figure out what are search to try to figure out what are search to try to figure out what are what is an optimal move for us and what what is an optimal move for us and what what is an optimal move for us and what is a desirable move that we would like is a desirable move that we would like is a desirable move that we would like this other player to play now the the this other player to play now the the this other player to play now the the difference between the way that we difference between the way that we difference between the way that we approached reinforcement learning and approached reinforcement learning and approached reinforcement learning and search in this game versus those search in this game versus those search in this game versus those previous games is that we have to keep previous games is that we have to keep previous games is that we have to keep it human compatible we have to it human compatible we have to it human compatible we have to understand how the other person is understand how the other person is understand how the other person is likely to play rather than just assuming likely to play rather than just assuming likely to play rather than just assuming that they're going to play like a that they're going to play like a that they're going to play like a machine and how language gets them to machine and how language gets them to machine and how language gets them to play play play um in a way that maximize the chance of um in a way that maximize the chance of um in a way that maximize the chance of following the intent you want them to following the intent you want them to following the intent you want them to follow okay how do you do that how do follow okay how do you do that how do follow okay how do you do that how do you how do you connect language to you how do you connect language to you how do you connect language to intent so the way that RL and and intent so the way that RL and and intent so the way that RL and and planning is done is actually not using planning is done is actually not using planning is done is actually not using language so we're coming up with this language so we're coming up with this language so we're coming up with this like plan for the action uh that we're like plan for the action uh that we're like plan for the action uh that we're gonna play and the other person's gonna gonna play and the other person's gonna gonna play and the other person's gonna play and then we feed that action into play and then we feed that action into play and then we feed that action into the dialogue model that will then send a the dialogue model that will then send a the dialogue model that will then send a message according to those plans so the message according to those plans so the message according to those plans so the language model there is mapping language model there is mapping language model there is mapping action to to message to message action to to message to message action to to message to message one word at a time one word at a time one word at a time uh basically one message at a time so
-
uh basically one message at a time so uh basically one message at a time so we'll we'll feed into the dialogue model we'll we'll feed into the dialogue model we'll we'll feed into the dialogue model like here are the actions that you like here are the actions that you like here are the actions that you should be discussing here's the message should be discussing here's the message should be discussing here's the message here's like the the here's like the the here's like the the content of the message that we would content of the message that we would content of the message that we would like you to send and then it will like you to send and then it will like you to send and then it will actually generate a message that actually generate a message that actually generate a message that corresponds to that okay does this corresponds to that okay does this corresponds to that okay does this actually work it works surprisingly well actually work it works surprisingly well actually work it works surprisingly well okay how okay how okay how oh man the the number of ways it oh man the the number of ways it oh man the the number of ways it probably goes horribly I would have probably goes horribly I would have probably goes horribly I would have imagined it goes horribly wrong imagined it goes horribly wrong imagined it goes horribly wrong um so how the heck is it effective at um so how the heck is it effective at um so how the heck is it effective at all I mean there are a lot of ways that all I mean there are a lot of ways that all I mean there are a lot of ways that this could fail so for example I mean this could fail so for example I mean this could fail so for example I mean you could have a situation where you're you could have a situation where you're you could have a situation where you're you're basically like you're basically like you're basically like we don't tell the the language model we don't tell the the language model we don't tell the the language model like here are the pieces of our action like here are the pieces of our action like here are the pieces of our action or the other person's action that you or the other person's action that you or the other person's action that you should be communicating and so like should be communicating and so like should be communicating and so like let's say you're about to attack let's say you're about to attack let's say you're about to attack somebody you probably don't want to tell somebody you probably don't want to tell somebody you probably don't want to tell them that you're going to attack them them that you're going to attack them them that you're going to attack them but there's nothing in the language like but there's nothing in the language like but there's nothing in the language like the language model is not very smart at the language model is not very smart at the language model is not very smart at the end of the day so it doesn't really the end of the day so it doesn't really the end of the day so it doesn't really have a way of knowing like well what have a way of knowing like well what have a way of knowing like well what should I be talking about should I tell should I be talking about should I tell should I be talking about should I tell this person I'm about to attack them or this person I'm about to attack them or this person I'm about to attack them or not not not um so we have to like develop a lot of um so we have to like develop a lot of um so we have to like develop a lot of other techniques that that deal with other techniques that that deal with other techniques that that deal with that that that um like one of the things we do for um like one of the things we do for um like one of the things we do for example is we try to calculate if I'm example is we try to calculate if I'm example is we try to calculate if I'm going to send this message what would I going to send this message what would I going to send this message what would I expect the other person to do in expect the other person to do in expect the other person to do in response so if it's a message like hey response so if it's a message like hey response so if it's a message like hey I'm going to attack you this turn I'm going to attack you this turn I'm going to attack you this turn they're probably gonna you know attack they're probably gonna you know attack they're probably gonna you know attack us or or defend against that attack and us or or defend against that attack and us or or defend against that attack and so we have a way of recognizing like hey so we have a way of recognizing like hey so we have a way of recognizing like hey sending this message is a negative sending this message is a negative sending this message is a negative expected value action and we should not expected value action and we should not expected value action and we should not send this message send this message send this message so yes for particular kinds of messages
-
so yes for particular kinds of messages so yes for particular kinds of messages you have like an extra function that you have like an extra function that you have like an extra function that does the uh estimates the value of that does the uh estimates the value of that does the uh estimates the value of that message yeah so we have these kinds of message yeah so we have these kinds of message yeah so we have these kinds of filters that like so it's a filter so filters that like so it's a filter so filters that like so it's a filter so there's a there's a good and is that there's a there's a good and is that there's a there's a good and is that filter in your network or is it rule filter in your network or is it rule filter in your network or is it rule based that's that's a that's a neural based that's that's a that's a neural based that's that's a that's a neural network so we're well it's a it's a network so we're well it's a it's a network so we're well it's a it's a combination it's a neural network but combination it's a neural network but combination it's a neural network but it's also using planning it's also using planning it's also using planning um it's trying to compute like what is um it's trying to compute like what is um it's trying to compute like what is the policy that the other players are the policy that the other players are the policy that the other players are going to play Given that this message going to play Given that this message going to play Given that this message um has been sent and then is that better um has been sent and then is that better um has been sent and then is that better than not sending the message or not I than not sending the message or not I than not sending the message or not I feel like that's how my brain works too feel like that's how my brain works too feel like that's how my brain works too like there's a language model that like there's a language model that like there's a language model that generates random crap and then there's generates random crap and then there's generates random crap and then there's these other neural Nets they're these other neural Nets they're these other neural Nets they're essentially filters at least that's when essentially filters at least that's when essentially filters at least that's when I tweet I tweet I tweet I'll usually my process of tweeting I'll I'll usually my process of tweeting I'll I'll usually my process of tweeting I'll think of something and it's hilarious to think of something and it's hilarious to think of something and it's hilarious to me and then about five seconds later the me and then about five seconds later the me and then about five seconds later the filter Network comes in and says no no filter Network comes in and says no no filter Network comes in and says no no that's not funny at all I mean there's that's not funny at all I mean there's that's not funny at all I mean there's some something interesting to that kind some something interesting to that kind some something interesting to that kind of process so you have a set of actions of process so you have a set of actions of process so you have a set of actions that you you want you have an intent that you you want you have an intent that you you want you have an intent that you want to achieve an intent that that you want to achieve an intent that that you want to achieve an intent that you want your opponent to achieve then you want your opponent to achieve then you want your opponent to achieve then you generate messages and then you you generate messages and then you you generate messages and then you evaluated those messages will achieve evaluated those messages will achieve evaluated those messages will achieve the the the the the the the the the uh the goal you want yeah and we're uh the goal you want yeah and we're uh the goal you want yeah and we're filtering for several things we're filtering for several things we're filtering for several things we're filtering like is this a sensible filtering like is this a sensible filtering like is this a sensible message you know so sometimes language message you know so sometimes language message you know so sometimes language models will send will generate messages models will send will generate messages models will send will generate messages that are just like totally nonsense that are just like totally nonsense that are just like totally nonsense um and we try to filter those out we um and we try to filter those out we um and we try to filter those out we also try to filter out messages that
-
also try to filter out messages that also try to filter out messages that that are basically lies that are basically lies that are basically lies um so you know diplomacy has this um so you know diplomacy has this um so you know diplomacy has this reputation as a game that's really about reputation as a game that's really about reputation as a game that's really about um deception and lying but we try to um deception and lying but we try to um deception and lying but we try to actually minimize the amount that the actually minimize the amount that the actually minimize the amount that the bot would lie bot would lie bot would lie um this was actually mostly or are you um this was actually mostly or are you um this was actually mostly or are you no I'm just kidding okay no I'm just kidding okay no I'm just kidding okay I mean like part of the reason for this I mean like part of the reason for this I mean like part of the reason for this is that we actually found that lying is that we actually found that lying is that we actually found that lying would make the bot perform worse in the would make the bot perform worse in the would make the bot perform worse in the long run it would end up with a lower long run it would end up with a lower long run it would end up with a lower score because once the bot lies score because once the bot lies score because once the bot lies um people would never trust it again um people would never trust it again um people would never trust it again and and trust is a huge aspect of the and and trust is a huge aspect of the and and trust is a huge aspect of the game of diplomacy taking notes here game of diplomacy taking notes here game of diplomacy taking notes here because I think this is applies to because I think this is applies to because I think this is applies to to life lessons too oh I think it's a to life lessons too oh I think it's a to life lessons too oh I think it's a really yeah really strong so like lying really yeah really strong so like lying really yeah really strong so like lying is a dangerous thing to do like you you is a dangerous thing to do like you you is a dangerous thing to do like you you want to avoid want to avoid want to avoid obvious lying yeah I mean I think when obvious lying yeah I mean I think when obvious lying yeah I mean I think when people play diplomacy for the first time people play diplomacy for the first time people play diplomacy for the first time they approach it as a game of deception they approach it as a game of deception they approach it as a game of deception and lying and and they and lying and and they and lying and and they ultimately if you talk to top diplomacy ultimately if you talk to top diplomacy ultimately if you talk to top diplomacy players what they'll tell you is that players what they'll tell you is that players what they'll tell you is that diplomacy is a game about trust and diplomacy is a game about trust and diplomacy is a game about trust and being able to build trust in an being able to build trust in an being able to build trust in an environment that encourages people to environment that encourages people to environment that encourages people to not trust anyone not trust anyone not trust anyone so so that's the ultimate tension in so so that's the ultimate tension in so so that's the ultimate tension in diplomacy how can this AI reason about diplomacy how can this AI reason about diplomacy how can this AI reason about whether you are being honest in your whether you are being honest in your whether you are being honest in your communication and how can the AI communication and how can the AI communication and how can the AI persuade you that it is being honest persuade you that it is being honest persuade you that it is being honest when it is telling you that hey I'm when it is telling you that hey I'm when it is telling you that hey I'm actually going to support you this turn actually going to support you this turn actually going to support you this turn is there some sense I don't know if you is there some sense I don't know if you is there some sense I don't know if you step back and think that this process step back and think that this process step back and think that this process well well well indirectly help us study human indirectly help us study human indirectly help us study human psychology
-
psychology psychology so like if trust is the ultimate goal so like if trust is the ultimate goal so like if trust is the ultimate goal wouldn't that help us understand what wouldn't that help us understand what wouldn't that help us understand what are the fundamental aspects of forming are the fundamental aspects of forming are the fundamental aspects of forming trust between humans and between humans trust between humans and between humans trust between humans and between humans and AI I mean that's a really really and AI I mean that's a really really and AI I mean that's a really really important question that's much bigger important question that's much bigger important question that's much bigger than the strategy games it's how can than the strategy games it's how can than the strategy games it's how can that that's fundamental to the human that that's fundamental to the human that that's fundamental to the human robot interaction problem how do we form robot interaction problem how do we form robot interaction problem how do we form Trust Trust Trust between intelligent entities between intelligent entities between intelligent entities so one of the things I'm really excited so one of the things I'm really excited so one of the things I'm really excited about with diplomacy about with diplomacy about with diplomacy um there's never really been a good um there's never really been a good um there's never really been a good domain to investigate these kinds of domain to investigate these kinds of domain to investigate these kinds of questions yeah questions yeah questions yeah um and diplomacy gives us a domain where um and diplomacy gives us a domain where um and diplomacy gives us a domain where trust is really at the center of it trust is really at the center of it trust is really at the center of it um and it's not just like you've hired a um and it's not just like you've hired a um and it's not just like you've hired a bunch of mechanical turkers that you bunch of mechanical turkers that you bunch of mechanical turkers that you know are being paid and trying to get know are being paid and trying to get know are being paid and trying to get through the task as quickly as possible through the task as quickly as possible through the task as quickly as possible you have these people that are really you have these people that are really you have these people that are really invested in the outcome of the game and invested in the outcome of the game and invested in the outcome of the game and they're really trying to do the best they're really trying to do the best they're really trying to do the best that they can that they can that they can um and so I'm really excited that we're um and so I'm really excited that we're um and so I'm really excited that we're able to we actually like have put able to we actually like have put able to we actually like have put together this we're open sourcing all of together this we're open sourcing all of together this we're open sourcing all of our models we're open sourcing uh all of our models we're open sourcing uh all of our models we're open sourcing uh all of the all the code and we're making the the all the code and we're making the the all the code and we're making the data that we've used available to data that we've used available to data that we've used available to researchers researchers researchers um so that they can investigate these um so that they can investigate these um so that they can investigate these kinds of questions so the data of the kinds of questions so the data of the kinds of questions so the data of the different the human and the AI play of different the human and the AI play of different the human and the AI play of diplomacy and the models that you use diplomacy and the models that you use diplomacy and the models that you use for the generation of the messages and for the generation of the messages and for the generation of the messages and the filtering yeah not not just even the the filtering yeah not not just even the the filtering yeah not not just even the data of the AI playing with the humans data of the AI playing with the humans data of the AI playing with the humans but all the training data that we that but all the training data that we that but all the training data that we that we had that we used to train the AI to we had that we used to train the AI to we had that we used to train the AI to understand how humans play the game
-
understand how humans play the game understand how humans play the game we're setting up a system where we're setting up a system where we're setting up a system where researchers will be able to apply researchers will be able to apply researchers will be able to apply um to be able to gain access to that um to be able to gain access to that um to be able to gain access to that data and be be able to to use it in data and be be able to to use it in data and be be able to to use it in their own research we should say what is their own research we should say what is their own research we should say what is the name of the system the name of the system the name of the system we're calling the bot Cicero Cicero and we're calling the bot Cicero Cicero and we're calling the bot Cicero Cicero and what's the name like you're open what's the name like you're open what's the name like you're open sourcing what's the name of the sourcing what's the name of the sourcing what's the name of the repository and and the like the the repository and and the like the the repository and and the like the the project is it also just called Cicero project is it also just called Cicero project is it also just called Cicero the big project or are you still coming the big project or are you still coming the big project or are you still coming up with the name the the data set comes up with the name the the data set comes up with the name the the data set comes from this website web diplomacy.net is from this website web diplomacy.net is from this website web diplomacy.net is this site that's been online for like 20 this site that's been online for like 20 this site that's been online for like 20 years now and uh it's one of the main years now and uh it's one of the main years now and uh it's one of the main sites that people use to play diplomacy sites that people use to play diplomacy sites that people use to play diplomacy on it we've got like 50 000 games of on it we've got like 50 000 games of on it we've got like 50 000 games of diplomacy with you know natural language diplomacy with you know natural language diplomacy with you know natural language communication communication communication um over 10 million messages so it's a um over 10 million messages so it's a um over 10 million messages so it's a pretty massive data set that people can pretty massive data set that people can pretty massive data set that people can use to um we're hoping that the the use to um we're hoping that the the use to um we're hoping that the the academic Community the research academic Community the research academic Community the research Community is able to use it for for all Community is able to use it for for all Community is able to use it for for all sorts of interesting research questions sorts of interesting research questions sorts of interesting research questions so do you from having studied this game so do you from having studied this game so do you from having studied this game is this is this is this a sufficiently rich problem space to a sufficiently rich problem space to a sufficiently rich problem space to explore this kind of human AI explore this kind of human AI explore this kind of human AI interaction yeah absolutely and I think interaction yeah absolutely and I think interaction yeah absolutely and I think it's it's it's I think it's maybe the best data set I think it's maybe the best data set I think it's maybe the best data set that I can think of out there to to that I can think of out there to to that I can think of out there to to investigate these kinds of questions of investigate these kinds of questions of investigate these kinds of questions of um negotiation trust um negotiation trust um negotiation trust um persuasion I wouldn't say it's the um persuasion I wouldn't say it's the um persuasion I wouldn't say it's the best data set in the world for best data set in the world for best data set in the world for um human AI interaction that's a very um human AI interaction that's a very um human AI interaction that's a very broad field but I think that it's broad field but I think that it's broad field but I think that it's definitely up there is like you know if definitely up there is like you know if definitely up there is like you know if you're really interested in language you're really interested in language you're really interested in language models interacting with humans in you models interacting with humans in you models interacting with humans in you know a setting where their incentives know a setting where their incentives know a setting where their incentives are not fully aligned this seems like an
-
are not fully aligned this seems like an are not fully aligned this seems like an ideal data set for investigating that ideal data set for investigating that ideal data set for investigating that so you have so you have so you have um you have a paper with some impressive um you have a paper with some impressive um you have a paper with some impressive results and just an impressive paper results and just an impressive paper results and just an impressive paper they're taking this problem on they're taking this problem on they're taking this problem on what's the most exciting thing to you in what's the most exciting thing to you in what's the most exciting thing to you in terms of the results from the the paper terms of the results from the the paper terms of the results from the the paper well I think there's ideas or results well I think there's ideas or results well I think there's ideas or results yeah I think there's a few aspects of yeah I think there's a few aspects of yeah I think there's a few aspects of the results and um that I think are the results and um that I think are the results and um that I think are really exciting so first of all the fact really exciting so first of all the fact really exciting so first of all the fact that we were able to achieve such strong that we were able to achieve such strong that we were able to achieve such strong performance performance performance um I was um I was um I was surprised by and pleasantly surprised by surprised by and pleasantly surprised by surprised by and pleasantly surprised by um so we played 40 games of diplomacy um so we played 40 games of diplomacy um so we played 40 games of diplomacy with real humans and the bot placed with real humans and the bot placed with real humans and the bot placed second out of all players that have second out of all players that have second out of all players that have played five or more games so it's about played five or more games so it's about played five or more games so it's about 80 players total 80 players total 80 players total um 19 of whom played five or more games um 19 of whom played five or more games um 19 of whom played five or more games and the bot was ranked second out of and the bot was ranked second out of and the bot was ranked second out of those players those players those players um and the bot was was really good in um and the bot was was really good in um and the bot was was really good in two Dimensions one being able to two Dimensions one being able to two Dimensions one being able to establish strong connections with the establish strong connections with the establish strong connections with the other players on the board being able to other players on the board being able to other players on the board being able to like persuade them to work with it like persuade them to work with it like persuade them to work with it um being able to coordinate with them um being able to coordinate with them um being able to coordinate with them about like how it's going to work with about like how it's going to work with about like how it's going to work with them and then also the Raw them and then also the Raw them and then also the Raw tactical and strategic aspects of the tactical and strategic aspects of the tactical and strategic aspects of the game you know being able to understand game you know being able to understand game you know being able to understand what the other players are likely to do what the other players are likely to do what the other players are likely to do being able to model their behavior and being able to model their behavior and being able to model their behavior and respond appropriately to that the bot respond appropriately to that the bot respond appropriately to that the bot also really excelled at what are some also really excelled at what are some also really excelled at what are some interesting things that the bot said interesting things that the bot said interesting things that the bot said by the way are you allowed to swear in by the way are you allowed to swear in by the way are you allowed to swear in the um okay are there rules to what the um okay are there rules to what the um okay are there rules to what you're allowed to say and not in you're allowed to say and not in you're allowed to say and not in diplomacy you can say whatever you want
-
diplomacy you can say whatever you want diplomacy you can say whatever you want I think the site will get very angry at I think the site will get very angry at I think the site will get very angry at you if you start like threatening you if you start like threatening you if you start like threatening somebody and if we actually like if somebody and if we actually like if somebody and if we actually like if you're threaten somebody you're supposed you're threaten somebody you're supposed you're threaten somebody you're supposed to do it politely yeah politely you know to do it politely yeah politely you know to do it politely yeah politely you know keep it in character keep it in character keep it in character um um um we actually had a researcher watching we actually had a researcher watching we actually had a researcher watching the bot 24 7 for well whenever we play a the bot 24 7 for well whenever we play a the bot 24 7 for well whenever we play a game we had a bot watching it to make game we had a bot watching it to make game we had a bot watching it to make sure that it wouldn't go off the rails sure that it wouldn't go off the rails sure that it wouldn't go off the rails and start like threatening somebody or and start like threatening somebody or and start like threatening somebody or something like that I would just love it something like that I would just love it something like that I would just love it if the boss started like mocking if the boss started like mocking if the boss started like mocking mocking everybody like some weird quirky mocking everybody like some weird quirky mocking everybody like some weird quirky strategies would emerge have you seen strategies would emerge have you seen strategies would emerge have you seen anything interesting that you huh that's anything interesting that you huh that's anything interesting that you huh that's a weird that's a a weird that's a a weird that's a that's a behavior either of the filter that's a behavior either of the filter that's a behavior either of the filter or the language model or the language model or the language model that was weird to you that was yeah they that was weird to you that was yeah they that was weird to you that was yeah they were definitely like things that the bot were definitely like things that the bot were definitely like things that the bot would would do that were not in line would would do that were not in line would would do that were not in line with like how humans would approach the with like how humans would approach the with like how humans would approach the game and that in a good way the humans game and that in a good way the humans game and that in a good way the humans actually you know we we've talked to actually you know we we've talked to actually you know we we've talked to some expert diplomacy players about some expert diplomacy players about some expert diplomacy players about these results and their takeaways that these results and their takeaways that these results and their takeaways that well maybe humans are approaching this well maybe humans are approaching this well maybe humans are approaching this the wrong way and this is actually like the wrong way and this is actually like the wrong way and this is actually like the right way to play the game the right way to play the game the right way to play the game um so what's required to win like what um so what's required to win like what um so what's required to win like what um what does it mean to mess up or to um what does it mean to mess up or to um what does it mean to mess up or to exploit the sub-optimal behavior of a exploit the sub-optimal behavior of a exploit the sub-optimal behavior of a player like uh is there is there player like uh is there is there player like uh is there is there optimally rational behavior and optimally rational behavior and optimally rational behavior and irrational behavior that you need to irrational behavior that you need to irrational behavior that you need to estimate that kind of stuff like what estimate that kind of stuff like what estimate that kind of stuff like what what stands out to you like is there a what stands out to you like is there a what stands out to you like is there a crack that you can exploit is there like crack that you can exploit is there like crack that you can exploit is there like um a weakness that you can exploit in um a weakness that you can exploit in um a weakness that you can exploit in the game that that everybody's looking the game that that everybody's looking the game that that everybody's looking for
-
for for well I I think well I I think well I I think you're asking kind of two questions you're asking kind of two questions you're asking kind of two questions there so one like modeling the there so one like modeling the there so one like modeling the irrationality and the suboptimality of irrationality and the suboptimality of irrationality and the suboptimality of humans humans humans um um um you can't in diplomacy you can't treat you can't in diplomacy you can't treat you can't in diplomacy you can't treat all the other players like they're all the other players like they're all the other players like they're machines and if you do that you're machines and if you do that you're machines and if you do that you're you're going to end up playing really you're going to end up playing really you're going to end up playing really poorly and so we actually ran this poorly and so we actually ran this poorly and so we actually ran this experiment so we we trained a bot in a experiment so we we trained a bot in a experiment so we we trained a bot in a two-player zero-sum version of diplomacy two-player zero-sum version of diplomacy two-player zero-sum version of diplomacy um the same way that you might approach um the same way that you might approach um the same way that you might approach a game like chess or poker and the bot a game like chess or poker and the bot a game like chess or poker and the bot was superhuman it would crush any was superhuman it would crush any was superhuman it would crush any competitor and then we took that same competitor and then we took that same competitor and then we took that same training approach and we trained a bot training approach and we trained a bot training approach and we trained a bot for the full Seven Player version of the for the full Seven Player version of the for the full Seven Player version of the game through self-play without any human game through self-play without any human game through self-play without any human data and we stuck it in a game with six data and we stuck it in a game with six data and we stuck it in a game with six humans and it got destroyed even in the humans and it got destroyed even in the humans and it got destroyed even in the version of the game where there's no version of the game where there's no version of the game where there's no explicit natural language communication explicit natural language communication explicit natural language communication it still got destroyed because it just it still got destroyed because it just it still got destroyed because it just wouldn't be able to understand how the wouldn't be able to understand how the wouldn't be able to understand how the other players were approaching the game other players were approaching the game other players were approaching the game and be able to to work with that and be able to to work with that and be able to to work with that can you just Linger on that meeting like can you just Linger on that meeting like can you just Linger on that meeting like there's an individual there's an there's an individual there's an there's an individual there's an individual personality each player and individual personality each player and individual personality each player and then you're supposed to remember that then you're supposed to remember that then you're supposed to remember that but Woody means it's not able to but Woody means it's not able to but Woody means it's not able to understand the the players well it would understand the the players well it would understand the the players well it would for example expect the human to support for example expect the human to support for example expect the human to support it in a certain way when the human men it in a certain way when the human men it in a certain way when the human men would simply like think like no I'm not would simply like think like no I'm not would simply like think like no I'm not supposed to support you here supposed to support you here supposed to support you here um it's kind of like you know if you um it's kind of like you know if you um it's kind of like you know if you develop a self-driving car and it's develop a self-driving car and it's develop a self-driving car and it's trained completely from scratch with trained completely from scratch with trained completely from scratch with other self-driving cars it might learn other self-driving cars it might learn other self-driving cars it might learn to drive on the left side of the road to drive on the left side of the road to drive on the left side of the road that's a totally reasonable thing to do that's a totally reasonable thing to do that's a totally reasonable thing to do if you're with these other self-driving if you're with these other self-driving if you're with these other self-driving cars that are also driving on the left cars that are also driving on the left cars that are also driving on the left side of the road but if you put it in an side of the road but if you put it in an side of the road but if you put it in an American city it's gonna crash but I
-
American city it's gonna crash but I American city it's gonna crash but I guess the intuition I'm trying to build guess the intuition I'm trying to build guess the intuition I'm trying to build up is why does it then crush a human up is why does it then crush a human up is why does it then crush a human play on heads up play on heads up play on heads up this is multiple this is an aspect of this is multiple this is an aspect of this is multiple this is an aspect of two player zero song versus games that two player zero song versus games that two player zero song versus games that involve cooperation so in a two-player involve cooperation so in a two-player involve cooperation so in a two-player zero-sum game zero-sum game zero-sum game um you can do self-play from scratch and um you can do self-play from scratch and um you can do self-play from scratch and you will arrive at the Nash equilibrium you will arrive at the Nash equilibrium you will arrive at the Nash equilibrium where you don't have to worry about the where you don't have to worry about the where you don't have to worry about the other player other player other player playing in a very human sub-optimal playing in a very human sub-optimal playing in a very human sub-optimal style that's just going to be that the style that's just going to be that the style that's just going to be that the only way that deviating from an ash only way that deviating from an ash only way that deviating from an ash equilibrium equilibrium equilibrium would would change things is if it would would change things is if it would would change things is if it helped you so I what's the dynamic of helped you so I what's the dynamic of helped you so I what's the dynamic of cooperation that's effective in cooperation that's effective in cooperation that's effective in diplomacy diplomacy diplomacy do you always have to to have one friend do you always have to to have one friend do you always have to to have one friend in the game you always want to maximize in the game you always want to maximize in the game you always want to maximize your friends and minimize your enemies got it and got it and boy in the the lying comes into play boy in the the lying comes into play boy in the the lying comes into play there there there so the more friends you have the better so the more friends you have the better so the more friends you have the better yeah I mean I guess you have to attack yeah I mean I guess you have to attack yeah I mean I guess you have to attack somebody or else you're not going to somebody or else you're not going to somebody or else you're not going to make progress all right so that's the make progress all right so that's the make progress all right so that's the tension but man this is too real this is tension but man this is too real this is tension but man this is too real this is too real to this is too too close to too real to this is too too close to too real to this is too too close to geopolitics of actual military conflict geopolitics of actual military conflict geopolitics of actual military conflict in the world okay in the world okay in the world okay uh that's fascinating so that uh that's fascinating so that uh that's fascinating so that cooperation element is what makes the cooperation element is what makes the cooperation element is what makes the game really really hard yeah and to give game really really hard yeah and to give game really really hard yeah and to give you an example of of how this you an example of of how this you an example of of how this sub-optimality and irrationality comes sub-optimality and irrationality comes sub-optimality and irrationality comes into play there's a really common into play there's a really common into play there's a really common situation in a game of diplomacy situation in a game of diplomacy situation in a game of diplomacy um that where one player starts to win
-
um that where one player starts to win um that where one player starts to win and they're like at the point where and they're like at the point where and they're like at the point where they're controlling about half the map they're controlling about half the map they're controlling about half the map yeah um and the remaining players who yeah um and the remaining players who yeah um and the remaining players who have all been fighting each other the have all been fighting each other the have all been fighting each other the whole game all have to like work whole game all have to like work whole game all have to like work together now to stop this other player together now to stop this other player together now to stop this other player from winning or else everybody's gonna from winning or else everybody's gonna from winning or else everybody's gonna lose lose lose um and it's kind of like you know Game um and it's kind of like you know Game um and it's kind of like you know Game of Thrones like I don't know if you've of Thrones like I don't know if you've of Thrones like I don't know if you've seen the show like you know you got the seen the show like you know you got the seen the show like you know you got the the others coming from the north and the others coming from the north and the others coming from the north and like all the people have to start work like all the people have to start work like all the people have to start work out their differences and stop them from out their differences and stop them from out their differences and stop them from from taking over from taking over from taking over um um um and the bot will do this like the bot and the bot will do this like the bot and the bot will do this like the bot will work with the other players to stop will work with the other players to stop will work with the other players to stop the superpower from winning but if it the superpower from winning but if it the superpower from winning but if it doesn't really if it's trained from doesn't really if it's trained from doesn't really if it's trained from scratch or it doesn't really have a good scratch or it doesn't really have a good scratch or it doesn't really have a good grounding in how humans approach it it grounding in how humans approach it it grounding in how humans approach it it will also at the same time attack the will also at the same time attack the will also at the same time attack the other players with its extra units so other players with its extra units so other players with its extra units so all the units that are not necessary to all the units that are not necessary to all the units that are not necessary to stop the superpower from winning it will stop the superpower from winning it will stop the superpower from winning it will use those to grab as many centers as use those to grab as many centers as use those to grab as many centers as possible from the other players and possible from the other players and possible from the other players and in totally rational play the other in totally rational play the other in totally rational play the other players should just live with that you players should just live with that you players should just live with that you know they have to understand like hey a know they have to understand like hey a know they have to understand like hey a score of one is better than a score of score of one is better than a score of score of one is better than a score of zero so zero so zero so um so okay he's grabbed my centers but I um so okay he's grabbed my centers but I um so okay he's grabbed my centers but I I'll just deal with it I'll just deal with it I'll just deal with it but humans don't act that way right the but humans don't act that way right the but humans don't act that way right the human gets really angry at the bot and human gets really angry at the bot and human gets really angry at the bot and ends up throwing the game because you ends up throwing the game because you ends up throwing the game because you know I'm gonna screw you over because know I'm gonna screw you over because know I'm gonna screw you over because you did something that's not fair to me you did something that's not fair to me you did something that's not fair to me got it and are you supposed to model got it and are you supposed to model got it and are you supposed to model that is the boss supposed to model that that is the boss supposed to model that that is the boss supposed to model that kind of human frustration yeah exactly kind of human frustration yeah exactly kind of human frustration yeah exactly and so that is something that seems and so that is something that seems and so that is something that seems almost impossible to model purely from almost impossible to model purely from almost impossible to model purely from scratch without any human data it's a scratch without any human data it's a scratch without any human data it's a very cultural thing yeah um and so you
-
very cultural thing yeah um and so you very cultural thing yeah um and so you need need need human data to be able to understand that human data to be able to understand that human data to be able to understand that hey that's how humans behave and you hey that's how humans behave and you hey that's how humans behave and you have to work around that it might be have to work around that it might be have to work around that it might be suboptimal it might be rational but but suboptimal it might be rational but but suboptimal it might be rational but but that's an aspect of humanity that you that's an aspect of humanity that you that's an aspect of humanity that you have to have you have to deal with so have to have you have to deal with so have to have you have to deal with so how difficult is it to train on human how difficult is it to train on human how difficult is it to train on human data given that human data is very data given that human data is very data given that human data is very limited versus what the US a purely limited versus what the US a purely limited versus what the US a purely self-play mechanism can generate that's self-play mechanism can generate that's self-play mechanism can generate that's actually one of the major challenges actually one of the major challenges actually one of the major challenges that we faced in the research that we that we faced in the research that we that we faced in the research that we had a good amount of human data we had had a good amount of human data we had had a good amount of human data we had about 50 000 games what we try to do is about 50 000 games what we try to do is about 50 000 games what we try to do is leverage as much soft play as possible leverage as much soft play as possible leverage as much soft play as possible while still leveraging the human data so while still leveraging the human data so while still leveraging the human data so what we do is we do self-play very what we do is we do self-play very what we do is we do self-play very similar to how it's been done in poker similar to how it's been done in poker similar to how it's been done in poker and go but we try to regularize the and go but we try to regularize the and go but we try to regularize the self-play towards the human data self-play towards the human data self-play towards the human data basically the way to think about it is basically the way to think about it is basically the way to think about it is um we penalize the bot for choosing um we penalize the bot for choosing um we penalize the bot for choosing actions that are very unlikely under how actions that are very unlikely under how actions that are very unlikely under how under the human data set under the human data set under the human data set and how do you know is there is this and how do you know is there is this and how do you know is there is this some kind of function that says this is some kind of function that says this is some kind of function that says this is human-like enough yeah so we we train a human-like enough yeah so we we train a human-like enough yeah so we we train a bot through supervised learning to model bot through supervised learning to model bot through supervised learning to model the human play as much as possible so we the human play as much as possible so we the human play as much as possible so we basically like train a neural net basically like train a neural net basically like train a neural net um on those 50 000 games and that gives um on those 50 000 games and that gives um on those 50 000 games and that gives us an approximate that gives us a policy us an approximate that gives us a policy us an approximate that gives us a policy that resembles to some extent how humans that resembles to some extent how humans that resembles to some extent how humans actually play the game now this isn't a actually play the game now this isn't a actually play the game now this isn't a perfect model of human play because we perfect model of human play because we perfect model of human play because we don't have unlimited data we don't have don't have unlimited data we don't have don't have unlimited data we don't have unlimited neural net capacity unlimited neural net capacity unlimited neural net capacity um but it gives us some approximation uh um but it gives us some approximation uh um but it gives us some approximation uh is there some data on the internet
-
is there some data on the internet is there some data on the internet that's useful besides just diplomacy so that's useful besides just diplomacy so that's useful besides just diplomacy so on the language side of things is there on the language side of things is there on the language side of things is there some can you go to like Reddit some can you go to like Reddit some can you go to like Reddit and and and um so sort of background model um so sort of background model um so sort of background model formulation that that's useful for the formulation that that's useful for the formulation that that's useful for the game of diplomacy yeah absolutely and so game of diplomacy yeah absolutely and so game of diplomacy yeah absolutely and so for the language model which for the language model which for the language model which um it's kind of like a separate question um it's kind of like a separate question um it's kind of like a separate question you know we didn't use the language you know we didn't use the language you know we didn't use the language model during self-play training but we model during self-play training but we model during self-play training but we pre-trained the language model on you pre-trained the language model on you pre-trained the language model on you know tons of internet data as much as know tons of internet data as much as know tons of internet data as much as possible and then we fine-tuned it possible and then we fine-tuned it possible and then we fine-tuned it specifically on the diplomacy games so specifically on the diplomacy games so specifically on the diplomacy games so we are able to like Leverage The Wider we are able to like Leverage The Wider we are able to like Leverage The Wider data set in order to fill in data set in order to fill in data set in order to fill in some of the gaps in like how some of the gaps in like how some of the gaps in like how communication happens more broadly communication happens more broadly communication happens more broadly besides just like specifically in these besides just like specifically in these besides just like specifically in these diplomacy games Okay cool so what what's diplomacy games Okay cool so what what's diplomacy games Okay cool so what what's some what are some interesting things some what are some interesting things some what are some interesting things that came to life from this from this that came to life from this from this that came to life from this from this work uh to you like what are some work uh to you like what are some work uh to you like what are some insights insights insights about about about um um um about games where natural language is about games where natural language is about games where natural language is involved and cooperation deep involved and cooperation deep involved and cooperation deep cooperation is involved well I think cooperation is involved well I think cooperation is involved well I think there's a few insights um so first of there's a few insights um so first of there's a few insights um so first of all all all the fact that you can't rely purely or the fact that you can't rely purely or the fact that you can't rely purely or even largely on self-play that you even largely on self-play that you even largely on self-play that you really have to have an understanding of really have to have an understanding of really have to have an understanding of how humans approach the game how humans approach the game how humans approach the game um I think that that's one of the major um I think that that's one of the major um I think that that's one of the major conclusions that I'm drawing from this conclusions that I'm drawing from this conclusions that I'm drawing from this work work work um and that is I think applicable more um and that is I think applicable more um and that is I think applicable more broadly to a lot of different games so broadly to a lot of different games so broadly to a lot of different games so we've actually already taken the we've actually already taken the we've actually already taken the approaches that we've used in diplomacy approaches that we've used in diplomacy approaches that we've used in diplomacy and tried them on uh Cooperative card and tried them on uh Cooperative card and tried them on uh Cooperative card game called Hanabi and we've had a lot game called Hanabi and we've had a lot game called Hanabi and we've had a lot of success in that game as well
-
of success in that game as well of success in that game as well um on the language side um on the language side um on the language side I think the fact that we were able to I think the fact that we were able to I think the fact that we were able to control the language model through this control the language model through this control the language model through this intense approach was very effective intense approach was very effective intense approach was very effective um and it allowed us instead of just um and it allowed us instead of just um and it allowed us instead of just imitating how humans would communicate imitating how humans would communicate imitating how humans would communicate were able to go beyond that and able to were able to go beyond that and able to were able to go beyond that and able to feed into it superhuman strategies that feed into it superhuman strategies that feed into it superhuman strategies that it can then um you know generate it can then um you know generate it can then um you know generate messages corresponding to messages corresponding to messages corresponding to is there something you could say about is there something you could say about is there something you could say about detecting whether a person or AI is detecting whether a person or AI is detecting whether a person or AI is lying or not lying or not lying or not the bot doesn't explicitly try to the bot doesn't explicitly try to the bot doesn't explicitly try to calculate whether somebody is lying or calculate whether somebody is lying or calculate whether somebody is lying or not but what it will do is try to not but what it will do is try to not but what it will do is try to predict what actions they're going to predict what actions they're going to predict what actions they're going to take given the communications given the take given the communications given the take given the communications given the messages that they've sent to us so messages that they've sent to us so messages that they've sent to us so given our conversation what do I think given our conversation what do I think given our conversation what do I think you're going to do and implicitly there you're going to do and implicitly there you're going to do and implicitly there is a calculation about whether you're is a calculation about whether you're is a calculation about whether you're lying to me in that lying to me in that lying to me in that you know if if you're based on your you know if if you're based on your you know if if you're based on your messages if I think you're going to messages if I think you're going to messages if I think you're going to attack me this turn attack me this turn attack me this turn um even though your messages say that um even though your messages say that um even though your messages say that you're not then you know essentially the you're not then you know essentially the you're not then you know essentially the bot is predicting that you're lying but bot is predicting that you're lying but bot is predicting that you're lying but it doesn't view it as as lying the same it doesn't view it as as lying the same it doesn't view it as as lying the same way that we would view it as lying way that we would view it as lying way that we would view it as lying but you could probably reformulate with but you could probably reformulate with but you could probably reformulate with all the same data and make a classifier all the same data and make a classifier all the same data and make a classifier lying or not yeah I think I think you lying or not yeah I think I think you lying or not yeah I think I think you could do that um that was not something could do that um that was not something could do that um that was not something that we were focused on but I think that that we were focused on but I think that that we were focused on but I think that it is possible that you know if you came it is possible that you know if you came it is possible that you know if you came up with some measurements of like what up with some measurements of like what up with some measurements of like what does it mean to tell a lie because does it mean to tell a lie because does it mean to tell a lie because there's there's a spectrum right like if
-
there's there's a spectrum right like if there's there's a spectrum right like if you're withholding some information is you're withholding some information is you're withholding some information is that a lie that a lie that a lie um if you're mostly telling the truth um if you're mostly telling the truth um if you're mostly telling the truth but you forgot to mention this like one but you forgot to mention this like one but you forgot to mention this like one action out of like 10 is that a lie action out of like 10 is that a lie action out of like 10 is that a lie um it's hard to draw the line but you um it's hard to draw the line but you um it's hard to draw the line but you know if you're willing to do that and know if you're willing to do that and know if you're willing to do that and then you could possibly use it to uh then you could possibly use it to uh then you could possibly use it to uh dude this feels like an argument inside dude this feels like an argument inside dude this feels like an argument inside a relationship now what constitutes a a relationship now what constitutes a a relationship now what constitutes a lie lie lie um depends what you mean by the um depends what you mean by the um depends what you mean by the definition of the word is okay definition of the word is okay definition of the word is okay um um um still it's fascinating because trust and still it's fascinating because trust and still it's fascinating because trust and lying is all intermixed into this and lying is all intermixed into this and lying is all intermixed into this and it's language models that are becoming it's language models that are becoming it's language models that are becoming more and more sophisticated it's just a more and more sophisticated it's just a more and more sophisticated it's just a fascinating space to explore fascinating space to explore fascinating space to explore um um um what do you see as the future of this what do you see as the future of this what do you see as the future of this work work work um that is inspired by the Breakthrough um that is inspired by the Breakthrough um that is inspired by the Breakthrough performance that you're getting here performance that you're getting here performance that you're getting here with diplomacy uh uh I think there's a few different I think there's a few different I think there's a few different directions to take this work directions to take this work directions to take this work um um um I I think really what it's showing us is I I think really what it's showing us is I I think really what it's showing us is the potential that language models have the potential that language models have the potential that language models have I mean I think a lot of people didn't I mean I think a lot of people didn't I mean I think a lot of people didn't think that this kind of result was think that this kind of result was think that this kind of result was possible even today despite all the possible even today despite all the possible even today despite all the progress that's been made in language progress that's been made in language progress that's been made in language models and so it shows us how we can models and so it shows us how we can models and so it shows us how we can Leverage The Power of things like Leverage The Power of things like Leverage The Power of things like self-play on top of language models to self-play on top of language models to self-play on top of language models to get get get um increasingly better performance and um increasingly better performance and um increasingly better performance and the ceiling is really the ceiling is really the ceiling is really much higher than what we have right now much higher than what we have right now much higher than what we have right now is this transferable somehow to is this transferable somehow to is this transferable somehow to to chat Bots
-
to chat Bots to chat Bots for the more General task of dialogue for the more General task of dialogue for the more General task of dialogue so because there is a kind of so because there is a kind of so because there is a kind of negotiation here a dance between negotiation here a dance between negotiation here a dance between entities that are trying to cooperate entities that are trying to cooperate entities that are trying to cooperate and at the same time a little bit and at the same time a little bit and at the same time a little bit adversarial which I think Maps somewhat adversarial which I think Maps somewhat adversarial which I think Maps somewhat to the general to the general to the general you know the entire process of Reddit or you know the entire process of Reddit or you know the entire process of Reddit or like internet communication you're like internet communication you're like internet communication you're cooperating you're adversarial you're cooperating you're adversarial you're cooperating you're adversarial you're having debates you're having uh having debates you're having uh having debates you're having uh camaraderie all that kind of stuff camaraderie all that kind of stuff camaraderie all that kind of stuff I think one of the things that's really I think one of the things that's really I think one of the things that's really useful about diplomacy is that we have a useful about diplomacy is that we have a useful about diplomacy is that we have a well-defined value function there is a well-defined value function there is a well-defined value function there is a well-defined score that the bot is well-defined score that the bot is well-defined score that the bot is trying to optimize and and in a in a trying to optimize and and in a in a trying to optimize and and in a in a setting like a general chatbot setting setting like a general chatbot setting setting like a general chatbot setting it needs it would need that kind of it needs it would need that kind of it needs it would need that kind of um objective in order to fully leverage um objective in order to fully leverage um objective in order to fully leverage the techniques that we've developed the techniques that we've developed the techniques that we've developed what about like what we talked about what about like what we talked about what about like what we talked about earlier with NPCs inside video games earlier with NPCs inside video games earlier with NPCs inside video games like how can it be used to create like how can it be used to create like how can it be used to create for Elder Scrolls 6 more compelling um for Elder Scrolls 6 more compelling um for Elder Scrolls 6 more compelling um NPCs NPCs NPCs that you could talk to instead of that you could talk to instead of that you could talk to instead of instead of committing all kinds of instead of committing all kinds of instead of committing all kinds of violence with a sword and fighting violence with a sword and fighting violence with a sword and fighting dragons just sitting in a Tavern and dragons just sitting in a Tavern and dragons just sitting in a Tavern and drink all day and talk to the chatbot drink all day and talk to the chatbot drink all day and talk to the chatbot the way that we've approached AI the way that we've approached AI the way that we've approached AI diplomacy is you condition the language diplomacy is you condition the language diplomacy is you condition the language on an intent now that intent and on an intent now that intent and on an intent now that intent and diplomacy is an is an action but it diplomacy is an is an action but it diplomacy is an is an action but it doesn't have to be and you can imagine doesn't have to be and you can imagine doesn't have to be and you can imagine you know you could have NPCs in video you know you could have NPCs in video you know you could have NPCs in video games or the metaverse or whatever where
-
games or the metaverse or whatever where games or the metaverse or whatever where there's some intent or there's some there's some intent or there's some there's some intent or there's some objective that they're trying to objective that they're trying to objective that they're trying to maximize and you can specify what that maximize and you can specify what that maximize and you can specify what that is is is um and and then the language can um and and then the language can um and and then the language can correspond to that intent now I'm not correspond to that intent now I'm not correspond to that intent now I'm not saying that this is you know happening saying that this is you know happening saying that this is you know happening imminently but um I'm saying that this imminently but um I'm saying that this imminently but um I'm saying that this is like a future application potentially is like a future application potentially is like a future application potentially of this direction of research so what's of this direction of research so what's of this direction of research so what's the more General formulation of this the more General formulation of this the more General formulation of this making self-play be able to scale the making self-play be able to scale the making self-play be able to scale the way self-play does and still maintain way self-play does and still maintain way self-play does and still maintain human-like Behavior human-like Behavior human-like Behavior the way that we've approached self-play the way that we've approached self-play the way that we've approached self-play in diplomacy is like in diplomacy is like in diplomacy is like we're we're trying to we're we're trying to we're we're trying to come up with good intents to condition come up with good intents to condition come up with good intents to condition the language model on and the space of the language model on and the space of the language model on and the space of intents is actions that can be played in intents is actions that can be played in intents is actions that can be played in the game now there is like the potential the game now there is like the potential the game now there is like the potential to have a broader set of intents things to have a broader set of intents things to have a broader set of intents things like you know long-term cooperation or like you know long-term cooperation or like you know long-term cooperation or long-term uh objectives or you know long-term uh objectives or you know long-term uh objectives or you know gossip about what another player was gossip about what another player was gossip about what another player was saying saying saying um these are things that we're currently um these are things that we're currently um these are things that we're currently not conditioning the language model on not conditioning the language model on not conditioning the language model on and so it's not able to we're not able and so it's not able to we're not able and so it's not able to we're not able to control it to say like oh you should to control it to say like oh you should to control it to say like oh you should be talking about this thing right now be talking about this thing right now be talking about this thing right now but it's quite possible that you could but it's quite possible that you could but it's quite possible that you could expand the scope of intents to be able expand the scope of intents to be able expand the scope of intents to be able to allow it to talk about those things to allow it to talk about those things to allow it to talk about those things now in the process of doing that the now in the process of doing that the now in the process of doing that the self-play would become much more self-play would become much more self-play would become much more complicated complicated complicated um and so that is a potential for for um and so that is a potential for for um and so that is a potential for for future work okay the increasing the future work okay the increasing the future work okay the increasing the number of intents I still am not quite number of intents I still am not quite number of intents I still am not quite clear clear clear how you keep the self-play how you keep the self-play how you keep the self-play integrated into the human world yeah I'm integrated into the human world yeah I'm integrated into the human world yeah I'm a little bit loose on the uh uh on
-
a little bit loose on the uh uh on a little bit loose on the uh uh on understanding how you do that so we understanding how you do that so we understanding how you do that so we train a neural Nets to train a neural Nets to train a neural Nets to um imitate the human data as closely as um imitate the human data as closely as um imitate the human data as closely as possible and that's what we call the possible and that's what we call the possible and that's what we call the anchor policy and now when we're doing anchor policy and now when we're doing anchor policy and now when we're doing self-play self-play self-play the the problem with the anchor policy the the problem with the anchor policy the the problem with the anchor policy is that it's not a perfect approximation is that it's not a perfect approximation is that it's not a perfect approximation of how humans actually play because we of how humans actually play because we of how humans actually play because we don't have infinite data because we don't have infinite data because we don't have infinite data because we don't have unlimited neural network don't have unlimited neural network don't have unlimited neural network capacity it's actually a relatively capacity it's actually a relatively capacity it's actually a relatively sub-optimal approximation of how humans sub-optimal approximation of how humans sub-optimal approximation of how humans actually play and we can improve that actually play and we can improve that actually play and we can improve that approximation by adding planning and RL approximation by adding planning and RL approximation by adding planning and RL and so what we do is we get a better and so what we do is we get a better and so what we do is we get a better approximation a better model of human approximation a better model of human approximation a better model of human play by play by play by during the self-play process we say you during the self-play process we say you during the self-play process we say you can deviate from this human anchor can deviate from this human anchor can deviate from this human anchor policy if there is an action that has policy if there is an action that has policy if there is an action that has you know particularly High expected you know particularly High expected you know particularly High expected value value value um but it would have to be a really high um but it would have to be a really high um but it would have to be a really high expected value in order to to deviate expected value in order to to deviate expected value in order to to deviate from from this human-like policy so you from from this human-like policy so you from from this human-like policy so you basically say try to maximize your basically say try to maximize your basically say try to maximize your expected value while at the same time expected value while at the same time expected value while at the same time stay as close as possible to the human stay as close as possible to the human stay as close as possible to the human policy and there is a parameter that policy and there is a parameter that policy and there is a parameter that controls those the the relative controls those the the relative controls those the the relative weighting of those Computing objectives weighting of those Computing objectives weighting of those Computing objectives so the question I have so the question I have so the question I have is how sophisticated can the anchor is how sophisticated can the anchor is how sophisticated can the anchor policy get policy get policy get to have a policy that approximates human to have a policy that approximates human to have a policy that approximates human behavior right yeah so as you increase behavior right yeah so as you increase behavior right yeah so as you increase the number of intents as you generalize the number of intents as you generalize the number of intents as you generalize the the space in which this is the the space in which this is the the space in which this is applicable applicable applicable and given that the human data is limited
-
and given that the human data is limited and given that the human data is limited try to anticipate a policy that works try to anticipate a policy that works try to anticipate a policy that works for in a much larger number of cases for in a much larger number of cases for in a much larger number of cases like how how difficult is the process of like how how difficult is the process of like how how difficult is the process of forming a damn good anchor policy forming a damn good anchor policy forming a damn good anchor policy well it really comes down to how much well it really comes down to how much well it really comes down to how much human data you have so it's all boss human data you have so it's all boss human data you have so it's all boss scale in the human data I think the more scale in the human data I think the more scale in the human data I think the more human data you have the better and I human data you have the better and I human data you have the better and I think that that's going to be the major think that that's going to be the major think that that's going to be the major bottleneck in in scaling to to more bottleneck in in scaling to to more bottleneck in in scaling to to more complicated complicated complicated um domains but that said um domains but that said um domains but that said you know there might be the potential you know there might be the potential you know there might be the potential just like in the language model where we just like in the language model where we just like in the language model where we leveraged you know tons of data on the leveraged you know tons of data on the leveraged you know tons of data on the internet and then specialized it for internet and then specialized it for internet and then specialized it for diplomacy diplomacy diplomacy um there is the future potential that um there is the future potential that um there is the future potential that you can leverage huge amounts of data you can leverage huge amounts of data you can leverage huge amounts of data across the board and then specialize it across the board and then specialize it across the board and then specialize it in the data set that you have for in the data set that you have for in the data set that you have for diplomacy and in that way you're diplomacy and in that way you're diplomacy and in that way you're essentially augmenting the amount of essentially augmenting the amount of essentially augmenting the amount of data that you have data that you have data that you have to what degree does this apply to what degree does this apply to what degree does this apply to the general the real world diplomacy to the general the real world diplomacy to the general the real world diplomacy the geopolitics the geopolitics the geopolitics you know there's a game theory has a you know there's a game theory has a you know there's a game theory has a history of being applied to understand history of being applied to understand history of being applied to understand and to give us hope about nuclear and to give us hope about nuclear and to give us hope about nuclear weapons for example the mutually assured weapons for example the mutually assured weapons for example the mutually assured destruction is a game theoretic concept destruction is a game theoretic concept destruction is a game theoretic concept that you can formulate some people say that you can formulate some people say that you can formulate some people say it's oversimplified but nevertheless it's oversimplified but nevertheless it's oversimplified but nevertheless here we are and we somehow haven't blown here we are and we somehow haven't blown here we are and we somehow haven't blown ourselves up do you see a future where ourselves up do you see a future where ourselves up do you see a future where this kind of this kind of this kind of this kind of system can be used to help this kind of system can be used to help this kind of system can be used to help us make decisions geopolitical decisions us make decisions geopolitical decisions us make decisions geopolitical decisions in the world in the world in the world well like I said the original motivation well like I said the original motivation well like I said the original motivation for the game of diplomacy was the
-
for the game of diplomacy was the for the game of diplomacy was the failures of World War One The Diplomatic failures of World War One The Diplomatic failures of World War One The Diplomatic failures that led to War uh and the real failures that led to War uh and the real failures that led to War uh and the real take-home message of diplomacy is that take-home message of diplomacy is that take-home message of diplomacy is that you know if people approach diplomacy you know if people approach diplomacy you know if people approach diplomacy the right way then war is ultimately the right way then war is ultimately the right way then war is ultimately unsuccessful unsuccessful unsuccessful um the way that I see at war is an um the way that I see at war is an um the way that I see at war is an inherently negative sum game right inherently negative sum game right inherently negative sum game right there's always a better outcome than War there's always a better outcome than War there's always a better outcome than War for all the parties involved and my hope for all the parties involved and my hope for all the parties involved and my hope is that you know as AI progresses then is that you know as AI progresses then is that you know as AI progresses then maybe this technology could be used to maybe this technology could be used to maybe this technology could be used to help people make better decisions help people make better decisions help people make better decisions um across the board and you know um across the board and you know um across the board and you know hopefully avoid negative some outcomes hopefully avoid negative some outcomes hopefully avoid negative some outcomes like War like War like War yeah I mean I just came back from yeah I mean I just came back from yeah I mean I just came back from Ukraine I'm going back there on deep Ukraine I'm going back there on deep Ukraine I'm going back there on deep personal levels personal levels personal levels think a lot about think a lot about think a lot about how peace can be achieved and I'm a big how peace can be achieved and I'm a big how peace can be achieved and I'm a big believer in conversation or leaders believer in conversation or leaders believer in conversation or leaders getting together and having getting together and having getting together and having conversations and trying to understand conversations and trying to understand conversations and trying to understand each other each other each other yeah it's fascinating to think um yeah it's fascinating to think um yeah it's fascinating to think um whether each one of those leaders can whether each one of those leaders can whether each one of those leaders can run a simulation ahead of time like if run a simulation ahead of time like if run a simulation ahead of time like if I'm an I'm an I'm an what are the possible consequences if what are the possible consequences if what are the possible consequences if I'm nice what are the possible I'm nice what are the possible I'm nice what are the possible consequences consequences consequences um my guess um my guess um my guess is that if the president of the United is that if the president of the United is that if the president of the United States got together with uh States got together with uh States got together with uh Vladimir zielinski and Vladimir Putin Vladimir zielinski and Vladimir Putin Vladimir zielinski and Vladimir Putin that there will be significant benefits that there will be significant benefits that there will be significant benefits to to to um the president United States not
-
um the president United States not um the president United States not having an ego having an ego having an ego of kind of of kind of of kind of playing down of giving away a lot of playing down of giving away a lot of playing down of giving away a lot of chips for the future successful world so chips for the future successful world so chips for the future successful world so giving a lot of power to the two giving a lot of power to the two giving a lot of power to the two presidents of the competing Nations to presidents of the competing Nations to presidents of the competing Nations to achieve peace that's my guess but it'd achieve peace that's my guess but it'd achieve peace that's my guess but it'd be nice to run a bunch of simulations be nice to run a bunch of simulations be nice to run a bunch of simulations but then you have to have human data but then you have to have human data but then you have to have human data right you really because it's like the right you really because it's like the right you really because it's like the game of diplomacy is fundamentally game of diplomacy is fundamentally game of diplomacy is fundamentally different than geopolitics you need data different than geopolitics you need data different than geopolitics you need data you need like I guess that's the you need like I guess that's the you need like I guess that's the question I have like how transferable is question I have like how transferable is question I have like how transferable is this to uh like I don't know any kind of this to uh like I don't know any kind of this to uh like I don't know any kind of negotiation right like to any kind of negotiation right like to any kind of negotiation right like to any kind of look some local I don't know a bunch of look some local I don't know a bunch of look some local I don't know a bunch of lawyers like arguing like at a divorce lawyers like arguing like at a divorce lawyers like arguing like at a divorce like divorce lawyers like how like divorce lawyers like how like divorce lawyers like how transferable this all kinds of human transferable this all kinds of human transferable this all kinds of human negotiation well I feel like this isn't negotiation well I feel like this isn't negotiation well I feel like this isn't a question that's unique to diplomacy I a question that's unique to diplomacy I a question that's unique to diplomacy I mean I think you look at RL mean I think you look at RL mean I think you look at RL breakthroughs reinforcement learning breakthroughs reinforcement learning breakthroughs reinforcement learning breakthroughs in previous games as well breakthroughs in previous games as well breakthroughs in previous games as well like you know AI for Starcraft AI for like you know AI for Starcraft AI for like you know AI for Starcraft AI for Atari you haven't really seen it Atari you haven't really seen it Atari you haven't really seen it deployed in the real world because you deployed in the real world because you deployed in the real world because you have these problems of it's really hard have these problems of it's really hard have these problems of it's really hard to collect a lot of data to collect a lot of data to collect a lot of data um and you don't have a you don't have a um and you don't have a you don't have a um and you don't have a you don't have a well-defined action space you don't have well-defined action space you don't have well-defined action space you don't have a well-defined reward function these are a well-defined reward function these are a well-defined reward function these are all things that you really need for all things that you really need for all things that you really need for reinforcement learning and planning to reinforcement learning and planning to reinforcement learning and planning to be really successful today now there are be really successful today now there are be really successful today now there are some domains where you do have that some domains where you do have that some domains where you do have that um code generation is one example um code generation is one example um code generation is one example theorem proving mathematics that's theorem proving mathematics that's theorem proving mathematics that's another example where you have a another example where you have a another example where you have a well-defined action space you have a well-defined action space you have a well-defined action space you have a well-defined reward function and those well-defined reward function and those well-defined reward function and those are the kinds of domains where I can see
-
are the kinds of domains where I can see are the kinds of domains where I can see RL in the short term being incredibly RL in the short term being incredibly RL in the short term being incredibly powerful but powerful but powerful but yeah I think that those are the barriers yeah I think that those are the barriers yeah I think that those are the barriers to deploying this at scale in the real to deploying this at scale in the real to deploying this at scale in the real world but and the hope is that in the world but and the hope is that in the world but and the hope is that in the long run we'll be able to get there yeah long run we'll be able to get there yeah long run we'll be able to get there yeah but you see diplomacy feels like closer but you see diplomacy feels like closer but you see diplomacy feels like closer to the real world than does Starcraft to the real world than does Starcraft to the real world than does Starcraft like because it's natural language right like because it's natural language right like because it's natural language right you're operating in a space of intense you're operating in a space of intense you're operating in a space of intense and in a space of natural language that and in a space of natural language that and in a space of natural language that feels very close to the real world and feels very close to the real world and feels very close to the real world and it also feels like you could get data on it also feels like you could get data on it also feels like you could get data on that from the internet that from the internet that from the internet yeah and that's why I do think that yeah and that's why I do think that yeah and that's why I do think that diplomacy is taking a big step closer to diplomacy is taking a big step closer to diplomacy is taking a big step closer to the real world than anything that's came the real world than anything that's came the real world than anything that's came before in terms of game AI breakthroughs before in terms of game AI breakthroughs before in terms of game AI breakthroughs the fact that the fact that the fact that you know we're we're communicating in you know we're we're communicating in you know we're we're communicating in natural language we've we're leveraging natural language we've we're leveraging natural language we've we're leveraging the fact that we have this like General the fact that we have this like General the fact that we have this like General data set of uh dialogue and data set of uh dialogue and data set of uh dialogue and communication from a breadth of the communication from a breadth of the communication from a breadth of the internet internet internet um that is that is a big step in that um that is that is a big step in that um that is that is a big step in that direction we're not 100 there but um but direction we're not 100 there but um but direction we're not 100 there but um but we're getting closer at least we're getting closer at least we're getting closer at least so if we actually return back to Poker so if we actually return back to Poker so if we actually return back to Poker and chess are some of the ideas that and chess are some of the ideas that and chess are some of the ideas that you're learning here with diplomacy you're learning here with diplomacy you're learning here with diplomacy could you construct AI systems that play could you construct AI systems that play could you construct AI systems that play like humans like humans like humans like um make for a fun opponent like um make for a fun opponent like um make for a fun opponent in a game with Jess yeah absolutely in a game with Jess yeah absolutely in a game with Jess yeah absolutely we've already started looking into this we've already started looking into this we've already started looking into this direction a bit so we tried to use the direction a bit so we tried to use the direction a bit so we tried to use the techniques that we've developed for techniques that we've developed for techniques that we've developed for diplomacy uh to make chess and go AIS diplomacy uh to make chess and go AIS diplomacy uh to make chess and go AIS and what we found is that it led to much and what we found is that it led to much and what we found is that it led to much more human-like strong chess and go
-
more human-like strong chess and go more human-like strong chess and go players the way that players the way that players the way that AIS like stockfish today play is in a AIS like stockfish today play is in a AIS like stockfish today play is in a very inhuman style it's very strong but very inhuman style it's very strong but very inhuman style it's very strong but it's very different from how humans play it's very different from how humans play it's very different from how humans play and so we can take the techniques that and so we can take the techniques that and so we can take the techniques that we've developed for diplomacy we do we've developed for diplomacy we do we've developed for diplomacy we do something similar in um in chess and go something similar in um in chess and go something similar in um in chess and go and we end up with the bot that's both and we end up with the bot that's both and we end up with the bot that's both strong and human-like strong and human-like strong and human-like um to elaborate on this a bit like one um to elaborate on this a bit like one um to elaborate on this a bit like one way to approach making a human-like AI way to approach making a human-like AI way to approach making a human-like AI for chess is to collect a bunch of human for chess is to collect a bunch of human for chess is to collect a bunch of human games like a bunch of human Grand Master games like a bunch of human Grand Master games like a bunch of human Grand Master games and just do supervised learning on games and just do supervised learning on games and just do supervised learning on those games but the problem is that if those games but the problem is that if those games but the problem is that if you do that what you end up with is an you do that what you end up with is an you do that what you end up with is an AI That's substantially weaker than the AI That's substantially weaker than the AI That's substantially weaker than the human Grand Masters that you've trained human Grand Masters that you've trained human Grand Masters that you've trained on because the neural net is not able to on because the neural net is not able to on because the neural net is not able to approximate the the Nuance of the approximate the the Nuance of the approximate the the Nuance of the strategy this goes back to the planning strategy this goes back to the planning strategy this goes back to the planning thing that I mentioned the search thing thing that I mentioned the search thing thing that I mentioned the search thing that I talked about before that these that I talked about before that these that I talked about before that these human Grand Masters when they're playing human Grand Masters when they're playing human Grand Masters when they're playing they're using search and they're using they're using search and they're using they're using search and they're using planning and the neural net alone unless planning and the neural net alone unless planning and the neural net alone unless you have a massive neural net that's you have a massive neural net that's you have a massive neural net that's like a thousand times bigger than what like a thousand times bigger than what like a thousand times bigger than what we have right now it's not able to we have right now it's not able to we have right now it's not able to approximate those details very approximate those details very approximate those details very effectively effectively effectively and on the other hand you can leverage and on the other hand you can leverage and on the other hand you can leverage search and planning very heavily but search and planning very heavily but search and planning very heavily but then what you end up with is an AI that then what you end up with is an AI that then what you end up with is an AI that plays in a very different style from how plays in a very different style from how plays in a very different style from how humans play the game humans play the game humans play the game now if you strike this intermediate now if you strike this intermediate now if you strike this intermediate balance by setting the um the balance by setting the um the balance by setting the um the regularization parameters correctly and regularization parameters correctly and regularization parameters correctly and say you can do planning but try to keep
-
say you can do planning but try to keep say you can do planning but try to keep it close to the human policy then you it close to the human policy then you it close to the human policy then you end up with an AI that plays in both a end up with an AI that plays in both a end up with an AI that plays in both a very human-like style and a very strong very human-like style and a very strong very human-like style and a very strong style and you can actually even tune it style and you can actually even tune it style and you can actually even tune it to have a certain ELO rating so you can to have a certain ELO rating so you can to have a certain ELO rating so you can say play in the style of like a 2800 say play in the style of like a 2800 say play in the style of like a 2800 elohuman elohuman elohuman um I wonder if you could do specific um I wonder if you could do specific um I wonder if you could do specific type of humans or categories of humans type of humans or categories of humans type of humans or categories of humans so not just skill but Style so not just skill but Style so not just skill but Style yeah I think so and so this is this is yeah I think so and so this is this is yeah I think so and so this is this is where the the research gets interesting where the the research gets interesting where the the research gets interesting like you know one of the things that I like you know one of the things that I like you know one of the things that I was thinking about is and this is was thinking about is and this is was thinking about is and this is actually already being done I think actually already being done I think actually already being done I think there's a researcher at the University there's a researcher at the University there's a researcher at the University of Toronto that's working on this of Toronto that's working on this of Toronto that's working on this um is to make an ad that plays in the um is to make an ad that plays in the um is to make an ad that plays in the style of a particular player like Magnus style of a particular player like Magnus style of a particular player like Magnus Carlson for example you can make an AI Carlson for example you can make an AI Carlson for example you can make an AI that plays like Magnus Carlson and then that plays like Magnus Carlson and then that plays like Magnus Carlson and then where I think this gets interesting is where I think this gets interesting is where I think this gets interesting is like maybe you're up against Magnus like maybe you're up against Magnus like maybe you're up against Magnus Carlson in the world championship or Carlson in the world championship or Carlson in the world championship or something you can play against this something you can play against this something you can play against this Magnus Carlson bot to prepare against Magnus Carlson bot to prepare against Magnus Carlson bot to prepare against the real Magnus Carlson and you can try the real Magnus Carlson and you can try the real Magnus Carlson and you can try to explore strategies that he might to explore strategies that he might to explore strategies that he might struggle with struggle with struggle with um and try to figure out like how do you um and try to figure out like how do you um and try to figure out like how do you beat this player in particular beat this player in particular beat this player in particular um on the other hand you can also have um on the other hand you can also have um on the other hand you can also have Magnus Carlson working with this bot to Magnus Carlson working with this bot to Magnus Carlson working with this bot to try to figure out where he's weak um and try to figure out where he's weak um and try to figure out where he's weak um and where he needs to improve his strategy where he needs to improve his strategy where he needs to improve his strategy um and so I can Envision this future um and so I can Envision this future um and so I can Envision this future where data on specific chess and go where data on specific chess and go where data on specific chess and go players becomes extremely valuable players becomes extremely valuable players becomes extremely valuable because you can use that data to create because you can use that data to create because you can use that data to create specific models of how these particular specific models of how these particular specific models of how these particular players play so increasingly human-like players play so increasingly human-like players play so increasingly human-like behavior and Bots however behavior and Bots however behavior and Bots however as you've mentioned makes cheating cheat
-
as you've mentioned makes cheating cheat as you've mentioned makes cheating cheat detection much harder it it does yeah detection much harder it it does yeah detection much harder it it does yeah the way that sheet detection Works in a the way that sheet detection Works in a the way that sheet detection Works in a game like poker and a game like chess game like poker and a game like chess game like poker and a game like chess and go from what I understand is trying and go from what I understand is trying and go from what I understand is trying to see like is this person making moves to see like is this person making moves to see like is this person making moves that are very common among chess AIS or that are very common among chess AIS or that are very common among chess AIS or you know AIS in general you know AIS in general you know AIS in general um but very uncommon among top human um but very uncommon among top human um but very uncommon among top human players players players and if you have the development of these and if you have the development of these and if you have the development of these AIS that play in a very strong style but AIS that play in a very strong style but AIS that play in a very strong style but also a very human-like style then that also a very human-like style then that also a very human-like style then that poses serious challenges for cheat poses serious challenges for cheat poses serious challenges for cheat detection and it makes you now ask detection and it makes you now ask detection and it makes you now ask yourself a hard question about what is yourself a hard question about what is yourself a hard question about what is the role of AI systems as they become the role of AI systems as they become the role of AI systems as they become more and more integrated in our society more and more integrated in our society more and more integrated in our society and this kind of human AI and this kind of human AI and this kind of human AI um um um integration has has some deep ethical integration has has some deep ethical integration has has some deep ethical issues that we should be aware of and issues that we should be aware of and issues that we should be aware of and also it's a kind of cyber security also it's a kind of cyber security also it's a kind of cyber security challenge right for to make you know one challenge right for to make you know one challenge right for to make you know one of the assumptions we have when we play of the assumptions we have when we play of the assumptions we have when we play games is that there's a trust that is games is that there's a trust that is games is that there's a trust that is only humans involved and there only humans involved and there only humans involved and there the better AI systems to create which the better AI systems to create which the better AI systems to create which makes it super exciting human-like AI makes it super exciting human-like AI makes it super exciting human-like AI systems with different styles of humans systems with different styles of humans systems with different styles of humans is really exciting but then we have to is really exciting but then we have to is really exciting but then we have to have the defenses better and better and have the defenses better and better and have the defenses better and better and better if we're to trust that we uh can better if we're to trust that we uh can better if we're to trust that we uh can enjoy human versus human game in a enjoy human versus human game in a enjoy human versus human game in a deeply Fair way it's fascinating so it's deeply Fair way it's fascinating so it's deeply Fair way it's fascinating so it's just uh it's humbling yeah I think there just uh it's humbling yeah I think there just uh it's humbling yeah I think there is a lot of like negative potential for
-
is a lot of like negative potential for is a lot of like negative potential for this kind of Technology but you know at this kind of Technology but you know at this kind of Technology but you know at the same time there's a lot of upside the same time there's a lot of upside the same time there's a lot of upside for it as well so you know for example for it as well so you know for example for it as well so you know for example right now it's really hard to learn how right now it's really hard to learn how right now it's really hard to learn how to get better in games like chess and to get better in games like chess and to get better in games like chess and poker and go because the way that the AI poker and go because the way that the AI poker and go because the way that the AI plays is so foreign and incomprehensible plays is so foreign and incomprehensible plays is so foreign and incomprehensible but if you have these AIS that are but if you have these AIS that are but if you have these AIS that are playing you know you can say like Oh I'm playing you know you can say like Oh I'm playing you know you can say like Oh I'm a 2000 ELO human how do I get to 2200 a 2000 ELO human how do I get to 2200 a 2000 ELO human how do I get to 2200 now you can have an AI that plays in the now you can have an AI that plays in the now you can have an AI that plays in the style of a 2200 elohiman and that will style of a 2200 elohiman and that will style of a 2200 elohiman and that will help you get better or you know you help you get better or you know you help you get better or you know you mentioned this problem of like how do mentioned this problem of like how do mentioned this problem of like how do you know that you're actually playing you know that you're actually playing you know that you're actually playing with humans when you're playing like with humans when you're playing like with humans when you're playing like online in video games well now we have online in video games well now we have online in video games well now we have the potential of populating these like the potential of populating these like the potential of populating these like Virtual Worlds with Virtual Worlds with Virtual Worlds with um agents like AI agents that are um agents like AI agents that are um agents like AI agents that are actually fun to play with and you don't actually fun to play with and you don't actually fun to play with and you don't have to always be playing with other have to always be playing with other have to always be playing with other humans to to you know have a fun time humans to to you know have a fun time humans to to you know have a fun time so yeah a lot a lot of upside potential so yeah a lot a lot of upside potential so yeah a lot a lot of upside potential too and I think you know with any sort too and I think you know with any sort too and I think you know with any sort of tool there's there's the potential of tool there's there's the potential of tool there's there's the potential for a lot of greatness and a lot of uh for a lot of greatness and a lot of uh for a lot of greatness and a lot of uh downsides as well so in the paper they downsides as well so in the paper they downsides as well so in the paper they got a chance to look at there's a got a chance to look at there's a got a chance to look at there's a section on uh ethical considerations section on uh ethical considerations section on uh ethical considerations what's in that section what are some what's in that section what are some what's in that section what are some ethical considerations here is it some ethical considerations here is it some ethical considerations here is it some of the stuff we've already talked about of the stuff we've already talked about of the stuff we've already talked about there's some things that we've already there's some things that we've already there's some things that we've already talked about talked about talked about um I think um I think um I think specific to diplomacy you know there's specific to diplomacy you know there's specific to diplomacy you know there's there's also the the challenge that the there's also the the challenge that the there's also the the challenge that the game is game is game is you know there is a deception aspect to you know there is a deception aspect to you know there is a deception aspect to the game the game the game um and so um and so um and so you know have developing language models you know have developing language models you know have developing language models that are capable of deception is I think that are capable of deception is I think that are capable of deception is I think a dicey issue and something that you a dicey issue and something that you a dicey issue and something that you know makes research on diplomacy
-
know makes research on diplomacy know makes research on diplomacy particularly challenging particularly challenging particularly challenging um and um and um and you know so so those kinds of issues of you know so so those kinds of issues of you know so so those kinds of issues of like should we even be developing AIS like should we even be developing AIS like should we even be developing AIS that are capable of lying to people that are capable of lying to people that are capable of lying to people that's something that we have to you that's something that we have to you that's something that we have to you know think carefully about uh that's so know think carefully about uh that's so know think carefully about uh that's so cool I mean that you have to do that cool I mean that you have to do that cool I mean that you have to do that kind of stuff in order to figure out kind of stuff in order to figure out kind of stuff in order to figure out where the ethical lines are but I can where the ethical lines are but I can where the ethical lines are but I can see in the future it being illegal to see in the future it being illegal to see in the future it being illegal to have a consumer product that have a consumer product that have a consumer product that lies lies lies yeah yeah like your personal assistant yeah yeah like your personal assistant yeah yeah like your personal assistant AI system is not a lot is always have to AI system is not a lot is always have to AI system is not a lot is always have to tell the truth but if if I ask it do I tell the truth but if if I ask it do I tell the truth but if if I ask it do I do I look did I get fatter over the past do I look did I get fatter over the past do I look did I get fatter over the past month I sure as hell want that AI system month I sure as hell want that AI system month I sure as hell want that AI system to lie to me uh so there's a trade-off to lie to me uh so there's a trade-off to lie to me uh so there's a trade-off between lying and being and being nice between lying and being and being nice between lying and being and being nice after somehow find after somehow find after somehow find where's the ethics in that and we're where's the ethics in that and we're where's the ethics in that and we're back to discussions inside relationships back to discussions inside relationships back to discussions inside relationships anyway what were you saying oh yeah I anyway what were you saying oh yeah I anyway what were you saying oh yeah I was getting like yeah this yeah that's was getting like yeah this yeah that's was getting like yeah this yeah that's kind of going to the question of like kind of going to the question of like kind of going to the question of like what what is a lot you know is a white what what is a lot you know is a white what what is a lot you know is a white lie a bad lie is it an ethical lie yeah lie a bad lie is it an ethical lie yeah lie a bad lie is it an ethical lie yeah you know those kinds of questions uh boy you know those kinds of questions uh boy you know those kinds of questions uh boy we return time and time again to deep we return time and time again to deep we return time and time again to deep human questions as we design AI systems human questions as we design AI systems human questions as we design AI systems that's exactly what they do they put a that's exactly what they do they put a that's exactly what they do they put a mirror to humanity to help us understand mirror to humanity to help us understand mirror to humanity to help us understand ourselves ourselves ourselves there's there's also the issue of like there's there's also the issue of like there's there's also the issue of like you know in these diplomacy experiments you know in these diplomacy experiments you know in these diplomacy experiments in order to do in order to do in order to do a fair comparison you know what we found a fair comparison you know what we found a fair comparison you know what we found is that there's an inherent anti-ai bias is that there's an inherent anti-ai bias is that there's an inherent anti-ai bias in these kinds of games so we actually in these kinds of games so we actually in these kinds of games so we actually played a tournament in a non-language played a tournament in a non-language played a tournament in a non-language version of the game where you know we we version of the game where you know we we version of the game where you know we we told the participants like hey in every told the participants like hey in every told the participants like hey in every single game there's going to be an AI
-
single game there's going to be an AI single game there's going to be an AI and what we found is that the humans and what we found is that the humans and what we found is that the humans would spend basically the entire game would spend basically the entire game would spend basically the entire game like trying to figure out who the bot like trying to figure out who the bot like trying to figure out who the bot was and then as soon as they thought was and then as soon as they thought was and then as soon as they thought they figured it out they would all team they figured it out they would all team they figured it out they would all team up and try to kill it up and try to kill it up and try to kill it um and you know overcoming that inherent um and you know overcoming that inherent um and you know overcoming that inherent anti-ai bias is is a challenge um on the anti-ai bias is is a challenge um on the anti-ai bias is is a challenge um on the flip side flip side flip side I think when robots become the enemy I think when robots become the enemy I think when robots become the enemy that's when we get to heal our human that's when we get to heal our human that's when we get to heal our human divisions and then we can become one as divisions and then we can become one as divisions and then we can become one as long as we have one enemy it's it's that long as we have one enemy it's it's that long as we have one enemy it's it's that Reagan thing when aliens show up that's Reagan thing when aliens show up that's Reagan thing when aliens show up that's when we we put our side our divisions when we we put our side our divisions when we we put our side our divisions we've become one one human species right we've become one one human species right we've become one one human species right you might have our differences but we're you might have our differences but we're you might have our differences but we're at least all human at least we all hate at least all human at least we all hate at least all human at least we all hate the robots no no no I think there will the robots no no no I think there will the robots no no no I think there will be actually in the future something like be actually in the future something like be actually in the future something like a civil rights movement for robots I a civil rights movement for robots I a civil rights movement for robots I think that's the fascinating things think that's the fascinating things think that's the fascinating things about AI systems and is they ask they about AI systems and is they ask they about AI systems and is they ask they force us to ask about force us to ask about force us to ask about ethical questions about what is ethical questions about what is ethical questions about what is sentience what is uh how do we feel sentience what is uh how do we feel sentience what is uh how do we feel about systems that are capable of about systems that are capable of about systems that are capable of suffering or capable of displaying suffering or capable of displaying suffering or capable of displaying suffering and how do we design products suffering and how do we design products suffering and how do we design products that show emotion and not how do we feel that show emotion and not how do we feel that show emotion and not how do we feel about that lying is another topic are we about that lying is another topic are we about that lying is another topic are we going to allow Bots to lie and not and going to allow Bots to lie and not and going to allow Bots to lie and not and where's the balance between being nice where's the balance between being nice where's the balance between being nice and and telling the truth I mean these and and telling the truth I mean these and and telling the truth I mean these are all fascinating human questions it's are all fascinating human questions it's are all fascinating human questions it's like so exciting to be in the century like so exciting to be in the century like so exciting to be in the century when we create systems that when we create systems that when we create systems that take these philosophical questions that take these philosophical questions that take these philosophical questions that have been asked for centuries and now we have been asked for centuries and now we have been asked for centuries and now we can engineer them inside systems where
-
can engineer them inside systems where can engineer them inside systems where like you really have to answer them like you really have to answer them like you really have to answer them because you'll have because you'll have because you'll have transformational impact on human society transformational impact on human society transformational impact on human society depending on what you design inside depending on what you design inside depending on what you design inside those systems it's fascinating and like those systems it's fascinating and like those systems it's fascinating and like you said I feel like diplomacy is a step you said I feel like diplomacy is a step you said I feel like diplomacy is a step towards the direction of the real world towards the direction of the real world towards the direction of the real world applying these RL methods towards the applying these RL methods towards the applying these RL methods towards the real world real world real world from from all the Breakthrough from from all the Breakthrough from from all the Breakthrough performances and go and chess and performances and go and chess and performances and go and chess and Starcraft and DOTA this is this feels Starcraft and DOTA this is this feels Starcraft and DOTA this is this feels like the real world especially now my like the real world especially now my like the real world especially now my mind's been on war in military conflict mind's been on war in military conflict mind's been on war in military conflict this feels like it can give us some deep this feels like it can give us some deep this feels like it can give us some deep insights about human behavior at the insights about human behavior at the insights about human behavior at the large geopolitical scale large geopolitical scale large geopolitical scale um um um what do you think what do you think what do you think is the um breakthrough is the um breakthrough is the um breakthrough or or or the directions of work that will take us the directions of work that will take us the directions of work that will take us towards solving intelligence towards towards solving intelligence towards towards solving intelligence towards creating AGI systems you've been a part creating AGI systems you've been a part creating AGI systems you've been a part of creating of creating of creating um by the way we should say a part of um by the way we should say a part of um by the way we should say a part of great teams that do this of creating great teams that do this of creating great teams that do this of creating systems that achieve breakthrough systems that achieve breakthrough systems that achieve breakthrough performances on before thought performances on before thought performances on before thought unsolvable problems like poker uh unsolvable problems like poker uh unsolvable problems like poker uh multiplayer poker diplomacy multiplayer poker diplomacy multiplayer poker diplomacy we're taking steps towards that we're taking steps towards that we're taking steps towards that direction what do you think it takes to direction what do you think it takes to direction what do you think it takes to go all the way to create superhuman go all the way to create superhuman go all the way to create superhuman level intelligence level intelligence level intelligence you know there's a lot of people trying you know there's a lot of people trying you know there's a lot of people trying to figure that out right now to figure that out right now to figure that out right now um and you know I should say like the um and you know I should say like the um and you know I should say like the amount of progress that's been made amount of progress that's been made amount of progress that's been made especially in the past few years is especially in the past few years is especially in the past few years is truly phenomenal I mean you look at truly phenomenal I mean you look at truly phenomenal I mean you look at where AI was 10 years ago and the idea
-
where AI was 10 years ago and the idea where AI was 10 years ago and the idea that you could have AIS that can that you could have AIS that can that you could have AIS that can generate language and generate images generate language and generate images generate language and generate images the way they're doing today and able to the way they're doing today and able to the way they're doing today and able to play a game like diplomacy was just like play a game like diplomacy was just like play a game like diplomacy was just like Unthinkable Unthinkable Unthinkable um even even five years ago let alone um even even five years ago let alone um even even five years ago let alone ten years ago ten years ago ten years ago um um um now there are there are aspects of AI now there are there are aspects of AI now there are there are aspects of AI that I think are still lacking um that I think are still lacking um that I think are still lacking um I think there's General agreements that I think there's General agreements that I think there's General agreements that one of the major issues with AI today is one of the major issues with AI today is one of the major issues with AI today is that it's very data inefficient it's that it's very data inefficient it's that it's very data inefficient it's very it requires a huge number of very it requires a huge number of very it requires a huge number of samples of training examples to be able samples of training examples to be able samples of training examples to be able to train you know you look at an AI that to train you know you look at an AI that to train you know you look at an AI that plays go and it needs millions of games plays go and it needs millions of games plays go and it needs millions of games of go to uh to learn how to play the of go to uh to learn how to play the of go to uh to learn how to play the game well whereas a human can pick it up game well whereas a human can pick it up game well whereas a human can pick it up in like you know I don't know how many in like you know I don't know how many in like you know I don't know how many games does a human go player go go Grand games does a human go player go go Grand games does a human go player go go Grand Master play in their lifetime probably Master play in their lifetime probably Master play in their lifetime probably you know in the thousands or tens of you know in the thousands or tens of you know in the thousands or tens of thousands I guess thousands I guess thousands I guess um um um so that's that's one issue overcome so that's that's one issue overcome so that's that's one issue overcome efficiency overcoming this challenge of efficiency overcoming this challenge of efficiency overcoming this challenge of data efficiency and this is particularly data efficiency and this is particularly data efficiency and this is particularly important if we want to deploy AI important if we want to deploy AI important if we want to deploy AI systems in real world settings systems in real world settings systems in real world settings um where they're interacting with humans um where they're interacting with humans um where they're interacting with humans because you know for example with because you know for example with because you know for example with robotics it's really hard to generate a robotics it's really hard to generate a robotics it's really hard to generate a huge number of samples it's it's a huge number of samples it's it's a huge number of samples it's it's a different story when you're working in different story when you're working in different story when you're working in these you know totally virtual games these you know totally virtual games these you know totally virtual games where you can play a million games and where you can play a million games and where you can play a million games and it's no big deal I was planning on just it's no big deal I was planning on just it's no big deal I was planning on just launching like a thousand of these launching like a thousand of these launching like a thousand of these robots in Austin I don't think it's robots in Austin I don't think it's robots in Austin I don't think it's illegal for Lego robots to roam the illegal for Lego robots to roam the illegal for Lego robots to roam the streets and just collect data that's not streets and just collect data that's not streets and just collect data that's not of course the worst that could happen of course the worst that could happen of course the worst that could happen yeah I kind of I mean that's one way to yeah I kind of I mean that's one way to yeah I kind of I mean that's one way to overcome the data efficiency problem is overcome the data efficiency problem is overcome the data efficiency problem is like scale it yeah like I actually tried
-
like scale it yeah like I actually tried like scale it yeah like I actually tried to see if there's a law against robots to see if there's a law against robots to see if there's a law against robots like Lego robots just operating like Lego robots just operating like Lego robots just operating in in this in the streets of a major in in this in the streets of a major in in this in the streets of a major city and there isn't I couldn't find any city and there isn't I couldn't find any city and there isn't I couldn't find any so so so um I'll take it all the way to the um I'll take it all the way to the um I'll take it all the way to the Supreme Court Supreme Court Supreme Court robot rights okay anyway sorry you were robot rights okay anyway sorry you were robot rights okay anyway sorry you were saying so the so what what are the ideas saying so the so what what are the ideas saying so the so what what are the ideas for getting becoming more data efficient for getting becoming more data efficient for getting becoming more data efficient uh I mean that's that's the trillion uh I mean that's that's the trillion uh I mean that's that's the trillion dollar question in AI today I mean if dollar question in AI today I mean if dollar question in AI today I mean if you can figure out how to make AI you can figure out how to make AI you can figure out how to make AI systems more more data efficient then systems more more data efficient then systems more more data efficient then that's that's that's a huge breakthrough so nobody really a huge breakthrough so nobody really a huge breakthrough so nobody really knows right now it could be just a knows right now it could be just a knows right now it could be just a gigantic background model language model gigantic background model language model gigantic background model language model and then you do and then you do and then you do um the training becomes like prompting um the training becomes like prompting um the training becomes like prompting that model that model that model to uh to uh to uh to essentially do a kind of querying a to essentially do a kind of querying a to essentially do a kind of querying a search into the space of the things it's search into the space of the things it's search into the space of the things it's learned to customize that to whatever learned to customize that to whatever learned to customize that to whatever problem you're trying to solve so maybe problem you're trying to solve so maybe problem you're trying to solve so maybe if you form a large enough language if you form a large enough language if you form a large enough language model you can go quite quite a long way model you can go quite quite a long way model you can go quite quite a long way that you know I think there's some truth that you know I think there's some truth that you know I think there's some truth to that I mean you look at the way to that I mean you look at the way to that I mean you look at the way humans approach humans approach humans approach um a game like poker they're not coming um a game like poker they're not coming um a game like poker they're not coming at it from scratch they're coming at it at it from scratch they're coming at it at it from scratch they're coming at it with a huge amount of background with a huge amount of background with a huge amount of background knowledge about you know how humans work knowledge about you know how humans work knowledge about you know how humans work how the world Works how the world Works how the world Works um the idea of money so they're able to um the idea of money so they're able to um the idea of money so they're able to leverage that kind of information to uh leverage that kind of information to uh leverage that kind of information to uh to to pick up the game faster uh so it's to to pick up the game faster uh so it's to to pick up the game faster uh so it's not really a fair comparison to then not really a fair comparison to then not really a fair comparison to then compare it to an AI That's like learning compare it to an AI That's like learning compare it to an AI That's like learning from scratch and maybe one of the ways from scratch and maybe one of the ways from scratch and maybe one of the ways that we address this uh sample that we address this uh sample that we address this uh sample complexity problem is by allowing AIS to complexity problem is by allowing AIS to complexity problem is by allowing AIS to leverage that general knowledge across a
-
leverage that general knowledge across a leverage that general knowledge across a ton of different domains ton of different domains ton of different domains so like I said you did uh a lot of so like I said you did uh a lot of so like I said you did uh a lot of incredible work in the space of research incredible work in the space of research incredible work in the space of research and actually Building Systems what and actually Building Systems what and actually Building Systems what advice would you give to uh let's start advice would you give to uh let's start advice would you give to uh let's start with beginners what advice would you with beginners what advice would you with beginners what advice would you give to beginners interested in machine give to beginners interested in machine give to beginners interested in machine learning just there at the very start of learning just there at the very start of learning just there at the very start of their Journey they're in high school and their Journey they're in high school and their Journey they're in high school and college thinking like this seems like a college thinking like this seems like a college thinking like this seems like a fascinating world what advice would you fascinating world what advice would you fascinating world what advice would you give them give them give them um I I would say um I I would say um I I would say that there are a lot of people working that there are a lot of people working that there are a lot of people working on similar aspects of machine learning on similar aspects of machine learning on similar aspects of machine learning and to not be afraid to try something a and to not be afraid to try something a and to not be afraid to try something a bit different my own path bit different my own path bit different my own path in AI is pretty atypical for a machine in AI is pretty atypical for a machine in AI is pretty atypical for a machine learning researcher today I mean I learning researcher today I mean I learning researcher today I mean I started out working on Game Theory and started out working on Game Theory and started out working on Game Theory and um and then shifting more towards um and then shifting more towards um and then shifting more towards reinforcement learning as time went on reinforcement learning as time went on reinforcement learning as time went on and that actually had a lot of benefits and that actually had a lot of benefits and that actually had a lot of benefits I think because it allowed me to look at I think because it allowed me to look at I think because it allowed me to look at these problems in a very different way these problems in a very different way these problems in a very different way from the way a lot of machine learning from the way a lot of machine learning from the way a lot of machine learning researchers view it and that comes with researchers view it and that comes with researchers view it and that comes with um drawbacks in some respects like I um drawbacks in some respects like I um drawbacks in some respects like I think there's definitely aspects of think there's definitely aspects of think there's definitely aspects of machine learning where you know I'm machine learning where you know I'm machine learning where you know I'm weaker than most of the researchers out weaker than most of the researchers out weaker than most of the researchers out there but I think that diversity of there but I think that diversity of there but I think that diversity of perspective perspective perspective um you know when I'm working with my um you know when I'm working with my um you know when I'm working with my teammates teammates teammates um there's something that I'm bringing um there's something that I'm bringing um there's something that I'm bringing to the table and there's something that to the table and there's something that to the table and there's something that they're bringing to the table and that they're bringing to the table and that they're bringing to the table and that kind of collaboration becomes very kind of collaboration becomes very kind of collaboration becomes very fruitful for that reason so there could fruitful for that reason so there could fruitful for that reason so there could be problems like like poker like you've be problems like like poker like you've be problems like like poker like you've chosen diplomacy that could be problems chosen diplomacy that could be problems chosen diplomacy that could be problems like that still out there that you can like that still out there that you can like that still out there that you can just tackle even if it seems extremely
-
just tackle even if it seems extremely just tackle even if it seems extremely difficult um I think that there's a lot of um I think that there's a lot of challenges challenges left and I think challenges challenges left and I think challenges challenges left and I think having a diversity of viewpoints and having a diversity of viewpoints and having a diversity of viewpoints and backgrounds is really helpful for backgrounds is really helpful for backgrounds is really helpful for working together to figure out how to working together to figure out how to working together to figure out how to tackle those kinds of challenges so as a tackle those kinds of challenges so as a tackle those kinds of challenges so as a beginner so that I would say that's beginner so that I would say that's beginner so that I would say that's that's more for like a grad student they that's more for like a grad student they that's more for like a grad student they already built up a base like a complete already built up a base like a complete already built up a base like a complete beginner what's a good journey so for beginner what's a good journey so for beginner what's a good journey so for you that was doing some more on the math you that was doing some more on the math you that was doing some more on the math side of things doing Game Theory all side of things doing Game Theory all side of things doing Game Theory all that good so it's basically build up a that good so it's basically build up a that good so it's basically build up a foundation in something so programming foundation in something so programming foundation in something so programming mathematics it could even be physics but mathematics it could even be physics but mathematics it could even be physics but build build that Foundation build build that Foundation build build that Foundation yeah I would say build a strong yeah I would say build a strong yeah I would say build a strong foundation in math and computer science foundation in math and computer science foundation in math and computer science and statistics in these kinds of areas and statistics in these kinds of areas and statistics in these kinds of areas but but don't be afraid to try something but but don't be afraid to try something but but don't be afraid to try something that's different and learn something that's different and learn something that's different and learn something that's different from you know the the that's different from you know the the that's different from you know the the thing that everybody else is doing to thing that everybody else is doing to thing that everybody else is doing to get into machine learning um you know get into machine learning um you know get into machine learning um you know there's there's value in having a there's there's value in having a there's there's value in having a different background than everybody else different background than everybody else different background than everybody else um yeah so but certainly having a strong um yeah so but certainly having a strong um yeah so but certainly having a strong math background especially in things math background especially in things math background especially in things like linear algebra and statistics and like linear algebra and statistics and like linear algebra and statistics and probability probability probability um are incredibly helpful today for for um are incredibly helpful today for for um are incredibly helpful today for for learning about and understanding machine learning about and understanding machine learning about and understanding machine learning do you think one day we'll be learning do you think one day we'll be learning do you think one day we'll be able to since you're taking steps from able to since you're taking steps from able to since you're taking steps from poker to diplomacy one day we'll be able poker to diplomacy one day we'll be able poker to diplomacy one day we'll be able to uh to uh to uh figure out how to live life optimally figure out how to live life optimally figure out how to live life optimally well what is it like in in poker and well what is it like in in poker and well what is it like in in poker and diplomacy you need a value function you diplomacy you need a value function you diplomacy you need a value function you need to have a reward system and so what need to have a reward system and so what need to have a reward system and so what does it mean to live a life that's does it mean to live a life that's does it mean to live a life that's optimal so okay so then you can exactly
-
optimal so okay so then you can exactly optimal so okay so then you can exactly like lay down a reward function being like lay down a reward function being like lay down a reward function being like I want to be rich or I want to be like I want to be rich or I want to be like I want to be rich or I want to be um um um I want to be in a happy relationship and I want to be in a happy relationship and I want to be in a happy relationship and then you'll say well then you'll say well then you'll say well do X do X do X you know there's there's a lot of uh you know there's there's a lot of uh you know there's there's a lot of uh talk today about in in AI safety circles talk today about in in AI safety circles talk today about in in AI safety circles about like this specification of you about like this specification of you about like this specification of you know reward functions so you you say know reward functions so you you say know reward functions so you you say like okay my objective is to be rich and like okay my objective is to be rich and like okay my objective is to be rich and maybe the AI tells you like okay well if maybe the AI tells you like okay well if maybe the AI tells you like okay well if you want to maximize the probability you want to maximize the probability you want to maximize the probability that you're rich go rob a bank sure and that you're rich go rob a bank sure and that you're rich go rob a bank sure and so you wanna is that is that really what so you wanna is that is that really what so you wanna is that is that really what you want is your objective really to be you want is your objective really to be you want is your objective really to be rich at all costs or is it more nuanced rich at all costs or is it more nuanced rich at all costs or is it more nuanced than that so the understand the than that so the understand the than that so the understand the consequences yeah consequences yeah consequences yeah yeah so yeah that that's so maybe life yeah so yeah that that's so maybe life yeah so yeah that that's so maybe life is more about defining the reward is more about defining the reward is more about defining the reward function that minimizes the unintended function that minimizes the unintended function that minimizes the unintended consequences consequences consequences than it is about the actual policy that than it is about the actual policy that than it is about the actual policy that gets you to the rewards function maybe gets you to the rewards function maybe gets you to the rewards function maybe life is just about constantly updating life is just about constantly updating life is just about constantly updating the reward function the reward function the reward function I think one of the challenges in life is I think one of the challenges in life is I think one of the challenges in life is is figuring out exactly what that reward is figuring out exactly what that reward is figuring out exactly what that reward function is sometimes it's pretty hard function is sometimes it's pretty hard function is sometimes it's pretty hard to specify the same way that you know to specify the same way that you know to specify the same way that you know trying to handcraft the optimal policy trying to handcraft the optimal policy trying to handcraft the optimal policy in a game like chess is really difficult in a game like chess is really difficult in a game like chess is really difficult it's not so clear-cut what the reward it's not so clear-cut what the reward it's not so clear-cut what the reward function is for for life function is for for life function is for for life I think one day AI will figure it out and I wonder what that would be until and I wonder what that would be until then then then I just really appreciate the kind of I just really appreciate the kind of I just really appreciate the kind of work you're doing and um it's it's work you're doing and um it's it's work you're doing and um it's it's really fascinating taking a leap into a
-
really fascinating taking a leap into a really fascinating taking a leap into a more and more real world like more and more real world like more and more real world like um problem space and just achieving um problem space and just achieving um problem space and just achieving incredible results by applying incredible results by applying incredible results by applying reinforcement learning no since I saw reinforcement learning no since I saw reinforcement learning no since I saw you work on poker you've been in you work on poker you've been in you work on poker you've been in constant inspiration it's an honor to constant inspiration it's an honor to constant inspiration it's an honor to get to finally talk to you and uh this get to finally talk to you and uh this get to finally talk to you and uh this is really fun thanks for having me is really fun thanks for having me is really fun thanks for having me thanks for listening to this thanks for listening to this thanks for listening to this conversation with no Brown conversation with no Brown conversation with no Brown to support this podcast please check out to support this podcast please check out to support this podcast please check out our sponsors in the description and now our sponsors in the description and now our sponsors in the description and now let me leave you with some words from let me leave you with some words from let me leave you with some words from Sun Tzu and the Art of War Sun Tzu and the Art of War Sun Tzu and the Art of War the whole secret lies in confusing the the whole secret lies in confusing the the whole secret lies in confusing the enemy so that he cannot fathom our real enemy so that he cannot fathom our real enemy so that he cannot fathom our real intent intent intent thank you for listening and hope to see thank you for listening and hope to see thank you for listening and hope to see you next time
Summary
The main theme is the advancement of AI in complex games, referencing Libratus, Pluribus, and Cicero's achievements in poker and diplomacy. These AI systems demonstrate that strategic success, even in human-like negotiations, can be achieved through approximating game theory principles like the Nash equilibrium rather than solely through adversarial "mind games." The takeaway is that AI can surpass human performance in games requiring deep strategy and negotiation by adhering to optimal game-theoretic approaches.