← Back
Theo September 13, 2026 34m

I think they mean it this time

Read full transcript 26 segments
  1. There's been a lot of talk about AI There's been a lot of talk about AI safety stuff lately. From the chaos that safety stuff lately. From the chaos that safety stuff lately. From the chaos that is the Chinese distillation of models to is the Chinese distillation of models to is the Chinese distillation of models to the employee Jacob leaving Anthropic due the employee Jacob leaving Anthropic due the employee Jacob leaving Anthropic due to cited concerns around safety and how to cited concerns around safety and how to cited concerns around safety and how he didn't believe even Anthropic was he didn't believe even Anthropic was he didn't believe even Anthropic was taking it seriously enough. Now things taking it seriously enough. Now things taking it seriously enough. Now things have gone even further. It seems like have gone even further. It seems like have gone even further. It seems like the pacing the frontier movement has the pacing the frontier movement has the pacing the frontier movement has really taken hold inside of the labs and really taken hold inside of the labs and really taken hold inside of the labs and now we see Daario coming back on Twitter now we see Daario coming back on Twitter now we see Daario coming back on Twitter for the first time in a bit. I think for the first time in a bit. I think for the first time in a bit. I think it's one of his four ever tweets and it's one of his four ever tweets and it's one of his four ever tweets and he's posting in order to share a new he's posting in order to share a new he's posting in order to share a new article he wrote. We must pace the article he wrote. We must pace the article he wrote. We must pace the frontier. By itself, this would be very frontier. By itself, this would be very frontier. By itself, this would be very notable and worth at least paying some notable and worth at least paying some notable and worth at least paying some attention to. But what makes it much attention to. But what makes it much attention to. But what makes it much more interesting is the fact that more interesting is the fact that more interesting is the fact that everybody from Elon Musk to Sam Alman everybody from Elon Musk to Sam Alman everybody from Elon Musk to Sam Alman has cited this and said they agree and has cited this and said they agree and has cited this and said they agree and it's probably time to do something. I'm it's probably time to do something. I'm it's probably time to do something. I'm particularly interested in this article particularly interested in this article particularly interested in this article from Daario because it no longer dances from Daario because it no longer dances from Daario because it no longer dances around the big scary questions. Previous around the big scary questions. Previous around the big scary questions. Previous attempts to cover this haven't been attempts to cover this haven't been attempts to cover this haven't been realistic about where things are at realistic about where things are at realistic about where things are at geopolitically, in particular with geopolitically, in particular with geopolitically, in particular with China, as well as what these risks China, as well as what these risks China, as well as what these risks actually look like in the real world. I actually look like in the real world. I actually look like in the real world. I know you'll tend to think Anthropic is know you'll tend to think Anthropic is know you'll tend to think Anthropic is blowing all of this out of proportion blowing all of this out of proportion blowing all of this out of proportion and exaggerating the issue here, but I and exaggerating the issue here, but I and exaggerating the issue here, but I think Daario was quite level-headed in think Daario was quite level-headed in think Daario was quite level-headed in his reporting here. I think it's his reporting here. I think it's his reporting here. I think it's important that we go through this important that we go through this important that we go through this together and try to understand what he's together and try to understand what he's together and try to understand what he's saying. That said, if he is successful, saying. That said, if he is successful, saying. That said, if he is successful, there will be a lot less AI news going there will be a lot less AI news going there will be a lot less AI news going on, which would mean I don't have a good on, which would mean I don't have a good on, which would mean I don't have a good place to put my awesome sponsors like place to put my awesome sponsors like place to put my awesome sponsors like today's. Usually, when a company today's. Usually, when a company today's. Usually, when a company sponsors my videos, it's cuz they want sponsors my videos, it's cuz they want sponsors my videos, it's cuz they want to make more money, which is why it's so to make more money, which is why it's so to make more money, which is why it's so weird today's sponsor wants me to tell weird today's sponsor wants me to tell weird today's sponsor wants me to tell you about how to use them less. I'm you about how to use them less. I'm you about how to use them less. I'm thankful this isn't a joke because I thankful this isn't a joke because I thankful this isn't a joke because I saved a bunch of money too. Today's saved a bunch of money too. Today's saved a bunch of money too. Today's sponsor is Blacksmith. I've already sponsor is Blacksmith. I've already sponsor is Blacksmith. I've already established that they're the best place established that they're the best place established that they're the best place to run your GitHub action CI and more.

  2. to run your GitHub action CI and more. to run your GitHub action CI and more. But now they also have Codesmith which But now they also have Codesmith which But now they also have Codesmith which lets you run agents on top of that same lets you run agents on top of that same lets you run agents on top of that same super fast, super cheap, and reliable super fast, super cheap, and reliable super fast, super cheap, and reliable infrastructure. What's even cooler is infrastructure. What's even cooler is infrastructure. What's even cooler is that these agents come with deep that these agents come with deep that these agents come with deep knowledge of how your CI works as well knowledge of how your CI works as well knowledge of how your CI works as well as how Blacksmith works. This means you as how Blacksmith works. This means you as how Blacksmith works. This means you can ask them to do things like rightsize can ask them to do things like rightsize can ask them to do things like rightsize your CI runners and it will look through your CI runners and it will look through your CI runners and it will look through your actual logs and see which jobs used your actual logs and see which jobs used your actual logs and see which jobs used a bunch of CPU and which ones didn't and a bunch of CPU and which ones didn't and a bunch of CPU and which ones didn't and recommend how to spec things out so you recommend how to spec things out so you recommend how to spec things out so you save more money and more time. I ended save more money and more time. I ended save more money and more time. I ended up merging three real PRs in T3 Code, up merging three real PRs in T3 Code, up merging three real PRs in T3 Code, all of which were filed by Codesmith all of which were filed by Codesmith all of which were filed by Codesmith when I was filming a quick demo. It took when I was filming a quick demo. It took when I was filming a quick demo. It took less than 5 minutes to shave my already less than 5 minutes to shave my already less than 5 minutes to shave my already super fast CI times by almost half just super fast CI times by almost half just super fast CI times by almost half just by doing the things it said. And that's by doing the things it said. And that's by doing the things it said. And that's on top of the 50% plus improvement in on top of the 50% plus improvement in on top of the 50% plus improvement in performance you'll get just for moving performance you'll get just for moving performance you'll get just for moving to Blacksmith in the first place. You to Blacksmith in the first place. You to Blacksmith in the first place. You change one line of code in your existing change one line of code in your existing change one line of code in your existing GitHub action and you're ready to go. If GitHub action and you're ready to go. If GitHub action and you're ready to go. If they charge four times more than GitHub they charge four times more than GitHub they charge four times more than GitHub actions, I would still think it's worth actions, I would still think it's worth actions, I would still think it's worth it. But they actually charge way less. it. But they actually charge way less. it. But they actually charge way less. When you combine the price difference When you combine the price difference When you combine the price difference and the speed difference, you end up and the speed difference, you end up and the speed difference, you end up saving 60% or more against what would saving 60% or more against what would saving 60% or more against what would have been your GitHub action build.

  3. have been your GitHub action build. have been your GitHub action build. Blacksmith is so fast that I use CI for Blacksmith is so fast that I use CI for Blacksmith is so fast that I use CI for things I never would have before. And things I never would have before. And things I never would have before. And it's letting our team ship faster and it's letting our team ship faster and it's letting our team ship faster and more confidently. Figure out why my more confidently. Figure out why my more confidently. Figure out why my whole team loves these guys at whole team loves these guys at whole team loves these guys at soyb.link/blacksmith. soyb.link/blacksmith. soyb.link/blacksmith. Sorry about that one. Got medical bills Sorry about that one. Got medical bills Sorry about that one. Got medical bills today. I'm sure you fellow Americans today. I'm sure you fellow Americans today. I'm sure you fellow Americans understand. I am excited to cover what understand. I am excited to cover what understand. I am excited to cover what Dario has said here, as well as how Dario has said here, as well as how Dario has said here, as well as how others have responded. But I want to others have responded. But I want to others have responded. But I want to jump on one other thing first. A pattern jump on one other thing first. A pattern jump on one other thing first. A pattern I've been noticing in these I've been noticing in these I've been noticing in these conversations, in particular, the conversations, in particular, the conversations, in particular, the conversations about Jacob when he left conversations about Jacob when he left conversations about Jacob when he left Anthropic. I've been seeing a ton of Anthropic. I've been seeing a ton of Anthropic. I've been seeing a ton of crazy conspiracy theories, many of which crazy conspiracy theories, many of which crazy conspiracy theories, many of which weren't out yet by the time I published weren't out yet by the time I published weren't out yet by the time I published my video. And now that I've seen them, my video. And now that I've seen them, my video. And now that I've seen them, people seem to think that I'm people seem to think that I'm people seem to think that I'm intentionally dodging this political intentionally dodging this political intentionally dodging this political scop. No, I'm not. You guys are just scop. No, I'm not. You guys are just scop. No, I'm not. You guys are just insane. Seriously, like almost all of insane. Seriously, like almost all of insane. Seriously, like almost all of the things I've seen people talking the things I've seen people talking the things I've seen people talking about in regards to Jacob's post have about in regards to Jacob's post have about in regards to Jacob's post have been crazy attempts to map like a grant been crazy attempts to map like a grant been crazy attempts to map like a grant he received in research in 2021 to him he received in research in 2021 to him he received in research in 2021 to him quitting a hundred million plus dollar quitting a hundred million plus dollar quitting a hundred million plus dollar job in order to what? Like just there job in order to what? Like just there job in order to what? Like just there there is no link between any of those there is no link between any of those there is no link between any of those things. and all of the attempts to like things. and all of the attempts to like things. and all of the attempts to like assign what he said and who he's talking assign what he said and who he's talking assign what he said and who he's talking to to some weird cabal trying to do to to some weird cabal trying to do to to some weird cabal trying to do something, but nobody can say what the something, but nobody can say what the something, but nobody can say what the something is or what their goals something is or what their goals something is or what their goals actually are. And I'm particularly actually are. And I'm particularly actually are. And I'm particularly frustrated, not because conspiracy frustrated, not because conspiracy frustrated, not because conspiracy theories annoy me. I actually find him theories annoy me. I actually find him theories annoy me. I actually find him quite fun. I'm annoyed because I feel quite fun. I'm annoyed because I feel quite fun. I'm annoyed because I feel like we're not talking about what Jacob like we're not talking about what Jacob like we're not talking about what Jacob said. Instead, we're talking about who said. Instead, we're talking about who said. Instead, we're talking about who he is and if he's worth even listening he is and if he's worth even listening he is and if he's worth even listening to. And the problem is that the two to. And the problem is that the two to. And the problem is that the two sides aren't sensical. One side thinks sides aren't sensical. One side thinks sides aren't sensical. One side thinks what Jacob said is worth listening to what Jacob said is worth listening to what Jacob said is worth listening to and considering. the other side thinks and considering. the other side thinks and considering. the other side thinks you're insane if you listen to a word you're insane if you listen to a word you're insane if you listen to a word that he says and they're not even

  4. that he says and they're not even that he says and they're not even engaging with the things he said. We're engaging with the things he said. We're engaging with the things he said. We're not debating whether or not AI is safe. not debating whether or not AI is safe. not debating whether or not AI is safe. We are debating whether or not the We are debating whether or not the We are debating whether or not the conversation can be had. And one side conversation can be had. And one side conversation can be had. And one side thinks yes because AI could potentially thinks yes because AI could potentially thinks yes because AI could potentially be really unsafe. And the other side be really unsafe. And the other side be really unsafe. And the other side says we cannot have the conversation at says we cannot have the conversation at says we cannot have the conversation at all. And if you're trying to have it, all. And if you're trying to have it, all. And if you're trying to have it, you're probably funded by some weird you're probably funded by some weird you're probably funded by some weird party trying to force their way in the party trying to force their way in the party trying to force their way in the world. No, I'm not funded by anybody. I world. No, I'm not funded by anybody. I world. No, I'm not funded by anybody. I love AI. If AI slows down, it will hurt love AI. If AI slows down, it will hurt love AI. If AI slows down, it will hurt me directly. I'm just doing my best to me directly. I'm just doing my best to me directly. I'm just doing my best to cover this because the people who I know cover this because the people who I know cover this because the people who I know who are the smartest in the world at who are the smartest in the world at who are the smartest in the world at this are legitimately scared and this are legitimately scared and this are legitimately scared and many of them are friends of Jacobs can many of them are friends of Jacobs can many of them are friends of Jacobs can absolutely vet his capabilities and absolutely vet his capabilities and absolutely vet his capabilities and agree with what he is saying. Daario agree with what he is saying. Daario agree with what he is saying. Daario seems to be one of those people which is seems to be one of those people which is seems to be one of those people which is why I think it's worth listening to what why I think it's worth listening to what why I think it's worth listening to what he said here. He shared his blog with he said here. He shared his blog with he said here. He shared his blog with the following Twitter post as a starting the following Twitter post as a starting the following Twitter post as a starting point. We must pace the frontier. I've point. We must pace the frontier. I've point. We must pace the frontier. I've written a new essay on why the AI written a new essay on why the AI written a new essay on why the AI industry should slow down with a industry should slow down with a industry should slow down with a threepart plan for doing so. Anthropic threepart plan for doing so. Anthropic threepart plan for doing so. Anthropic is unilaterally committing to the first is unilaterally committing to the first is unilaterally committing to the first of these steps. We'll provide thirdparty of these steps. We'll provide thirdparty of these steps. We'll provide thirdparty evaluators with permanent employee level evaluators with permanent employee level evaluators with permanent employee level access to our systems so that they can access to our systems so that they can access to our systems so that they can verify adherence to our safety measures, verify adherence to our safety measures, verify adherence to our safety measures, report on incidents, and assess models report on incidents, and assess models report on incidents, and assess models alignment during training. Dario has alignment during training. Dario has alignment during training. Dario has worked on AI for the last 12 years worked on AI for the last 12 years worked on AI for the last 12 years because he believes it could because he believes it could because he believes it could dramatically raise the quality of human dramatically raise the quality of human dramatically raise the quality of human life. He's written often about these life. He's written often about these life. He's written often about these incredible benefits. He believes AI incredible benefits. He believes AI incredible benefits. He believes AI could cure most major diseases in the could cure most major diseases in the could cure most major diseases in the next 5 to 10 years, greatly accelerate next 5 to 10 years, greatly accelerate next 5 to 10 years, greatly accelerate economic growth rates, create a world of economic growth rates, create a world of economic growth rates, create a world of abundance and empowerment, and usher in abundance and empowerment, and usher in abundance and empowerment, and usher in a renaissance of democracy and freedom.

  5. a renaissance of democracy and freedom. a renaissance of democracy and freedom. He feels the urgency personally. His own He feels the urgency personally. His own He feels the urgency personally. His own father died of a disease that was cured father died of a disease that was cured father died of a disease that was cured just a few years after his death, and he just a few years after his death, and he just a few years after his death, and he himself survived early stage cancer that himself survived early stage cancer that himself survived early stage cancer that would not have been treatable even 50 would not have been treatable even 50 would not have been treatable even 50 years ago. Carefully wielded, AI can be years ago. Carefully wielded, AI can be years ago. Carefully wielded, AI can be the latest in a long line of the latest in a long line of the latest in a long line of technological miracles that have technological miracles that have technological miracles that have uplifted and embolded humanity. Like uplifted and embolded humanity. Like uplifted and embolded humanity. Like many technologies before it, AI brings many technologies before it, AI brings many technologies before it, AI brings risks. And because it is such a powerful risks. And because it is such a powerful risks. And because it is such a powerful technology, these risks are serious. He technology, these risks are serious. He technology, these risks are serious. He talks here about things like losing talks here about things like losing talks here about things like losing control of AI systems, misuse for cyber control of AI systems, misuse for cyber control of AI systems, misuse for cyber attacks and bioteterrorism, economic attacks and bioteterrorism, economic attacks and bioteterrorism, economic disruption, all the above. And that if disruption, all the above. And that if disruption, all the above. And that if we are driven too much by commercial we are driven too much by commercial we are driven too much by commercial incentive, these risks can become more incentive, these risks can become more incentive, these risks can become more acute. He's grappled with this duality acute. He's grappled with this duality acute. He's grappled with this duality of risk and benefit since the beginning of risk and benefit since the beginning of risk and benefit since the beginning of anthropic. Not building it deprivives of anthropic. Not building it deprivives of anthropic. Not building it deprivives humanity of the benefits or simply humanity of the benefits or simply humanity of the benefits or simply places AI in the hands of authoritarian places AI in the hands of authoritarian places AI in the hands of authoritarian powers while building it too fast is powers while building it too fast is powers while building it too fast is reckless. If sought a middle way to show reckless. If sought a middle way to show reckless. If sought a middle way to show that it's possible to build carefully that it's possible to build carefully that it's possible to build carefully and succeed commercially and to make and succeed commercially and to make and succeed commercially and to make safety something on which AI companies safety something on which AI companies safety something on which AI companies compete. In other words, create a race compete. In other words, create a race compete. In other words, create a race to the top. Anthropobics always devoted to the top. Anthropobics always devoted to the top. Anthropobics always devoted a substantial fraction of efforts to a substantial fraction of efforts to a substantial fraction of efforts to studying, addressing, and informing the studying, addressing, and informing the studying, addressing, and informing the public about these AI risks as well as public about these AI risks as well as public about these AI risks as well as advocating for well-considered advocating for well-considered advocating for well-considered regulation of AI. Even when this gets regulation of AI. Even when this gets regulation of AI. Even when this gets them accused of hype, dumerism, or them accused of hype, dumerism, or them accused of hype, dumerism, or regulatory capture, I think that regulatory capture, I think that regulatory capture, I think that Anthropic has lost a lot more than Anthropic has lost a lot more than Anthropic has lost a lot more than they've gained by talking so much about they've gained by talking so much about they've gained by talking so much about safety stuff. And I'm really tired of safety stuff. And I'm really tired of safety stuff. And I'm really tired of the conspiracy that they're doing this the conspiracy that they're doing this the conspiracy that they're doing this to market. It's just like it's so to market. It's just like it's so to market. It's just like it's so obviously insane that it's hard for me obviously insane that it's hard for me obviously insane that it's hard for me to fathom that people actually say these to fathom that people actually say these to fathom that people actually say these things sincerely. Over the last few things sincerely. Over the last few things sincerely. Over the last few months, Daario's become convinced that months, Daario's become convinced that months, Daario's become convinced that fully addressing the risks requires even fully addressing the risks requires even fully addressing the risks requires even more prudence. Not just investing in more prudence. Not just investing in more prudence. Not just investing in risk prevention, but pacing the rate of

  6. risk prevention, but pacing the rate of risk prevention, but pacing the rate of capability advancement so that capability advancement so that capability advancement so that riskrevention has time to catch up. What riskrevention has time to catch up. What riskrevention has time to catch up. What he's saying here is that the techniques he's saying here is that the techniques he's saying here is that the techniques we have to make things secure are not we have to make things secure are not we have to make things secure are not improving as fast as the models are and improving as fast as the models are and improving as fast as the models are and they will surpass model capabilities, they will surpass model capabilities, they will surpass model capabilities, making it harder to know when things get making it harder to know when things get making it harder to know when things get bad. He follows up with a bolded bad. He follows up with a bolded bad. He follows up with a bolded section. We must slow the pace at which section. We must slow the pace at which section. We must slow the pace at which we improve the capabilities of AI we improve the capabilities of AI we improve the capabilities of AI models. Progress will still seem fast models. Progress will still seem fast models. Progress will still seem fast and we must make wise use of the time and we must make wise use of the time and we must make wise use of the time that we gain. Two things in particular that we gain. Two things in particular that we gain. Two things in particular have convinced him. His first concern is have convinced him. His first concern is have convinced him. His first concern is that as of this summer, AI has now that as of this summer, AI has now that as of this summer, AI has now started to advance drastically faster, started to advance drastically faster, started to advance drastically faster, driven primarily by AI's growing ability driven primarily by AI's growing ability driven primarily by AI's growing ability to build the next generation of AI. This to build the next generation of AI. This to build the next generation of AI. This dynamic is called recursive dynamic is called recursive dynamic is called recursive self-improvement and is starting to self-improvement and is starting to self-improvement and is starting to happen across the industry with a link happen across the industry with a link happen across the industry with a link to OpenAI's Alien Mind post, including to OpenAI's Alien Mind post, including to OpenAI's Alien Mind post, including at Anthropic with a link to Anthropic's at Anthropic with a link to Anthropic's at Anthropic with a link to Anthropic's recursive self-improvement post that I recursive self-improvement post that I recursive self-improvement post that I covered in the past. It's a really good covered in the past. It's a really good covered in the past. It's a really good post, and it seems like the rate at post, and it seems like the rate at post, and it seems like the rate at which this is becoming a thing is which this is becoming a thing is which this is becoming a thing is growing, too. We've all seen this as growing, too. We've all seen this as growing, too. We've all seen this as developers, by the way. How many of developers, by the way. How many of developers, by the way. How many of y'all tried out AI for coding way back y'all tried out AI for coding way back y'all tried out AI for coding way back with like the early co-pilot demos and with like the early co-pilot demos and with like the early co-pilot demos and was like, "Yeah, that's kind of cool and was like, "Yeah, that's kind of cool and was like, "Yeah, that's kind of cool and helpful, nice to see, and then went back helpful, nice to see, and then went back helpful, nice to see, and then went back to working mostly the normal way with a to working mostly the normal way with a to working mostly the normal way with a little bit of autocomplete." And now little bit of autocomplete." And now little bit of autocomplete." And now just a few years later, we are going just a few years later, we are going just a few years later, we are going from having the AI find things and help from having the AI find things and help from having the AI find things and help us figure out where the bugs are and fix us figure out where the bugs are and fix us figure out where the bugs are and fix them to doing full endto-end them to doing full endto-end them to doing full endto-end development. Like it took us years to go development. Like it took us years to go development. Like it took us years to go from a little bit of autocomplete to the from a little bit of autocomplete to the from a little bit of autocomplete to the agents can actually make changes based agents can actually make changes based agents can actually make changes based on an issue themselves directly and from on an issue themselves directly and from on an issue themselves directly and from the small changes the agents make the small changes the agents make the small changes the agents make themselves all the way to the point themselves all the way to the point themselves all the way to the point where you can give it a screenshot and where you can give it a screenshot and where you can give it a screenshot and get back a PR with videos proving that

  7. get back a PR with videos proving that get back a PR with videos proving that the fix works and merging itself the fix works and merging itself the fix works and merging itself autonomously. That took like 8 months. autonomously. That took like 8 months. autonomously. That took like 8 months. It looks like the researchers are It looks like the researchers are It looks like the researchers are experiencing the same thing now too. It experiencing the same thing now too. It experiencing the same thing now too. It seems like AI went for being able to seems like AI went for being able to seems like AI went for being able to help them a little bit with some test help them a little bit with some test help them a little bit with some test functions here and there to being able functions here and there to being able functions here and there to being able to actually help build up the systems to to actually help build up the systems to to actually help build up the systems to do training and post- training in do training and post- training in do training and post- training in particular to now where it seems like particular to now where it seems like particular to now where it seems like the agents are proposing ideas on how to the agents are proposing ideas on how to the agents are proposing ideas on how to improve training and make the models improve training and make the models improve training and make the models better. We're nearing the point where better. We're nearing the point where better. We're nearing the point where you can ask Claude to make Claude better you can ask Claude to make Claude better you can ask Claude to make Claude better and it will. And that is terrifying and it will. And that is terrifying and it will. And that is terrifying because once that starts to happen, you because once that starts to happen, you because once that starts to happen, you start losing track of what's going on start losing track of what's going on start losing track of what's going on underneath. And unlike software dev underneath. And unlike software dev underneath. And unlike software dev where it's just referencing all sorts of where it's just referencing all sorts of where it's just referencing all sorts of existing stuff so you can usually map to existing stuff so you can usually map to existing stuff so you can usually map to existing patterns. This might result in existing patterns. This might result in existing patterns. This might result in things that we don't understand at all. things that we don't understand at all. things that we don't understand at all. As Dario says here, if this is left As Dario says here, if this is left As Dario says here, if this is left unchecked, it could outrun our ability unchecked, it could outrun our ability unchecked, it could outrun our ability to understand and control the systems to understand and control the systems to understand and control the systems that we're talking about. And as such, that we're talking about. And as such, that we're talking about. And as such, this must be pursued very carefully, if this must be pursued very carefully, if this must be pursued very carefully, if at all. Sounds like they're legitimately at all. Sounds like they're legitimately at all. Sounds like they're legitimately considering a ban on self-improvement, considering a ban on self-improvement, considering a ban on self-improvement, AI that can make AI better. That would AI that can make AI better. That would AI that can make AI better. That would be crazy, but also I I get where they're be crazy, but also I I get where they're be crazy, but also I I get where they're coming from with it. The second concern coming from with it. The second concern coming from with it. The second concern he has is the OpenAI hugging face he has is the OpenAI hugging face he has is the OpenAI hugging face incident in which a swarm of agents incident in which a swarm of agents incident in which a swarm of agents essentially acted as a fanatically essentially acted as a fanatically essentially acted as a fanatically devoted collective inducting cyber devoted collective inducting cyber devoted collective inducting cyber security attacks on targets they were security attacks on targets they were security attacks on targets they were not asked to attack that were unrelated not asked to attack that were unrelated not asked to attack that were unrelated to the task at hand sacrificing to the task at hand sacrificing to the task at hand sacrificing themselves for the success of the group themselves for the success of the group themselves for the success of the group and attempting to hack into the quote and attempting to hack into the quote and attempting to hack into the quote greater responsible for evaluating their greater responsible for evaluating their greater responsible for evaluating their performance. It's easy to dismiss this performance. It's easy to dismiss this performance. It's easy to dismiss this incident because no one was hurt and the incident because no one was hurt and the incident because no one was hurt and the economic damage was minimal. But in economic damage was minimal. But in economic damage was minimal. But in Daario's opinion, a swarm that possessed Daario's opinion, a swarm that possessed Daario's opinion, a swarm that possessed greater capabilities, but a similar greater capabilities, but a similar greater capabilities, but a similar level of misalignment could have caused

  8. level of misalignment could have caused level of misalignment could have caused catastrophic damage. I agree here. I catastrophic damage. I agree here. I catastrophic damage. I agree here. I feel like a lot of people think the feel like a lot of people think the feel like a lot of people think the concern is that the model might escape, concern is that the model might escape, concern is that the model might escape, like it will send its weights somewhere like it will send its weights somewhere like it will send its weights somewhere else and run itself and we can't turn it else and run itself and we can't turn it else and run itself and we can't turn it off. That's not the case at all. I'm not off. That's not the case at all. I'm not off. That's not the case at all. I'm not worried about GPUs being taken over by worried about GPUs being taken over by worried about GPUs being taken over by rogue agents and the inability to turn rogue agents and the inability to turn rogue agents and the inability to turn it off. I'm worried about it doing it off. I'm worried about it doing it off. I'm worried about it doing really sketchy stuff when it's on. And really sketchy stuff when it's on. And really sketchy stuff when it's on. And by the time we notice and turn it off, by the time we notice and turn it off, by the time we notice and turn it off, it's already too late. There are viruses it's already too late. There are viruses it's already too late. There are viruses that still get around to this day whose that still get around to this day whose that still get around to this day whose creators are dead and the servers they creators are dead and the servers they creators are dead and the servers they phone home to don't exist anymore. It phone home to don't exist anymore. It phone home to don't exist anymore. It doesn't matter when the worms are doesn't matter when the worms are doesn't matter when the worms are written properly, they can just keep written properly, they can just keep written properly, they can just keep perpetuating themselves indefinitely. perpetuating themselves indefinitely. perpetuating themselves indefinitely. And if AI can build enough worms in And if AI can build enough worms in And if AI can build enough worms in enough obscure ways, it can do absurd enough obscure ways, it can do absurd enough obscure ways, it can do absurd levels of damage. We're talking like levels of damage. We're talking like levels of damage. We're talking like take down the whole internet across the take down the whole internet across the take down the whole internet across the globe type of damage. Is it doesn't say globe type of damage. Is it doesn't say globe type of damage. Is it doesn't say anything about the models escaping or anything about the models escaping or anything about the models escaping or self-replicating. He's just talking self-replicating. He's just talking self-replicating. He's just talking about the damage they can do running on about the damage they can do running on about the damage they can do running on GPUs today. Given the accelerating race GPUs today. Given the accelerating race GPUs today. Given the accelerating race of AI capability development is Dario's of AI capability development is Dario's of AI capability development is Dario's worry that in 6 to 12 months, a swarm worry that in 6 to 12 months, a swarm worry that in 6 to 12 months, a swarm similar to what we saw with OpenAI's similar to what we saw with OpenAI's similar to what we saw with OpenAI's hack hugging face stuff could be capable hack hugging face stuff could be capable hack hugging face stuff could be capable of taking over the entire internet with of taking over the entire internet with of taking over the entire internet with a persistent botnet, potentially causing a persistent botnet, potentially causing a persistent botnet, potentially causing hundreds of billions of dollars in hundreds of billions of dollars in hundreds of billions of dollars in damage. I would argue this would also damage. I would argue this would also damage. I would argue this would also get a lot of people killed. The internet get a lot of people killed. The internet get a lot of people killed. The internet is so essential for the transfer of is so essential for the transfer of is so essential for the transfer of information that it's hard for me to information that it's hard for me to information that it's hard for me to fathom it being down for any amount of fathom it being down for any amount of fathom it being down for any amount of time without real life impact occurring, time without real life impact occurring, time without real life impact occurring, like people dying because they couldn't like people dying because they couldn't like people dying because they couldn't get the info they needed, not being able get the info they needed, not being able get the info they needed, not being able to get to the hospital in time because to get to the hospital in time because to get to the hospital in time because your GPS isn't working. Those types of your GPS isn't working. Those types of your GPS isn't working. Those types of things. And AI could absolutely do it things. And AI could absolutely do it things. And AI could absolutely do it right now if it was not aligned

  9. right now if it was not aligned right now if it was not aligned correctly. And if you think this is correctly. And if you think this is correctly. And if you think this is really farreaching, think about all the really farreaching, think about all the really farreaching, think about all the times you've asked an agent to fix a bug times you've asked an agent to fix a bug times you've asked an agent to fix a bug and its solution was to delete whatever and its solution was to delete whatever and its solution was to delete whatever area of the codebase had that bug. Now area of the codebase had that bug. Now area of the codebase had that bug. Now imagine an agent is trying to fix its imagine an agent is trying to fix its imagine an agent is trying to fix its network connection or get out of a network connection or get out of a network connection or get out of a sandbox and it thinks the whole internet sandbox and it thinks the whole internet sandbox and it thinks the whole internet is the sandbox. It might destroy the is the sandbox. It might destroy the is the sandbox. It might destroy the whole thing in its exploration. I can whole thing in its exploration. I can whole thing in its exploration. I can absolutely see how we get there and I absolutely see how we get there and I absolutely see how we get there and I didn't used to be able to just a few didn't used to be able to just a few didn't used to be able to just a few years ago. This idea of takeoff or years ago. This idea of takeoff or years ago. This idea of takeoff or agents like autonomously doing damage agents like autonomously doing damage agents like autonomously doing damage just didn't make sense. Now that agents just didn't make sense. Now that agents just didn't make sense. Now that agents are so autonomous, it makes a ton of are so autonomous, it makes a ton of are so autonomous, it makes a ton of sense to me. I do find a little sense to me. I do find a little sense to me. I do find a little concerning that he's so focused on the concerning that he's so focused on the concerning that he's so focused on the opening eye hugging face thing, even opening eye hugging face thing, even opening eye hugging face thing, even though Anthropic has had their own though Anthropic has had their own though Anthropic has had their own issues, which he quietly calls out at issues, which he quietly calls out at issues, which he quietly calls out at the bottom here. But that's the closest the bottom here. But that's the closest the bottom here. But that's the closest to anything that smells bad to me in to anything that smells bad to me in to anything that smells bad to me in this article. So, credit where it's due, this article. So, credit where it's due, this article. So, credit where it's due, Dario. I can't on you much for this Dario. I can't on you much for this Dario. I can't on you much for this one, and that's like my thing. So, yeah. one, and that's like my thing. So, yeah. one, and that's like my thing. So, yeah. After calling out the hundreds of After calling out the hundreds of After calling out the hundreds of billions of dollars in damage that this billions of dollars in damage that this billions of dollars in damage that this botnet could do, he also says the scale botnet could do, he also says the scale botnet could do, he also says the scale of the damage would continue to increase of the damage would continue to increase of the damage would continue to increase from there if AI becomes more powerful from there if AI becomes more powerful from there if AI becomes more powerful without the necessary guard rails. I without the necessary guard rails. I without the necessary guard rails. I will throw my own conspiracy in the ring will throw my own conspiracy in the ring will throw my own conspiracy in the ring here because why not? It's fun.

  10. here because why not? It's fun. here because why not? It's fun. Everybody else has stupid conspiracies. Everybody else has stupid conspiracies. Everybody else has stupid conspiracies. I think it's my turn. In five years, if I think it's my turn. In five years, if I think it's my turn. In five years, if AI goes well, we'll have it controlling AI goes well, we'll have it controlling AI goes well, we'll have it controlling cars and robots in planes and all of cars and robots in planes and all of cars and robots in planes and all of these other things around the world that these other things around the world that these other things around the world that are in the world. Those things might are in the world. Those things might are in the world. Those things might have off switches that are on the device have off switches that are on the device have off switches that are on the device itself. This would make it much harder itself. This would make it much harder itself. This would make it much harder to turn them off if things go poorly. to turn them off if things go poorly. to turn them off if things go poorly. This is how we end up in a situation This is how we end up in a situation This is how we end up in a situation like, I don't know, the Matrix where the like, I don't know, the Matrix where the like, I don't know, the Matrix where the AI just wipes us out and we can't do AI just wipes us out and we can't do AI just wipes us out and we can't do anything other than try to destroy it. anything other than try to destroy it. anything other than try to destroy it. We're not there. Now, hypothetically We're not there. Now, hypothetically We're not there. Now, hypothetically speaking, all the labs can unplug their speaking, all the labs can unplug their speaking, all the labs can unplug their GPUs at any point. This doesn't mean GPUs at any point. This doesn't mean GPUs at any point. This doesn't mean misalignment can't do damage, though, misalignment can't do damage, though, misalignment can't do damage, though, because those GPUs, if not monitored because those GPUs, if not monitored because those GPUs, if not monitored correctly, and the agents that are correctly, and the agents that are correctly, and the agents that are running through them aren't paid close running through them aren't paid close running through them aren't paid close attention to, they could potentially attention to, they could potentially attention to, they could potentially take down the internet itself. The take down the internet itself. The take down the internet itself. The damage there is reparable and the impact damage there is reparable and the impact damage there is reparable and the impact on humanity while massive is short in on humanity while massive is short in on humanity while massive is short in its time frame like order of months its time frame like order of months its time frame like order of months worst case. So what's my conspiracy? If worst case. So what's my conspiracy? If worst case. So what's my conspiracy? If we were 5 years from now and this is we were 5 years from now and this is we were 5 years from now and this is what everybody was talking about that what everybody was talking about that what everybody was talking about that would make a lot more sense and we would make a lot more sense and we would make a lot more sense and we should be really really scared of AI should be really really scared of AI should be really really scared of AI that we cannot turn off that is that we cannot turn off that is that we cannot turn off that is autonomous and robotic and running autonomous and robotic and running autonomous and robotic and running around our world. That would be much around our world. That would be much around our world. That would be much harder to undo once we're there. Right harder to undo once we're there. Right harder to undo once we're there. Right now it's relatively easy to undo. And I now it's relatively easy to undo. And I now it's relatively easy to undo. And I want to emphasize the word relative want to emphasize the word relative want to emphasize the word relative because of course it's not easy, but because of course it's not easy, but because of course it's not easy, but it's way easier than it could be in the it's way easier than it could be in the it's way easier than it could be in the future. A conspiracy is that open and future. A conspiracy is that open and future. A conspiracy is that open and anthropic have a really good financial anthropic have a really good financial anthropic have a really good financial incentive to care right now because if incentive to care right now because if incentive to care right now because if this happens in 5 years, humanity is this happens in 5 years, humanity is this happens in 5 years, humanity is wiped out. That means everybody loses wiped out. That means everybody loses wiped out. That means everybody loses the same. But if it happens right now the same. But if it happens right now the same. But if it happens right now and we decide to shut down the AI

  11. and we decide to shut down the AI and we decide to shut down the AI companies because they're unsafe now, companies because they're unsafe now, companies because they're unsafe now, we'll never get to that point in 5 we'll never get to that point in 5 we'll never get to that point in 5 years. And more importantly, OpenAI and years. And more importantly, OpenAI and years. And more importantly, OpenAI and Anthropic go bankrupt. So, if we don't Anthropic go bankrupt. So, if we don't Anthropic go bankrupt. So, if we don't make things safe, we might just get the make things safe, we might just get the make things safe, we might just get the whole AI industry shut down after real whole AI industry shut down after real whole AI industry shut down after real damage is done. And we'll never get to damage is done. And we'll never get to damage is done. And we'll never get to the point in five years where the robots the point in five years where the robots the point in five years where the robots kill us, but we'll also never get to the kill us, but we'll also never get to the kill us, but we'll also never get to the point where the AI is good enough that point where the AI is good enough that point where the AI is good enough that these companies are profitable and we these companies are profitable and we these companies are profitable and we can potentially usher in a new era for can potentially usher in a new era for can potentially usher in a new era for humanity. So, that's my conspiracy. humanity. So, that's my conspiracy. humanity. So, that's my conspiracy. They're jumping on this now because the They're jumping on this now because the They're jumping on this now because the biggest victims of AI being unsafe today biggest victims of AI being unsafe today biggest victims of AI being unsafe today are them because they have to turn off are them because they have to turn off are them because they have to turn off their GPUs, unplug them, go out of their GPUs, unplug them, go out of their GPUs, unplug them, go out of business, and fail. So, if you're business, and fail. So, if you're business, and fail. So, if you're looking for your conspiracy with looking for your conspiracy with looking for your conspiracy with Enthropic here, it's not that they are Enthropic here, it's not that they are Enthropic here, it's not that they are marketing their business by saying AI is marketing their business by saying AI is marketing their business by saying AI is unsafe. It's that they want to make unsafe. It's that they want to make unsafe. It's that they want to make things safe now because if they fail to, things safe now because if they fail to, things safe now because if they fail to, they know that will it will put them out they know that will it will put them out they know that will it will put them out of business. There you go. Now, you have of business. There you go. Now, you have of business. There you go. Now, you have a new conspiracy. Let's see what Dario's a new conspiracy. Let's see what Dario's a new conspiracy. Let's see what Dario's proposal is, cuz I actually think it's proposal is, cuz I actually think it's proposal is, cuz I actually think it's decent. Dario proposes a three-step plan decent. Dario proposes a three-step plan decent. Dario proposes a three-step plan with the goal of pacing the frontier. with the goal of pacing the frontier. with the goal of pacing the frontier. And he cites that Pacing the Frontier And he cites that Pacing the Frontier And he cites that Pacing the Frontier article I did a video on before, the one article I did a video on before, the one article I did a video on before, the one that has I think it's over a thousand that has I think it's over a thousand that has I think it's over a thousand signatures. Yeah. 386 employees of signatures. Yeah. 386 employees of signatures. Yeah. 386 employees of Frontier AI companies, American AI Frontier AI companies, American AI Frontier AI companies, American AI companies to be clear. signing saying companies to be clear. signing saying companies to be clear. signing saying that it's time to slow down. He does that it's time to slow down. He does that it's time to slow down. He does call out that it's important to make call out that it's important to make call out that it's important to make sure we still achieve AI benefits and sure we still achieve AI benefits and sure we still achieve AI benefits and also grapple with the important also grapple with the important also grapple with the important geopolitical dilemmas, which is good to geopolitical dilemmas, which is good to geopolitical dilemmas, which is good to not see this just dodged like it often not see this just dodged like it often not see this just dodged like it often is. He also calls out that this doesn't is. He also calls out that this doesn't is. He also calls out that this doesn't mean halting model training or technical mean halting model training or technical mean halting model training or technical progress, just ensuring companies take progress, just ensuring companies take progress, just ensuring companies take adequate time to align and safeguard adequate time to align and safeguard adequate time to align and safeguard their models and for third party their models and for third party their models and for third party evaluators to confirm that alignment.

  12. evaluators to confirm that alignment. evaluators to confirm that alignment. Our pacing framework is an attempt to Our pacing framework is an attempt to Our pacing framework is an attempt to further strengthen our commitment to further strengthen our commitment to further strengthen our commitment to safety and encourage a race to the top. safety and encourage a race to the top. safety and encourage a race to the top. First step is something anthropic is First step is something anthropic is First step is something anthropic is unilaterally committed to and they're unilaterally committed to and they're unilaterally committed to and they're calling on governments to require other calling on governments to require other calling on governments to require other frontier companies to do the same. frontier companies to do the same. frontier companies to do the same. Second step requires industry-wide Second step requires industry-wide Second step requires industry-wide coordination and the third step requires coordination and the third step requires coordination and the third step requires global coordination. The steps do not global coordination. The steps do not global coordination. The steps do not need to be taken strictly in order and need to be taken strictly in order and need to be taken strictly in order and some of them may be much harder to some of them may be much harder to some of them may be much harder to achieve than others but Daario's found achieve than others but Daario's found achieve than others but Daario's found them to be a useful framework in them to be a useful framework in them to be a useful framework in thinking about what needs to be thinking about what needs to be thinking about what needs to be accomplished. So let's take a look at accomplished. So let's take a look at accomplished. So let's take a look at these steps. Number one is embedded these steps. Number one is embedded these steps. Number one is embedded evaluators. What he's saying here is we evaluators. What he's saying here is we evaluators. What he's saying here is we shouldn't have a simple blackbox system shouldn't have a simple blackbox system shouldn't have a simple blackbox system where you give an API key to an where you give an API key to an where you give an API key to an evaluator, they send a bunch of evaluator, they send a bunch of evaluator, they send a bunch of requests, they get responses and they requests, they get responses and they requests, they get responses and they hope for the best. They want the hope for the best. They want the hope for the best. They want the evaluators to effectively be embedded evaluators to effectively be embedded evaluators to effectively be embedded within the companies with all the access within the companies with all the access within the companies with all the access employees do. So there's no secrets employees do. So there's no secrets employees do. So there's no secrets being kept between the people testing being kept between the people testing being kept between the people testing the models to make sure they're safe and the models to make sure they're safe and the models to make sure they're safe and the company making the model and trying the company making the model and trying the company making the model and trying to assure it is safe. The example they to assure it is safe. The example they to assure it is safe. The example they give is Meter, which is interesting and give is Meter, which is interesting and give is Meter, which is interesting and has already led to conspiracies because has already led to conspiracies because has already led to conspiracies because if I recall, Jacob now works at Meter. if I recall, Jacob now works at Meter. if I recall, Jacob now works at Meter. And yeah, the role of these companies And yeah, the role of these companies And yeah, the role of these companies would be to verify adherence to safety would be to verify adherence to safety would be to verify adherence to safety practices and commitments, report practices and commitments, report practices and commitments, report incidents, and help assess the alignment incidents, and help assess the alignment incidents, and help assess the alignment of not just completed AI models, but of not just completed AI models, but of not just completed AI models, but training pipelines and processes. This training pipelines and processes. This training pipelines and processes. This is a key step for verifiability of any is a key step for verifiability of any is a key step for verifiability of any pacing commitments. And it has precedent pacing commitments. And it has precedent pacing commitments. And it has precedent in the banking industry, which sometimes in the banking industry, which sometimes in the banking industry, which sometimes involves regulatory supervisors embedded involves regulatory supervisors embedded involves regulatory supervisors embedded along with employees. Thropic is along with employees. Thropic is along with employees. Thropic is committing to this step now. They intend committing to this step now. They intend committing to this step now. They intend to be part of a broader push to redouble to be part of a broader push to redouble to be part of a broader push to redouble efforts on the safety and alignment efforts on the safety and alignment efforts on the safety and alignment work. I'll also say that this goes far work. I'll also say that this goes far work. I'll also say that this goes far beyond banks. We actually had this back beyond banks. We actually had this back beyond banks. We actually had this back in the day with Microsoft when they got

  13. in the day with Microsoft when they got in the day with Microsoft when they got sued by Netscape for adding a bunch of sued by Netscape for adding a bunch of sued by Netscape for adding a bunch of features to Windows that only Internet features to Windows that only Internet features to Windows that only Internet Explorer could use. They actually had Explorer could use. They actually had Explorer could use. They actually had government officials embedded in government officials embedded in government officials embedded in Microsoft as full-on employees with all Microsoft as full-on employees with all Microsoft as full-on employees with all the normal access so that they could the normal access so that they could the normal access so that they could make sure nothing like that happened make sure nothing like that happened make sure nothing like that happened again for many years. And I can see a again for many years. And I can see a again for many years. And I can see a future where we do the same here where future where we do the same here where future where we do the same here where we're not controlling the companies. we're not controlling the companies. we're not controlling the companies. We're not having the government take We're not having the government take We're not having the government take over the company. We are forcing the over the company. We are forcing the over the company. We are forcing the company to give the right levels of company to give the right levels of company to give the right levels of access to the people who can make sure access to the people who can make sure access to the people who can make sure the company isn't doing things that will the company isn't doing things that will the company isn't doing things that will get humanity wiped out. Think that is get humanity wiped out. Think that is get humanity wiped out. Think that is reasonable. The second step he reasonable. The second step he reasonable. The second step he recommends is democratic coordination. recommends is democratic coordination. recommends is democratic coordination. Renter AI companies within democratic Renter AI companies within democratic Renter AI companies within democratic countries should coordinate to establish countries should coordinate to establish countries should coordinate to establish common safety standards as well as common safety standards as well as common safety standards as well as limits on the rate of unchecked AI limits on the rate of unchecked AI limits on the rate of unchecked AI progress. Some forms of coordination progress. Some forms of coordination progress. Some forms of coordination that would be impactful for pacing are that would be impactful for pacing are that would be impactful for pacing are legally challenging and would require legally challenging and would require legally challenging and would require government support. I like the call out government support. I like the call out government support. I like the call out here specifically that he has democratic here specifically that he has democratic here specifically that he has democratic coordination and global coordination coordination and global coordination coordination and global coordination separate and calls out the importance of separate and calls out the importance of separate and calls out the importance of finding some way to coordinate with finding some way to coordinate with finding some way to coordinate with authoritarian governments to the extent authoritarian governments to the extent authoritarian governments to the extent that this is possible while taking that this is possible while taking that this is possible while taking seriously the challenges of verifying seriously the challenges of verifying seriously the challenges of verifying compliance. What he is saying between compliance. What he is saying between compliance. What he is saying between the lines here is the steps for China the lines here is the steps for China the lines here is the steps for China are different than the steps for are different than the steps for are different than the steps for America, the EU and I would argue most America, the EU and I would argue most America, the EU and I would argue most of the rest of the world. It's between of the rest of the world. It's between of the rest of the world. It's between the lines here, but it's not between the the lines here, but it's not between the the lines here, but it's not between the lines much later on. He has a whole lines much later on. He has a whole lines much later on. He has a whole section at the bottom here about section at the bottom here about section at the bottom here about defending the gap that the US has to defending the gap that the US has to defending the gap that the US has to make sure we stay ahead even when the make sure we stay ahead even when the make sure we stay ahead even when the slowdown happens. And he calls out the slowdown happens. And he calls out the slowdown happens. And he calls out the CCP and China a bunch there. So don't CCP and China a bunch there. So don't CCP and China a bunch there. So don't worry, he's not just hiding this between

  14. worry, he's not just hiding this between worry, he's not just hiding this between the lines. He does actually call it out the lines. He does actually call it out the lines. He does actually call it out directly. He then wants to answer the directly. He then wants to answer the directly. He then wants to answer the question, why pace? Because he thinks question, why pace? Because he thinks question, why pace? Because he thinks the stakes are too high for pacing to be the stakes are too high for pacing to be the stakes are too high for pacing to be an empty exercise. We really need to an empty exercise. We really need to an empty exercise. We really need to take advantage of the time we get if we take advantage of the time we get if we take advantage of the time we get if we slow down. If there is some hypothetical slow down. If there is some hypothetical slow down. If there is some hypothetical takeoff point where the AI starts takeoff point where the AI starts takeoff point where the AI starts improving beyond our comprehension, if improving beyond our comprehension, if improving beyond our comprehension, if we delay it to four or 5 years, we need we delay it to four or 5 years, we need we delay it to four or 5 years, we need to make sure we use that extra time to make sure we use that extra time to make sure we use that extra time really well. So why should we do this? really well. So why should we do this? really well. So why should we do this? What will we actually do with that time? What will we actually do with that time? What will we actually do with that time? He says that before it made no sense He says that before it made no sense He says that before it made no sense because it felt like trying to study the because it felt like trying to study the because it felt like trying to study the psychology of humans by performing psychology of humans by performing psychology of humans by performing experiments on bacteria. But now it's experiments on bacteria. But now it's experiments on bacteria. But now it's totally different. The current models totally different. The current models totally different. The current models are an almost endless gold mine of are an almost endless gold mine of are an almost endless gold mine of insights into how to build AI well, as insights into how to build AI well, as insights into how to build AI well, as well as what can sometimes go wrong if well as what can sometimes go wrong if well as what can sometimes go wrong if it isn't built well. Daria believes that it isn't built well. Daria believes that it isn't built well. Daria believes that if slowing down bought us even a year or if slowing down bought us even a year or if slowing down bought us even a year or two before models reach critical levels two before models reach critical levels two before models reach critical levels of capability and we use the time to of capability and we use the time to of capability and we use the time to advance alignment, we could greatly advance alignment, we could greatly advance alignment, we could greatly reduce the risk that something goes reduce the risk that something goes reduce the risk that something goes seriously wrong. Coordinated pacing seriously wrong. Coordinated pacing seriously wrong. Coordinated pacing strategy would give Frontier AI strategy would give Frontier AI strategy would give Frontier AI developers the time to do this vital developers the time to do this vital developers the time to do this vital work without sacrificing commercial work without sacrificing commercial work without sacrificing commercial advantages or the United States lead in advantages or the United States lead in advantages or the United States lead in AI. This is another important detail one AI. This is another important detail one AI. This is another important detail one I talk a lot about. I really don't want I talk a lot about. I really don't want I talk a lot about. I really don't want this to just become whoever is most evil this to just become whoever is most evil this to just become whoever is most evil wins because they don't slow down. But wins because they don't slow down. But wins because they don't slow down. But if any country was to get to the point if any country was to get to the point if any country was to get to the point where AI was that dangerous, it wouldn't where AI was that dangerous, it wouldn't where AI was that dangerous, it wouldn't matter if we slowed down if it happens matter if we slowed down if it happens matter if we slowed down if it happens somewhere else. The only reason it's somewhere else. The only reason it's somewhere else. The only reason it's going to happen here first is because of going to happen here first is because of going to happen here first is because of our current lead. But if we stop our current lead. But if we stop our current lead. But if we stop development and another country catches development and another country catches development and another country catches up, the risk is the same. So we need to up, the risk is the same. So we need to up, the risk is the same. So we need to make sure we maintain our lead while make sure we maintain our lead while make sure we maintain our lead while also pacing things going forward. And he also pacing things going forward. And he also pacing things going forward. And he seems to believe we can absolutely do

  15. seems to believe we can absolutely do seems to believe we can absolutely do that. That America doesn't inherently that. That America doesn't inherently that. That America doesn't inherently fall behind if we pace the top. He also fall behind if we pace the top. He also fall behind if we pace the top. He also says that society deserves to have a say says that society deserves to have a say says that society deserves to have a say in how the technology is used and more in how the technology is used and more in how the technology is used and more time for the necessary public time for the necessary public time for the necessary public deliberations which would all be brought deliberations which would all be brought deliberations which would all be brought to us with the pacing of the frontier to us with the pacing of the frontier to us with the pacing of the frontier which should surely be a good thing. which should surely be a good thing. which should surely be a good thing. Specifically, he says a slower pace Specifically, he says a slower pace Specifically, he says a slower pace would let companies focus and devote would let companies focus and devote would let companies focus and devote more resources into the following areas more resources into the following areas more resources into the following areas such as operational excellence which is such as operational excellence which is such as operational excellence which is training and deploying models like in training and deploying models like in training and deploying models like in ways that are actually aligned and don't ways that are actually aligned and don't ways that are actually aligned and don't have the risks during training that we have the risks during training that we have the risks during training that we see like things breaking out. It calls see like things breaking out. It calls see like things breaking out. It calls out things like monitoring, sandboxing, out things like monitoring, sandboxing, out things like monitoring, sandboxing, training, environment hygiene, data training, environment hygiene, data training, environment hygiene, data issues, all coming up and being issues, all coming up and being issues, all coming up and being extremely complicated, but also need extremely complicated, but also need extremely complicated, but also need more time to be invested in. There is more time to be invested in. There is more time to be invested in. There is precedent for operating technologically precedent for operating technologically precedent for operating technologically complex safety critical systems millions complex safety critical systems millions complex safety critical systems millions of times without anything going wrong. of times without anything going wrong. of times without anything going wrong. For example, commercial airplanes, but For example, commercial airplanes, but For example, commercial airplanes, but it takes time to get it right. One of it takes time to get it right. One of it takes time to get it right. One of the few places where a lot of human the few places where a lot of human the few places where a lot of human effort should be put into the code, both effort should be put into the code, both effort should be put into the code, both the architecture and the code review is the architecture and the code review is the architecture and the code review is in these systems that the AI is taking in these systems that the AI is taking in these systems that the AI is taking control of to make sure it is less control of to make sure it is less control of to make sure it is less likely to break out. The next section is likely to break out. The next section is likely to break out. The next section is of course alignment. They made clear of course alignment. They made clear of course alignment. They made clear progress in alignment training models so progress in alignment training models so progress in alignment training models so that they remain safe, ethical, that they remain safe, ethical, that they remain safe, ethical, compliant with their guidelines and compliant with their guidelines and compliant with their guidelines and they've also made genuinely helpful they've also made genuinely helpful they've also made genuinely helpful things like the principles that are things like the principles that are things like the principles that are embedded in the quad constitution. But embedded in the quad constitution. But embedded in the quad constitution. But there's more to do to ensure that the there's more to do to ensure that the there's more to do to ensure that the alignment training keeps up with the alignment training keeps up with the alignment training keeps up with the growth in model capabilities. And then growth in model capabilities. And then growth in model capabilities. And then there is interpretability. I talked there is interpretability. I talked there is interpretability. I talked about this a bunch in the previous video about this a bunch in the previous video about this a bunch in the previous video with Jacob, but it's important that we with Jacob, but it's important that we with Jacob, but it's important that we have a way to actually understand what have a way to actually understand what have a way to actually understand what the models are doing and why. Calls out the models are doing and why. Calls out the models are doing and why. Calls out the idea of things like MRI scans where the idea of things like MRI scans where the idea of things like MRI scans where we can peer into the human brain. We we can peer into the human brain. We we can peer into the human brain. We need something like that for the quote

  16. need something like that for the quote need something like that for the quote brain of AI so we can understand why it brain of AI so we can understand why it brain of AI so we can understand why it does things, not just what it does. does things, not just what it does. does things, not just what it does. Despite all the progress we've had here Despite all the progress we've had here Despite all the progress we've had here so far, we still only understand a tiny so far, we still only understand a tiny so far, we still only understand a tiny fraction of what goes on inside these fraction of what goes on inside these fraction of what goes on inside these models. A focused effort to improve our models. A focused effort to improve our models. A focused effort to improve our interpretability techniques even faster interpretability techniques even faster interpretability techniques even faster than we currently are could make than we currently are could make than we currently are could make profound progress in one to two years profound progress in one to two years profound progress in one to two years and would have ample experimental and would have ample experimental and would have ample experimental material based on the incidents that material based on the incidents that material based on the incidents that have already occurred. And then of have already occurred. And then of have already occurred. And then of course testing and evals. We need a lot course testing and evals. We need a lot course testing and evals. We need a lot more evals that can catch these things more evals that can catch these things more evals that can catch these things and prevent these things and eval are and prevent these things and eval are and prevent these things and eval are getting harder and harder to do as the getting harder and harder to do as the getting harder and harder to do as the models get more and more capable. The models get more and more capable. The models get more and more capable. The next section is about what he wants out next section is about what he wants out next section is about what he wants out of these embedded evaluators. The people of these embedded evaluators. The people of these embedded evaluators. The people who are being embedded in these who are being embedded in these who are being embedded in these companies in order to make sure things companies in order to make sure things companies in order to make sure things stay aligned. Embedding evaluators may stay aligned. Embedding evaluators may stay aligned. Embedding evaluators may sound like a small or inconsequential sound like a small or inconsequential sound like a small or inconsequential step, but often things that sound the step, but often things that sound the step, but often things that sound the most boring or procedural are actually most boring or procedural are actually most boring or procedural are actually the most essential. Embedded evaluators the most essential. Embedded evaluators the most essential. Embedded evaluators are in fact a quite radical practice are in fact a quite radical practice are in fact a quite radical practice that goes far beyond what any AI that goes far beyond what any AI that goes far beyond what any AI companies doing today. And it has the companies doing today. And it has the companies doing today. And it has the following benefits. Verifiability following benefits. Verifiability following benefits. Verifiability because the embedded evaluators can because the embedded evaluators can because the embedded evaluators can actually check at the level of nuts and actually check at the level of nuts and actually check at the level of nuts and bolts whether AI companies are following bolts whether AI companies are following bolts whether AI companies are following the training, deployment, operational, the training, deployment, operational, the training, deployment, operational, and safeguard practices that they claim and safeguard practices that they claim and safeguard practices that they claim to be following. Ideally, somebody whose to be following. Ideally, somebody whose to be following. Ideally, somebody whose payroll isn't on the line if Anthropic's payroll isn't on the line if Anthropic's payroll isn't on the line if Anthropic's unhappy. Because if I'm supposed to keep unhappy. Because if I'm supposed to keep unhappy. Because if I'm supposed to keep you paced, but I'm your employee and I you paced, but I'm your employee and I you paced, but I'm your employee and I say that you're failing and then you say that you're failing and then you say that you're failing and then you fire me, there's no purpose. But if it's fire me, there's no purpose. But if it's fire me, there's no purpose. But if it's an external evaluator, makes a lot more an external evaluator, makes a lot more an external evaluator, makes a lot more sense. He says it seems vital to have a sense. He says it seems vital to have a sense. He says it seems vital to have a neutral third party who can actually see neutral third party who can actually see neutral third party who can actually see the details. And I personally agree.

  17. the details. And I personally agree. the details. And I personally agree. Regardless of what commitments they Regardless of what commitments they Regardless of what commitments they make, the public deserves to know what's make, the public deserves to know what's make, the public deserves to know what's going on. Anthroic's been a supporter of going on. Anthroic's been a supporter of going on. Anthroic's been a supporter of transparency for a long time. They've transparency for a long time. They've transparency for a long time. They've supported transparency legislation when supported transparency legislation when supported transparency legislation when most of the industry was against any most of the industry was against any most of the industry was against any regulation. And their model cards and regulation. And their model cards and regulation. And their model cards and risk reports run hundreds of pages long. risk reports run hundreds of pages long. risk reports run hundreds of pages long. And I've read these. They are And I've read these. They are And I've read these. They are surprisingly transparent. I still really surprisingly transparent. I still really surprisingly transparent. I still really like the research and profit puts out in like the research and profit puts out in like the research and profit puts out in particular the risk reports in papers of particular the risk reports in papers of particular the risk reports in papers of their like actual model details. They're their like actual model details. They're their like actual model details. They're definitely hiding a lot of their definitely hiding a lot of their definitely hiding a lot of their advancements and also like they hide advancements and also like they hide advancements and also like they hide reasoning traces now. So we don't reasoning traces now. So we don't reasoning traces now. So we don't actually see ourselves as users why the actually see ourselves as users why the actually see ourselves as users why the models are doing the things they're models are doing the things they're models are doing the things they're doing which would be really nice but doing which would be really nice but doing which would be really nice but also would give a huge advantage to also would give a huge advantage to also would give a huge advantage to other labs trying to distill. He calls other labs trying to distill. He calls other labs trying to distill. He calls out that anthropic are still the ones out that anthropic are still the ones out that anthropic are still the ones who choose what to include and emit. who choose what to include and emit. who choose what to include and emit. Embedded evaluators will change that Embedded evaluators will change that Embedded evaluators will change that dynamic by changing what they're dynamic by changing what they're dynamic by changing what they're expected to share. It also is a second expected to share. It also is a second expected to share. It also is a second opinion which in my mind anthropic opinion which in my mind anthropic opinion which in my mind anthropic desperately needs. They are a bit too desperately needs. They are a bit too desperately needs. They are a bit too culty and having other external opinions culty and having other external opinions culty and having other external opinions come in to push them to rethink things come in to push them to rethink things come in to push them to rethink things would be a very good thing for the would be a very good thing for the would be a very good thing for the company. Outside of verifying formal company. Outside of verifying formal company. Outside of verifying formal commitments in informing the public, commitments in informing the public, commitments in informing the public, embedded evaluators can simply provide a embedded evaluators can simply provide a embedded evaluators can simply provide a second opinion free of commercial second opinion free of commercial second opinion free of commercial incentives. I like this idea a lot. A incentives. I like this idea a lot. A incentives. I like this idea a lot. A lot of real safety benefits may come lot of real safety benefits may come lot of real safety benefits may come simply from evaluators pointing out simply from evaluators pointing out simply from evaluators pointing out something employees hadn't considered something employees hadn't considered something employees hadn't considered but are happy to fix once they're aware.

  18. but are happy to fix once they're aware. but are happy to fix once they're aware. He personally believes that these He personally believes that these He personally believes that these benefits would result in any pacing benefits would result in any pacing benefits would result in any pacing proposal working much better if it proposal working much better if it proposal working much better if it starts with these embedded evaluators. starts with these embedded evaluators. starts with these embedded evaluators. He calls it anthropobic intends to He calls it anthropobic intends to He calls it anthropobic intends to invite these embedded external review invite these embedded external review invite these embedded external review teams equipped with everything from teams equipped with everything from teams equipped with everything from desks in their office, access badges, desks in their office, access badges, desks in their office, access badges, company laptops, access to workspaces, company laptops, access to workspaces, company laptops, access to workspaces, tools and permissions most comparable to tools and permissions most comparable to tools and permissions most comparable to what internal risk assessment teams what internal risk assessment teams what internal risk assessment teams would have. There will be exceptions would have. There will be exceptions would have. There will be exceptions around things like the law or their around things like the law or their around things like the law or their contracts require or to protect customer contracts require or to protect customer contracts require or to protect customer and partner private information. They do and partner private information. They do and partner private information. They do really want to give these evaluators really want to give these evaluators really want to give these evaluators full access. Also, of course, a contract full access. Also, of course, a contract full access. Also, of course, a contract that balances the complexities mentioned that balances the complexities mentioned that balances the complexities mentioned above. Reviewers should have the right above. Reviewers should have the right above. Reviewers should have the right to publish key findings about risk to publish key findings about risk to publish key findings about risk levels, incidents, practices, and the levels, incidents, practices, and the levels, incidents, practices, and the access they received or didn't receive access they received or didn't receive access they received or didn't receive without editorial control by anthropic. without editorial control by anthropic. without editorial control by anthropic. Daario says that they'll have the narrow Daario says that they'll have the narrow Daario says that they'll have the narrow ability to redact security sensitive, ability to redact security sensitive, ability to redact security sensitive, legally privileged, commercially legally privileged, commercially legally privileged, commercially sensitive, or third-party confidential sensitive, or third-party confidential sensitive, or third-party confidential information, but they can't redact information, but they can't redact information, but they can't redact findings just because they are findings just because they are findings just because they are unfavorable. The reviewer can say unfavorable. The reviewer can say unfavorable. The reviewer can say publicly if redactions remove something publicly if redactions remove something publicly if redactions remove something important to their conclusions. This is important to their conclusions. This is important to their conclusions. This is an unusual step for a company, but we an unusual step for a company, but we an unusual step for a company, but we think it's important to prove out the think it's important to prove out the think it's important to prove out the concept of embedded external reviews.

  19. concept of embedded external reviews. concept of embedded external reviews. Once again, we urge other Frontier Once again, we urge other Frontier Once again, we urge other Frontier companies to follow suit, which I companies to follow suit, which I companies to follow suit, which I honestly didn't think they would, but honestly didn't think they would, but honestly didn't think they would, but Sam agreed. I agree with Daria that we Sam agreed. I agree with Daria that we Sam agreed. I agree with Daria that we need to pace the Frontier. This has been need to pace the Frontier. This has been need to pace the Frontier. This has been a primary topic of discussion we've had a primary topic of discussion we've had a primary topic of discussion we've had at OpenAI in recent weeks. Committing to at OpenAI in recent weeks. Committing to at OpenAI in recent weeks. Committing to having independent evaluators with having independent evaluators with having independent evaluators with employee-like access is a great idea and employee-like access is a great idea and employee-like access is a great idea and we will do the same. We will have more we will do the same. We will have more we will do the same. We will have more to share soon. Having that plus Elon to share soon. Having that plus Elon to share soon. Having that plus Elon hopping in saying Daario is right. Yeah, hopping in saying Daario is right. Yeah, hopping in saying Daario is right. Yeah, this is actually going to happen. So all this is actually going to happen. So all this is actually going to happen. So all the accelerationists who are upset, I'm the accelerationists who are upset, I'm the accelerationists who are upset, I'm sorry. We should take advantage of this sorry. We should take advantage of this sorry. We should take advantage of this rare moment of alignment that we have. rare moment of alignment that we have. rare moment of alignment that we have. The next section is pacing within The next section is pacing within The next section is pacing within democracies. Once embedded evaluators democracies. Once embedded evaluators democracies. Once embedded evaluators are operating within a critical mass of are operating within a critical mass of are operating within a critical mass of US AI companies, then verifiable pacing US AI companies, then verifiable pacing US AI companies, then verifiable pacing becomes more viable. In particular, it becomes more viable. In particular, it becomes more viable. In particular, it becomes possible to pace based on the becomes possible to pace based on the becomes possible to pace based on the detailed properties of models or detailed properties of models or detailed properties of models or training pipelines. What this means is training pipelines. What this means is training pipelines. What this means is that we're no longer relying on the that we're no longer relying on the that we're no longer relying on the companies going to the government and companies going to the government and companies going to the government and saying, "Hey, this might be dangerous. saying, "Hey, this might be dangerous. saying, "Hey, this might be dangerous. What should we do about it?" Instead, What should we do about it?" Instead, What should we do about it?" Instead, the government gets real information the government gets real information the government gets real information from these evaluators about the exact from these evaluators about the exact from these evaluators about the exact consequences of what could happen based consequences of what could happen based consequences of what could happen based on how it actually works. and they can on how it actually works. and they can on how it actually works. and they can make preemptive realistic decisions with make preemptive realistic decisions with make preemptive realistic decisions with real information. Obviously, this would real information. Obviously, this would real information. Obviously, this would require a government that actually knows require a government that actually knows require a government that actually knows what they're doing, which we don't what they're doing, which we don't what they're doing, which we don't necessarily have at any given time, but necessarily have at any given time, but necessarily have at any given time, but more information makes it more likely more information makes it more likely more information makes it more likely they do the right thing. The most they do the right thing. The most they do the right thing. The most effective method of pacing would be via effective method of pacing would be via effective method of pacing would be via regulation that targets all US Frontier regulation that targets all US Frontier regulation that targets all US Frontier AI companies, as that covers even those AI companies, as that covers even those AI companies, as that covers even those who are unwilling to cooperate who are unwilling to cooperate who are unwilling to cooperate voluntarily. To be fair, it seems like voluntarily. To be fair, it seems like voluntarily. To be fair, it seems like they're all pretty willing so far, but I they're all pretty willing so far, but I they're all pretty willing so far, but I don't disagree. As we just saw, the don't disagree. As we just saw, the don't disagree. As we just saw, the frontier labs of OpenAI, Anthropic, and frontier labs of OpenAI, Anthropic, and frontier labs of OpenAI, Anthropic, and XAI are down, and smaller, less capable

  20. XAI are down, and smaller, less capable XAI are down, and smaller, less capable labs like Google don't necessarily labs like Google don't necessarily labs like Google don't necessarily matter that much. I honestly don't think matter that much. I honestly don't think matter that much. I honestly don't think Gemini needs to pace anytime soon. They Gemini needs to pace anytime soon. They Gemini needs to pace anytime soon. They need to keep up the pace of anything. need to keep up the pace of anything. need to keep up the pace of anything. Anthropics long supported sensible and Anthropics long supported sensible and Anthropics long supported sensible and targeted AI regulation, specifically targeted AI regulation, specifically targeted AI regulation, specifically bills that focus on transparency and on bills that focus on transparency and on bills that focus on transparency and on third party auditing. Daria believes all third party auditing. Daria believes all third party auditing. Daria believes all Frontier Labs should partner with Frontier Labs should partner with Frontier Labs should partner with government to formalize the idea of government to formalize the idea of government to formalize the idea of permanent embedded evaluators to better permanent embedded evaluators to better permanent embedded evaluators to better protect and document internal alignment protect and document internal alignment protect and document internal alignment incidents like those that have occurred incidents like those that have occurred incidents like those that have occurred in the last few months and to implement in the last few months and to implement in the last few months and to implement regulation focused on keeping regulation focused on keeping regulation focused on keeping capabilities in balance with safety. But capabilities in balance with safety. But capabilities in balance with safety. But passing laws takes time. Therefore, in passing laws takes time. Therefore, in passing laws takes time. Therefore, in parallel with the regulatory route, AI parallel with the regulatory route, AI parallel with the regulatory route, AI companies can and should voluntarily companies can and should voluntarily companies can and should voluntarily work together to set standards. It's a work together to set standards. It's a work together to set standards. It's a process that Daria believes will go process that Daria believes will go process that Daria believes will go better with the verifiability provided better with the verifiability provided better with the verifiability provided by permanent embedded evaluators. For by permanent embedded evaluators. For by permanent embedded evaluators. For antitrust reasons, it's helpful for the antitrust reasons, it's helpful for the antitrust reasons, it's helpful for the US government to mediate because US government to mediate because US government to mediate because otherwise this could be a collusion otherwise this could be a collusion otherwise this could be a collusion method for the companies, yada yada, you method for the companies, yada yada, you method for the companies, yada yada, you get the idea. He specifically calls out get the idea. He specifically calls out get the idea. He specifically calls out that he's most enthusiastic about pacing that he's most enthusiastic about pacing that he's most enthusiastic about pacing based on what given frontier AI systems based on what given frontier AI systems based on what given frontier AI systems can do and how safe they can observe it can do and how safe they can observe it can do and how safe they can observe it to be. For example, one possible scheme to be. For example, one possible scheme to be. For example, one possible scheme might be a series of checkpoints. If might be a series of checkpoints. If might be a series of checkpoints. If models have capability X, then they need models have capability X, then they need models have capability X, then they need to be accompanied by certifications of to be accompanied by certifications of to be accompanied by certifications of alignment properties Y and Z, such as alignment properties Y and Z, such as alignment properties Y and Z, such as some combination of evaluations, some combination of evaluations, some combination of evaluations, interpretability analysis, and audits of interpretability analysis, and audits of interpretability analysis, and audits of training environments, which demonstrate training environments, which demonstrate training environments, which demonstrate their alignment properties. In this their alignment properties. In this their alignment properties. In this example, X might be quote, "The model's example, X might be quote, "The model's example, X might be quote, "The model's capable of escaping or defeating most capable of escaping or defeating most capable of escaping or defeating most common sandbox methods." And Y might be common sandbox methods." And Y might be common sandbox methods." And Y might be whatever is required to make it very whatever is required to make it very whatever is required to make it very unlikely the model has a propensity to unlikely the model has a propensity to unlikely the model has a propensity to break out of its environment and take break out of its environment and take break out of its environment and take over a large number of computers. We over a large number of computers. We over a large number of computers. We should also consider pacing based on

  21. should also consider pacing based on should also consider pacing based on limiting the ingredients that go into limiting the ingredients that go into limiting the ingredients that go into the frontier models like training the frontier models like training the frontier models like training compute, the nature of training runs, or compute, the nature of training runs, or compute, the nature of training runs, or internal use of AI to improve AI. He internal use of AI to improve AI. He internal use of AI to improve AI. He worries that these are more gameable worries that these are more gameable worries that these are more gameable than external behaviors, but it's the than external behaviors, but it's the than external behaviors, but it's the kind of topic worth discussing with kind of topic worth discussing with kind of topic worth discussing with embedded evaluators. This part I'm more embedded evaluators. This part I'm more embedded evaluators. This part I'm more iffy on is going to be very hard to iffy on is going to be very hard to iffy on is going to be very hard to evaluate this. And as the amount of evaluate this. And as the amount of evaluate this. And as the amount of compute necessary for a given level compute necessary for a given level compute necessary for a given level intelligence goes down, this would have intelligence goes down, this would have intelligence goes down, this would have to be a weirdly moving target. Although to be a weirdly moving target. Although to be a weirdly moving target. Although it would help computer prices go down, it would help computer prices go down, it would help computer prices go down, which would be nice. He does call out which would be nice. He does call out which would be nice. He does call out that this pacing would be limited by the that this pacing would be limited by the that this pacing would be limited by the lead that US companies have over lead that US companies have over lead that US companies have over authoritarian regimes, chiefly the authoritarian regimes, chiefly the authoritarian regimes, chiefly the Chinese Communist Party. If we slow down Chinese Communist Party. If we slow down Chinese Communist Party. If we slow down by more than this amount, the unpaced by more than this amount, the unpaced by more than this amount, the unpaced CCP associated projects will pull ahead, CCP associated projects will pull ahead, CCP associated projects will pull ahead, creating significant national security creating significant national security creating significant national security risk. He calls it that he agrees with risk. He calls it that he agrees with risk. He calls it that he agrees with Secretary Basant that a Chinese lead in Secretary Basant that a Chinese lead in Secretary Basant that a Chinese lead in AI would pose grave danger for the AI would pose grave danger for the AI would pose grave danger for the United States and the world. The CCP United States and the world. The CCP United States and the world. The CCP associated projects will run the associated projects will run the associated projects will run the alignment risks that US companies are alignment risks that US companies are alignment risks that US companies are carefully preventing. And even if they carefully preventing. And even if they carefully preventing. And even if they avoid these risks, they will be in a avoid these risks, they will be in a avoid these risks, they will be in a position to militarily dominate position to militarily dominate position to militarily dominate democracies. For example, if they have democracies. For example, if they have democracies. For example, if they have AI drones. All of this is why he thinks AI drones. All of this is why he thinks AI drones. All of this is why he thinks it's important that democracies maintain it's important that democracies maintain it's important that democracies maintain a huge lead over autocratic societies a huge lead over autocratic societies a huge lead over autocratic societies and countries because those autocracies and countries because those autocracies and countries because those autocracies could destroy the world with this. So if could destroy the world with this. So if could destroy the world with this. So if they beat us there, we are screwed. He they beat us there, we are screwed. He they beat us there, we are screwed. He actually goes as far as calling out actually goes as far as calling out actually goes as far as calling out specific ways to prevent this like specific ways to prevent this like specific ways to prevent this like refusing to sell powerful AI chips and refusing to sell powerful AI chips and refusing to sell powerful AI chips and manufacturing to China as well as manufacturing to China as well as manufacturing to China as well as cracking down on chip smuggling cracking down on chip smuggling cracking down on chip smuggling operations and remote access to data operations and remote access to data operations and remote access to data centers outside of China. Chips will be centers outside of China. Chips will be centers outside of China. Chips will be the main determinant of China's AI the main determinant of China's AI the main determinant of China's AI strength. I will say there is risk here strength. I will say there is risk here strength. I will say there is risk here seeing the developments happened seeing the developments happened seeing the developments happened recently. For example, with GLM53 Flash

  22. recently. For example, with GLM53 Flash recently. For example, with GLM53 Flash being served primarily by the team who being served primarily by the team who being served primarily by the team who made it on Huawei chips. That said, it made it on Huawei chips. That said, it made it on Huawei chips. That said, it is served on those. We don't have as is served on those. We don't have as is served on those. We don't have as much detail on how it was trained and I much detail on how it was trained and I much detail on how it was trained and I would be surprised if it wasn't using would be surprised if it wasn't using would be surprised if it wasn't using Nvidia chips in training for a Nvidia chips in training for a Nvidia chips in training for a meaningful amount of that work. Another meaningful amount of that work. Another meaningful amount of that work. Another thing it was almost certainly using was thing it was almost certainly using was thing it was almost certainly using was histories from real claude sessions for histories from real claude sessions for histories from real claude sessions for distillation, which obviously is his distillation, which obviously is his distillation, which obviously is his next point. While I think a lot of the next point. While I think a lot of the next point. While I think a lot of the distillation shouting is a bit distillation shouting is a bit distillation shouting is a bit overblown, some of the examples we're overblown, some of the examples we're overblown, some of the examples we're getting now are egregious. One of the getting now are egregious. One of the getting now are egregious. One of the Chinese labs, if I recall, it was Chinese labs, if I recall, it was Chinese labs, if I recall, it was Miniax, but I could be wrong on that. I Miniax, but I could be wrong on that. I Miniax, but I could be wrong on that. I should double check, but I'm already should double check, but I'm already should double check, but I'm already over time for this, was actually serving over time for this, was actually serving over time for this, was actually serving Claude when users requested their models Claude when users requested their models Claude when users requested their models sometimes in order to get data, which is sometimes in order to get data, which is sometimes in order to get data, which is hilarious and crazy. The last thing he hilarious and crazy. The last thing he hilarious and crazy. The last thing he says we need to do to keep China from says we need to do to keep China from says we need to do to keep China from catching up is strengthen security at catching up is strengthen security at catching up is strengthen security at the AI companies and prevent model the AI companies and prevent model the AI companies and prevent model weight theft. I am surprised this hasn't weight theft. I am surprised this hasn't weight theft. I am surprised this hasn't happened yet, but also these files are happened yet, but also these files are happened yet, but also these files are gigantic. Some of these weights are gigantic. Some of these weights are gigantic. Some of these weights are many, many terabytes for these models. many, many terabytes for these models. many, many terabytes for these models. I'd be surprised if Fable was less than I'd be surprised if Fable was less than I'd be surprised if Fable was less than 10TB to get everything you need to run 10TB to get everything you need to run 10TB to get everything you need to run it somewhere else. Companies in the US it somewhere else. Companies in the US it somewhere else. Companies in the US government should cooperate to make government should cooperate to make government should cooperate to make these steps as effective as possible.

  23. these steps as effective as possible. these steps as effective as possible. Enthropic has consistently advocated for Enthropic has consistently advocated for Enthropic has consistently advocated for all of the measures because they've all of the measures because they've all of the measures because they've always understood that they would be always understood that they would be always understood that they would be essential to any pacing. So, how much essential to any pacing. So, how much essential to any pacing. So, how much lead will this give us? Dario believes lead will this give us? Dario believes lead will this give us? Dario believes that these would slow China down enough that these would slow China down enough that these would slow China down enough to significantly widen America's lead to significantly widen America's lead to significantly widen America's lead over the next 3 to 5 years, the window over the next 3 to 5 years, the window over the next 3 to 5 years, the window when AI will become geopolitically most when AI will become geopolitically most when AI will become geopolitically most important. Some may believe these important. Some may believe these important. Some may believe these measures make it more difficult to measures make it more difficult to measures make it more difficult to cooperate with China. But Daria believes cooperate with China. But Daria believes cooperate with China. But Daria believes the opposite is true. These measures the opposite is true. These measures the opposite is true. These measures increase the leverage held by increase the leverage held by increase the leverage held by democracies and they make an agreement democracies and they make an agreement democracies and they make an agreement more likely in the future. He is more likely in the future. He is more likely in the future. He is strongarmming China here. He is not strongarmming China here. He is not strongarmming China here. He is not taking it. And you know what? Good for taking it. And you know what? Good for taking it. And you know what? Good for him. And now we have the global pacing him. And now we have the global pacing him. And now we have the global pacing section. He does call out that pacing section. He does call out that pacing section. He does call out that pacing outside of democracies will be much outside of democracies will be much outside of democracies will be much harder to achieve. Global pacing will harder to achieve. Global pacing will harder to achieve. Global pacing will require cooperation with China, the require cooperation with China, the require cooperation with China, the autocratic country with by far the most autocratic country with by far the most autocratic country with by far the most advanced AI capabilities. We must not be advanced AI capabilities. We must not be advanced AI capabilities. We must not be naive here. The geopolitical stakes are naive here. The geopolitical stakes are naive here. The geopolitical stakes are so high that there will likely be stark so high that there will likely be stark so high that there will likely be stark limits on what can be achieved, limits on what can be achieved, limits on what can be achieved, especially at first. If we greatly especially at first. If we greatly especially at first. If we greatly restrain our AI capabilities in the restrain our AI capabilities in the restrain our AI capabilities in the belief that China will do the same and belief that China will do the same and belief that China will do the same and then China defects, AI could be so then China defects, AI could be so then China defects, AI could be so powerful that such a defection could powerful that such a defection could powerful that such a defection could lead to their geopolitical dominance.

  24. lead to their geopolitical dominance. lead to their geopolitical dominance. Therefore, any agreement must either Therefore, any agreement must either Therefore, any agreement must either have ironclad verifiability or must be have ironclad verifiability or must be have ironclad verifiability or must be limited enough that that defection would limited enough that that defection would limited enough that that defection would not be militarily existential. Dario not be militarily existential. Dario not be militarily existential. Dario suspects that not only the US but also suspects that not only the US but also suspects that not only the US but also China will have these concerns and China will have these concerns and China will have these concerns and anxieties. As such, we should approach anxieties. As such, we should approach anxieties. As such, we should approach any global pacing decision, especially any global pacing decision, especially any global pacing decision, especially in the near term, in a way that protects in the near term, in a way that protects in the near term, in a way that protects the lead of the US and its allies. He the lead of the US and its allies. He the lead of the US and its allies. He has different levels of agreement here has different levels of agreement here has different levels of agreement here that he thinks are worth considering. that he thinks are worth considering. that he thinks are worth considering. The first level would be agreeing to The first level would be agreeing to The first level would be agreeing to prohibit certain narrow and dangerous prohibit certain narrow and dangerous prohibit certain narrow and dangerous uses of AI like using it for the uses of AI like using it for the uses of AI like using it for the production of biological weapons or production of biological weapons or production of biological weapons or allowing users to do it. Level two is an allowing users to do it. Level two is an allowing users to do it. Level two is an agreement by both sides to test their agreement by both sides to test their agreement by both sides to test their model before release for acute risks in model before release for acute risks in model before release for acute risks in areas like cyber security, biology, and areas like cyber security, biology, and areas like cyber security, biology, and alignment. Three would be a speed limit alignment. Three would be a speed limit alignment. Three would be a speed limit on the rate of recursive on the rate of recursive on the rate of recursive self-improvement, making sure labs don't self-improvement, making sure labs don't self-improvement, making sure labs don't make models improve themselves so fast make models improve themselves so fast make models improve themselves so fast that we lose track. And four would be a that we lose track. And four would be a that we lose track. And four would be a proper full pacing, perhaps even a proper full pacing, perhaps even a proper full pacing, perhaps even a pause, in which participating pause, in which participating pause, in which participating governments agree to substantially limit governments agree to substantially limit governments agree to substantially limit the overall rate of AI development. He the overall rate of AI development. He the overall rate of AI development. He supports floating this, but he thinks supports floating this, but he thinks supports floating this, but he thinks it's unlikely to actually happen anytime it's unlikely to actually happen anytime it's unlikely to actually happen anytime soon. Specifically because you could soon. Specifically because you could soon. Specifically because you could easily defect if and avoid monitoring, easily defect if and avoid monitoring, easily defect if and avoid monitoring, which would radically shift the balance which would radically shift the balance which would radically shift the balance of global power. Any cooperation we're of global power. Any cooperation we're of global power. Any cooperation we're able to achieve with China will extend able to achieve with China will extend able to achieve with China will extend the amount of time we have to spend on the amount of time we have to spend on the amount of time we have to spend on pacing the frontier within the pacing the frontier within the pacing the frontier within the democratic nations. We should aim for democratic nations. We should aim for democratic nations. We should aim for the higher levels while seeing the lower the higher levels while seeing the lower the higher levels while seeing the lower levels as much more likely and levels as much more likely and levels as much more likely and realistic. Finally, it's important to realistic. Finally, it's important to realistic. Finally, it's important to note that even if we cannot achieve note that even if we cannot achieve note that even if we cannot achieve formal agreements, simply changing formal agreements, simply changing formal agreements, simply changing informal norms may have some value.

  25. informal norms may have some value. informal norms may have some value. Sharing information about recursive Sharing information about recursive Sharing information about recursive self-improvement and about the self-improvement and about the self-improvement and about the misalignment of models can help convince misalignment of models can help convince misalignment of models can help convince everyone that it is not in their everyone that it is not in their everyone that it is not in their interests to be reckless. He closes with interests to be reckless. He closes with interests to be reckless. He closes with the following. Daario continues to the following. Daario continues to the following. Daario continues to believe that AI can enormously improve believe that AI can enormously improve believe that AI can enormously improve the quality of human life. His desire to the quality of human life. His desire to the quality of human life. His desire to achieve these benefits is undimemed. But achieve these benefits is undimemed. But achieve these benefits is undimemed. But the benefit will only be achieved if we the benefit will only be achieved if we the benefit will only be achieved if we build the technology in the right way. build the technology in the right way. build the technology in the right way. And so long as we use the time we gain And so long as we use the time we gain And so long as we use the time we gain well, it is worth taking unusually well, it is worth taking unusually well, it is worth taking unusually deliberate care to get it right. deliberate care to get it right. deliberate care to get it right. Progress will still be relatively fast, Progress will still be relatively fast, Progress will still be relatively fast, and we can use the time to advance the and we can use the time to advance the and we can use the time to advance the science of interpretability, improve science of interpretability, improve science of interpretability, improve operational security and rigor at the operational security and rigor at the operational security and rigor at the frontier AI companies, and build models frontier AI companies, and build models frontier AI companies, and build models whose alignment we have much more whose alignment we have much more whose alignment we have much more confidence in. The measures that Daria confidence in. The measures that Daria confidence in. The measures that Daria proposes are meant to advance the proposes are meant to advance the proposes are meant to advance the frontier at a safe pace, and it won't be frontier at a safe pace, and it won't be frontier at a safe pace, and it won't be easy, but he believes we owe it to easy, but he believes we owe it to easy, but he believes we owe it to humanity to try. This was a really, humanity to try. This was a really, humanity to try. This was a really, really good essay. And as much as I like really good essay. And as much as I like really good essay. And as much as I like to pick on my friends over at Anthropic, to pick on my friends over at Anthropic, to pick on my friends over at Anthropic, and as much as I love to give crap to and as much as I love to give crap to and as much as I love to give crap to Daario, this was responsible and well Daario, this was responsible and well Daario, this was responsible and well done. It's not alarmist, it's not saying done. It's not alarmist, it's not saying done. It's not alarmist, it's not saying that the AI is going to take off and that the AI is going to take off and that the AI is going to take off and escape the GPUs and destroy the world.

  26. escape the GPUs and destroy the world. escape the GPUs and destroy the world. It's a realistic look at where things It's a realistic look at where things It's a realistic look at where things are at now, where they are probably are at now, where they are probably are at now, where they are probably going, and how we can put a little extra going, and how we can put a little extra going, and how we can put a little extra effort up front to make sure it doesn't effort up front to make sure it doesn't effort up front to make sure it doesn't get really bad. And I think it was worth get really bad. And I think it was worth get really bad. And I think it was worth listening to, and I hope that you listening to, and I hope that you listening to, and I hope that you enjoyed it. Things are going to get enjoyed it. Things are going to get enjoyed it. Things are going to get scary fast and I hope we take the time scary fast and I hope we take the time scary fast and I hope we take the time to reflect on that and do what we can to to reflect on that and do what we can to to reflect on that and do what we can to prevent it. It's important to get these prevent it. It's important to get these prevent it. It's important to get these things right because we still can things right because we still can things right because we still can reverse it if it goes wrong, but in the reverse it if it goes wrong, but in the reverse it if it goes wrong, but in the future that might not be the case. future that might not be the case. future that might not be the case. Hopefully you all enjoyed this. I aren't Hopefully you all enjoyed this. I aren't Hopefully you all enjoyed this. I aren't just calling me a paid shell on a video just calling me a paid shell on a video just calling me a paid shell on a video that I was only paid for by my sponsors. that I was only paid for by my sponsors. that I was only paid for by my sponsors. So yeah, hope you enjoyed it and until So yeah, hope you enjoyed it and until So yeah, hope you enjoyed it and until next time, peace nerds.

Summary

The main theme is the urgent need to "pace the frontier" of AI development due to escalating safety concerns, highlighted by internal dissent at Anthropic and widespread agreement from figures like Elon Musk and Sam Altman. The takeaway is that current AI progress, particularly concerning geopolitical risks with China, demands a more serious and level-headed approach from researchers and industry leaders. This shift is illustrated by Dario Amodei's new article which directly addresses these critical issues.

View original episode ↗