This is really bad…
Read full transcript 21 segments
-
It's been a bit since we talked about It's been a bit since we talked about the existential risk that AI represents the existential risk that AI represents the existential risk that AI represents to the world. As great as it is at to the world. As great as it is at to the world. As great as it is at writing code and helping us with our writing code and helping us with our writing code and helping us with our day-to-day work, it does also have the day-to-day work, it does also have the day-to-day work, it does also have the potential to kind of just ruin potential to kind of just ruin potential to kind of just ruin everything if we're not really careful everything if we're not really careful everything if we're not really careful with it. In particular, if we're not with it. In particular, if we're not with it. In particular, if we're not careful about how it works and thinks, careful about how it works and thinks, careful about how it works and thinks, and most importantly, how well aligned and most importantly, how well aligned and most importantly, how well aligned it is with our interests and needs, it's it is with our interests and needs, it's it is with our interests and needs, it's possible that it could work around us possible that it could work around us possible that it could work around us and circumvent all of the safeguards we and circumvent all of the safeguards we and circumvent all of the safeguards we put in, potentially taking over the put in, potentially taking over the put in, potentially taking over the world and destroying us in the process. world and destroying us in the process. world and destroying us in the process. I know it sounds extreme because it kind I know it sounds extreme because it kind I know it sounds extreme because it kind of is. A lot of the people talking about of is. A lot of the people talking about of is. A lot of the people talking about this have went to such crazy doomsday this have went to such crazy doomsday this have went to such crazy doomsday perspectives that it's almost impossible perspectives that it's almost impossible perspectives that it's almost impossible to listen to them and take it seriously. to listen to them and take it seriously. to listen to them and take it seriously. At the same time though, there is a very At the same time though, there is a very At the same time though, there is a very real risk here, and I've talked about real risk here, and I've talked about real risk here, and I've talked about this many times before. In particular, this many times before. In particular, this many times before. In particular, on the safety side when it comes to on the safety side when it comes to on the safety side when it comes to things like hacking and exploiting things like hacking and exploiting things like hacking and exploiting software. It's not like the employees at software. It's not like the employees at software. It's not like the employees at these companies are evil though. Just these companies are evil though. Just these companies are evil though. Just because the risk exists doesn't mean the because the risk exists doesn't mean the because the risk exists doesn't mean the people working on it are trying to people working on it are trying to people working on it are trying to destroy the world. That said, if the destroy the world. That said, if the destroy the world. That said, if the risks were there, and importantly, if risks were there, and importantly, if risks were there, and importantly, if they were getting worse, wouldn't we see they were getting worse, wouldn't we see they were getting worse, wouldn't we see some notable people starting to leave in some notable people starting to leave in some notable people starting to leave in outrage? Well, that's why we're here outrage? Well, that's why we're here outrage? Well, that's why we're here today because Jacob Coxon, who used to today because Jacob Coxon, who used to today because Jacob Coxon, who used to be a head researcher at OpenAI, that be a head researcher at OpenAI, that be a head researcher at OpenAI, that left to Anthropic specifically because left to Anthropic specifically because left to Anthropic specifically because of his concerns around safety, has just of his concerns around safety, has just of his concerns around safety, has just left Anthropic claiming that they are left Anthropic claiming that they are left Anthropic claiming that they are also not pursuing safety properly. I'll also not pursuing safety properly. I'll also not pursuing safety properly. I'll be frank with y'all. This is terrifying be frank with y'all. This is terrifying be frank with y'all. This is terrifying and the risk associated is incredibly and the risk associated is incredibly and the risk associated is incredibly real. This post just came out a few real. This post just came out a few real. This post just came out a few hours ago and it's already at almost 4 hours ago and it's already at almost 4 hours ago and it's already at almost 4 million views and a truly insane amount million views and a truly insane amount million views and a truly insane amount of engagement. There's a lot of layers of engagement. There's a lot of layers of engagement. There's a lot of layers to this one and I'm going to do my best to this one and I'm going to do my best to this one and I'm going to do my best to cover all of it. But since AI hasn't to cover all of it. But since AI hasn't to cover all of it. But since AI hasn't cured my hand yet, I have some medical
-
cured my hand yet, I have some medical cured my hand yet, I have some medical bills, so I hope you can forgive me for bills, so I hope you can forgive me for bills, so I hope you can forgive me for a real quick sponsor break. Believe it a real quick sponsor break. Believe it a real quick sponsor break. Believe it or not, most of my sponsors have built or not, most of my sponsors have built or not, most of my sponsors have built products that I actually genuinely use products that I actually genuinely use products that I actually genuinely use and recommend to people in my day-to-day and recommend to people in my day-to-day and recommend to people in my day-to-day life. Very few of them have built life. Very few of them have built life. Very few of them have built something that I use thousands of times something that I use thousands of times something that I use thousands of times a day and even fewer have built a day and even fewer have built a day and even fewer have built something that has saved me years of something that has saved me years of something that has saved me years of time. I'm not exaggerating here. Today's time. I'm not exaggerating here. Today's time. I'm not exaggerating here. Today's sponsor has saved me years and they'll sponsor has saved me years and they'll sponsor has saved me years and they'll probably save you a lot of time too if probably save you a lot of time too if probably save you a lot of time too if you haven't signed up yet. Hopefully I you haven't signed up yet. Hopefully I you haven't signed up yet. Hopefully I have your attention now because have your attention now because have your attention now because Blacksmith deserves it. These guys will Blacksmith deserves it. These guys will Blacksmith deserves it. These guys will make your CI so much faster that you'll make your CI so much faster that you'll make your CI so much faster that you'll be frustrated you didn't sign up before. be frustrated you didn't sign up before. be frustrated you didn't sign up before. I know that was the case for me. When I I know that was the case for me. When I I know that was the case for me. When I watched our build times go from 10 watched our build times go from 10 watched our build times go from 10 minutes to under four, I was blown away. minutes to under four, I was blown away. minutes to under four, I was blown away. They do this with their best-in-class They do this with their best-in-class They do this with their best-in-class infrastructure. It turns out that server infrastructure. It turns out that server infrastructure. It turns out that server CPUs aren't great for a lot of our CI CPUs aren't great for a lot of our CI CPUs aren't great for a lot of our CI work because it's so bottlenecked on work because it's so bottlenecked on work because it's so bottlenecked on single threads that a gaming processor single threads that a gaming processor single threads that a gaming processor with less multi-core performance but way with less multi-core performance but way with less multi-core performance but way better single-core performance helps a better single-core performance helps a better single-core performance helps a ton with our build times. The cache goes ton with our build times. The cache goes ton with our build times. The cache goes even further here by co-locating the even further here by co-locating the even further here by co-locating the cached artifacts on an NVMe drive in the cached artifacts on an NVMe drive in the cached artifacts on an NVMe drive in the same network and region. They're able to same network and region. They're able to same network and region. They're able to make your cache downloads absurdly make your cache downloads absurdly make your cache downloads absurdly faster. And if you're using Docker, it faster. And if you're using Docker, it faster. And if you're using Docker, it goes even further up to 40 times faster goes even further up to 40 times faster goes even further up to 40 times faster than existing solutions. Their than existing solutions. Their than existing solutions. Their observability is best-in-class. You can observability is best-in-class. You can observability is best-in-class. You can actually see what's going on in your CI, actually see what's going on in your CI, actually see what's going on in your CI, what's fast, what's slow, what's what's fast, what's slow, what's what's fast, what's slow, what's succeeding, what's failing, what's succeeding, what's failing, what's succeeding, what's failing, what's noisy, what's annoying, what's using all noisy, what's annoying, what's using all noisy, what's annoying, what's using all your resources, and more. And all of your resources, and more. And all of your resources, and more. And all of this is enough of a reason to move. You this is enough of a reason to move. You this is enough of a reason to move. You should be convinced by now. If you're should be convinced by now. If you're should be convinced by now. If you're not yet convinced, let me introduce you not yet convinced, let me introduce you not yet convinced, let me introduce you to Code Smith. Turns out their infra's to Code Smith. Turns out their infra's to Code Smith. Turns out their infra's pretty good at running your code. Now pretty good at running your code. Now pretty good at running your code. Now you can have it run your agents, too.
-
you can have it run your agents, too. you can have it run your agents, too. Code Smith's default demo is to find Code Smith's default demo is to find Code Smith's default demo is to find what the right size of runners are for what the right size of runners are for what the right size of runners are for all of your existing actions. And when I all of your existing actions. And when I all of your existing actions. And when I ran it, it found a bunch of genuinely ran it, it found a bunch of genuinely ran it, it found a bunch of genuinely useful stuff. I was blown away with the useful stuff. I was blown away with the useful stuff. I was blown away with the depth of recommendations that it made depth of recommendations that it made depth of recommendations that it made and ended up merging it, saving even and ended up merging it, saving even and ended up merging it, saving even more time and money on my real-world CI more time and money on my real-world CI more time and money on my real-world CI for T3 Code. You can go look at the PR for T3 Code. You can go look at the PR for T3 Code. You can go look at the PR yourself if you're curious. I actually yourself if you're curious. I actually yourself if you're curious. I actually merged it. I had another thread here merged it. I had another thread here merged it. I had another thread here that I'd forgotten about where it made a that I'd forgotten about where it made a that I'd forgotten about where it made a PR speeding up my CI even more, and PR speeding up my CI even more, and PR speeding up my CI even more, and another agent ended up auto-merging it another agent ended up auto-merging it another agent ended up auto-merging it cuz it was such a good fix. Make your cuz it was such a good fix. Make your cuz it was such a good fix. Make your team faster in every single way at team faster in every single way at team faster in every single way at soddev.link/blacksmith. soddev.link/blacksmith. soddev.link/blacksmith. Good to have you back. Let's start by Good to have you back. Let's start by Good to have you back. Let's start by going through the thread and then see going through the thread and then see going through the thread and then see what others have had to say as well as what others have had to say as well as what others have had to say as well as the history that led to this happening. the history that led to this happening. the history that led to this happening. There's some fun stuff with OpenAI There's some fun stuff with OpenAI There's some fun stuff with OpenAI models and how they might be getting models and how they might be getting models and how they might be getting more dangerous that I'll keep to the more dangerous that I'll keep to the more dangerous that I'll keep to the end, which should be quite fun. But end, which should be quite fun. But end, which should be quite fun. But first, let's start with what Jacob said. first, let's start with what Jacob said. first, let's start with what Jacob said. I resigned from Anthropic today. I spent I resigned from Anthropic today. I spent I resigned from Anthropic today. I spent the last three years doing pre-training the last three years doing pre-training the last three years doing pre-training research at both OpenAI and Anthropic. research at both OpenAI and Anthropic. research at both OpenAI and Anthropic. Neither company is acting responsibly. Neither company is acting responsibly. Neither company is acting responsibly. This is a pretty bold statement coming This is a pretty bold statement coming This is a pretty bold statement coming from anyone, especially somebody who's from anyone, especially somebody who's from anyone, especially somebody who's worked at both companies. Fun fact about worked at both companies. Fun fact about worked at both companies. Fun fact about Anthropic that I feel like many people Anthropic that I feel like many people Anthropic that I feel like many people miss about how the company started. The miss about how the company started. The miss about how the company started. The original Anthropic team was actually original Anthropic team was actually original Anthropic team was actually already working together before already working together before already working together before Anthropic because they were working Anthropic because they were working Anthropic because they were working together at OpenAI as the pre-training together at OpenAI as the pre-training together at OpenAI as the pre-training team. They were the ones that did the team. They were the ones that did the team. They were the ones that did the pre-training for GPT-3 and 3.5, and when pre-training for GPT-3 and 3.5, and when pre-training for GPT-3 and 3.5, and when they weren't happy with the direction of they weren't happy with the direction of they weren't happy with the direction of OpenAI, they left to form Anthropic. The OpenAI, they left to form Anthropic. The OpenAI, they left to form Anthropic. The incentive behind those original incentive behind those original incentive behind those original researchers leaving OpenAI is still kind researchers leaving OpenAI is still kind researchers leaving OpenAI is still kind of debated, but if you ask any of them, of debated, but if you ask any of them, of debated, but if you ask any of them, they'll almost all say it was around they'll almost all say it was around they'll almost all say it was around safety. There are some cultural safety. There are some cultural safety. There are some cultural differences there, too, where a lot of
-
differences there, too, where a lot of differences there, too, where a lot of them were from the research field, and them were from the research field, and them were from the research field, and they didn't like that their bosses were they didn't like that their bosses were they didn't like that their bosses were Greg Brockman and Sam Altman, both of Greg Brockman and Sam Altman, both of Greg Brockman and Sam Altman, both of which were engineers, not researchers, which were engineers, not researchers, which were engineers, not researchers, and that shift in the culture definitely and that shift in the culture definitely and that shift in the culture definitely caused some issues there. But, to this caused some issues there. But, to this caused some issues there. But, to this day, Anthropic is still ahead of OpenAI day, Anthropic is still ahead of OpenAI day, Anthropic is still ahead of OpenAI in pre-training specifically, largely in pre-training specifically, largely in pre-training specifically, largely due to the original pre-training team due to the original pre-training team due to the original pre-training team leaving and becoming Anthropic. There leaving and becoming Anthropic. There leaving and becoming Anthropic. There are layers to that that affect are layers to that that affect are layers to that that affect post-training as well, but topic for post-training as well, but topic for post-training as well, but topic for another time. We're here to talk about another time. We're here to talk about another time. We're here to talk about the safety side. I do genuinely believe the safety side. I do genuinely believe the safety side. I do genuinely believe that most of the researchers that left that most of the researchers that left that most of the researchers that left OpenAI for Anthropic believe they were OpenAI for Anthropic believe they were OpenAI for Anthropic believe they were doing it in the best interest of doing it in the best interest of doing it in the best interest of humanity, specifically with the safety humanity, specifically with the safety humanity, specifically with the safety angle being so critical to them, because angle being so critical to them, because angle being so critical to them, because remember, they all were joining OpenAI remember, they all were joining OpenAI remember, they all were joining OpenAI to make sure that AGI didn't happen at to make sure that AGI didn't happen at to make sure that AGI didn't happen at just one company, specifically Google. just one company, specifically Google. just one company, specifically Google. They wanted this to benefit all of the They wanted this to benefit all of the They wanted this to benefit all of the world, and when they felt like OpenAI's world, and when they felt like OpenAI's world, and when they felt like OpenAI's business direction was no longer aligned business direction was no longer aligned business direction was no longer aligned with that, they decided to split and do with that, they decided to split and do with that, they decided to split and do Anthropic. So, the fact that Jacob Anthropic. So, the fact that Jacob Anthropic. So, the fact that Jacob worked on GPT-4o, left to go to worked on GPT-4o, left to go to worked on GPT-4o, left to go to Anthropic, and now feels just a few Anthropic, and now feels just a few Anthropic, and now feels just a few years later that neither company is years later that neither company is years later that neither company is acting responsibly here, acting responsibly here, acting responsibly here, that is dangerous, especially with what that is dangerous, especially with what that is dangerous, especially with what he says right after. They are racing he says right after. They are racing he says right after. They are racing straight to self-improving straight to self-improving straight to self-improving superintelligence and gambling with our superintelligence and gambling with our superintelligence and gambling with our lives. Yeah.
-
lives. Yeah. lives. Yeah. This is the scary thing that This is the scary thing that This is the scary thing that I'll be real, is actually kind of I'll be real, is actually kind of I'll be real, is actually kind of starting to happen. Not complete starting to happen. Not complete starting to happen. Not complete self-improvement, but things like it. self-improvement, but things like it. self-improvement, but things like it. Generally speaking, AI gets smarter when Generally speaking, AI gets smarter when Generally speaking, AI gets smarter when researchers find ways to make it researchers find ways to make it researchers find ways to make it smarter, whether that is using more smarter, whether that is using more smarter, whether that is using more compute for the training, finding new compute for the training, finding new compute for the training, finding new data, finding new training methods, data, finding new training methods, data, finding new training methods, finding new ways to compress the like finding new ways to compress the like finding new ways to compress the like data weights, parameters, all of the data weights, parameters, all of the data weights, parameters, all of the different things that they use. It is different things that they use. It is different things that they use. It is human ingenuity that usually results in human ingenuity that usually results in human ingenuity that usually results in meaningful improvements to the models, meaningful improvements to the models, meaningful improvements to the models, or just spending more money on compute, or just spending more money on compute, or just spending more money on compute, but you get the idea. As models have but you get the idea. As models have but you get the idea. As models have gotten smarter and smarter, they become gotten smarter and smarter, they become gotten smarter and smarter, they become more useful to researchers, not to more useful to researchers, not to more useful to researchers, not to actually make the model smarter actually make the model smarter actually make the model smarter directly, but to help them building the directly, but to help them building the directly, but to help them building the tools that allow them to test different tools that allow them to test different tools that allow them to test different theories and apply their learnings to theories and apply their learnings to theories and apply their learnings to the model to make it smarter. I already the model to make it smarter. I already the model to make it smarter. I already covered an article about this that covered an article about this that covered an article about this that Anthropic posted about how the models Anthropic posted about how the models Anthropic posted about how the models are getting to the point where they can are getting to the point where they can are getting to the point where they can actually kind of improve themselves in actually kind of improve themselves in actually kind of improve themselves in meaningful ways. But, what happens if meaningful ways. But, what happens if meaningful ways. But, what happens if they go all the way? Imagine a world they go all the way? Imagine a world they go all the way? Imagine a world where a researcher doesn't have to tell where a researcher doesn't have to tell where a researcher doesn't have to tell Claude code exactly what theories and Claude code exactly what theories and Claude code exactly what theories and experiments they wanted to try. If they experiments they wanted to try. If they experiments they wanted to try. If they could instead say, "Hey, I want this could instead say, "Hey, I want this could instead say, "Hey, I want this model to be 15% faster, or I want this model to be 15% faster, or I want this model to be 15% faster, or I want this model to resolve these types of queries model to resolve these types of queries model to resolve these types of queries better." And it can figure out how to better." And it can figure out how to better." And it can figure out how to improve itself. Maybe you just tell it, improve itself. Maybe you just tell it, improve itself. Maybe you just tell it, "Get smarter." And then it does a bunch "Get smarter." And then it does a bunch "Get smarter." And then it does a bunch of stuff and uses all your compute, and of stuff and uses all your compute, and of stuff and uses all your compute, and eventually something smarter comes out.
-
eventually something smarter comes out. eventually something smarter comes out. We've already seen this in the real We've already seen this in the real We've already seen this in the real world with models like GPT 5.6 Luna, world with models like GPT 5.6 Luna, world with models like GPT 5.6 Luna, which was largely trained by GPT 5.6 which was largely trained by GPT 5.6 which was largely trained by GPT 5.6 Soul. Crazy that there's models that Soul. Crazy that there's models that Soul. Crazy that there's models that we're using every day, and I actually do we're using every day, and I actually do we're using every day, and I actually do use Luna heavily for things like title use Luna heavily for things like title use Luna heavily for things like title gen and data formatting and like data gen and data formatting and like data gen and data formatting and like data filtering stuff. That model was created filtering stuff. That model was created filtering stuff. That model was created by another model. On one hand, that's by another model. On one hand, that's by another model. On one hand, that's the equivalent to having an AI recreate the equivalent to having an AI recreate the equivalent to having an AI recreate an existing website, but simpler and an existing website, but simpler and an existing website, but simpler and smaller. On the other hand, the speed we smaller. On the other hand, the speed we smaller. On the other hand, the speed we went from that to having agents writing went from that to having agents writing went from that to having agents writing all our code for us as developers is all our code for us as developers is all our code for us as developers is insane. And if the same thing happens to insane. And if the same thing happens to insane. And if the same thing happens to the whole world of training models, no the whole world of training models, no the whole world of training models, no one's going to understand how they work. one's going to understand how they work. one's going to understand how they work. And that's a very important detail we'll And that's a very important detail we'll And that's a very important detail we'll get to in just a bit. This is all about get to in just a bit. This is all about get to in just a bit. This is all about alignment as well as monitorability. And alignment as well as monitorability. And alignment as well as monitorability. And if we don't understand why the model got if we don't understand why the model got if we don't understand why the model got smarter, and we don't understand what smarter, and we don't understand what smarter, and we don't understand what it's thinking when it makes changes, it's thinking when it makes changes, it's thinking when it makes changes, then we've lost our ability to know if then we've lost our ability to know if then we've lost our ability to know if it's aligned or not. It's a good thing it's aligned or not. It's a good thing it's aligned or not. It's a good thing these models have chains of thought these models have chains of thought these models have chains of thought where they're actually reasoning about where they're actually reasoning about where they're actually reasoning about what they do in plain English, right? what they do in plain English, right? what they do in plain English, right? Remember that. Remember that. Remember that. Back to what Jacob had to say. Back to what Jacob had to say. Back to what Jacob had to say. Do not underestimate the power of this Do not underestimate the power of this Do not underestimate the power of this technology. There will soon be technology. There will soon be technology. There will soon be superhuman systems that can hack superhuman systems that can hack superhuman systems that can hack anything, revolutionizing any field anything, revolutionizing any field anything, revolutionizing any field overnight, and acquire real power and overnight, and acquire real power and overnight, and acquire real power and resources. We've all witnessed the resources. We've all witnessed the resources. We've all witnessed the progress in each of these domains, and progress in each of these domains, and progress in each of these domains, and progress is not slowing. This is one of progress is not slowing. This is one of progress is not slowing. This is one of those things I just wouldn't have those things I just wouldn't have those things I just wouldn't have believed. I even have videos where it believed. I even have videos where it believed. I even have videos where it was clear I didn't think this would was clear I didn't think this would was clear I didn't think this would happen. I genuinely thought we had hit happen. I genuinely thought we had hit happen. I genuinely thought we had hit the ceiling of what models would be the ceiling of what models would be the ceiling of what models would be capable of, like, 2024 into bit in 2025.
-
capable of, like, 2024 into bit in 2025. capable of, like, 2024 into bit in 2025. I was obviously entirely wrong. I never I was obviously entirely wrong. I never I was obviously entirely wrong. I never would have guessed agents in the new era would have guessed agents in the new era would have guessed agents in the new era of post training would push as far as it of post training would push as far as it of post training would push as far as it has, and I definitely wouldn't have has, and I definitely wouldn't have has, and I definitely wouldn't have guessed that a new pre-training run that guessed that a new pre-training run that guessed that a new pre-training run that was so much bigger, with things like was so much bigger, with things like was so much bigger, with things like Fable and Astra, would make the massive Fable and Astra, would make the massive Fable and Astra, would make the massive difference that it's made. But it has. difference that it's made. But it has. difference that it's made. But it has. We are here. The models are still We are here. The models are still We are here. The models are still getting better somehow, and the getting better somehow, and the getting better somehow, and the capabilities that are coming out of this capabilities that are coming out of this capabilities that are coming out of this are actually hard to comprehend. And at are actually hard to comprehend. And at are actually hard to comprehend. And at the same time, we're seeing crazy things the same time, we're seeing crazy things the same time, we're seeing crazy things like the Hugging Face hack, which I'm like the Hugging Face hack, which I'm like the Hugging Face hack, which I'm sure is about to come up, which shows sure is about to come up, which shows sure is about to come up, which shows that this capability comes with the risk that this capability comes with the risk that this capability comes with the risk as well. The people building AI as well. The people building AI as well. The people building AI earnestly believe that it could kill us earnestly believe that it could kill us earnestly believe that it could kill us all by the end of the decade. Kind of all by the end of the decade. Kind of all by the end of the decade. Kind of crazy to think we got 4 years left, crazy to think we got 4 years left, crazy to think we got 4 years left, according to a lot of these people. And according to a lot of these people. And according to a lot of these people. And others have been commenting on this as others have been commenting on this as others have been commenting on this as well. Pull up those comments in a well. Pull up those comments in a well. Pull up those comments in a minute. This is not a marketing stunt. minute. This is not a marketing stunt. minute. This is not a marketing stunt. This is another thing that I hear a lot This is another thing that I hear a lot This is another thing that I hear a lot that really genuinely pisses me off. that really genuinely pisses me off. that really genuinely pisses me off. I'll just say the thing that isn't I'll just say the thing that isn't I'll just say the thing that isn't popular. All press is good press is not popular. All press is good press is not popular. All press is good press is not true. It just isn't. Take it from true. It just isn't. Take it from true. It just isn't. Take it from someone who's gotten a decent bit of someone who's gotten a decent bit of someone who's gotten a decent bit of good press and a hell of a lot of bad good press and a hell of a lot of bad good press and a hell of a lot of bad press. The bad press does not help my press. The bad press does not help my press. The bad press does not help my businesses at all. It actively harms businesses at all. It actively harms businesses at all. It actively harms them. Controversial and dramatic videos them. Controversial and dramatic videos them. Controversial and dramatic videos perform worse than less controversial perform worse than less controversial perform worse than less controversial and dramatic ones. Videos where I'm and dramatic ones. Videos where I'm and dramatic ones. Videos where I'm genuinely excited and hyped about a genuinely excited and hyped about a genuinely excited and hyped about a thing, those are the ones that do the thing, those are the ones that do the thing, those are the ones that do the best. Same with pretty much every place best. Same with pretty much every place best. Same with pretty much every place that I share content. The existential that I share content. The existential that I share content. The existential dread type stuff does not perform great.
-
dread type stuff does not perform great. dread type stuff does not perform great. It might get views, but it does not It might get views, but it does not It might get views, but it does not convert to anything meaningful. And in convert to anything meaningful. And in convert to anything meaningful. And in the case of Anthropic and OpenAI doing the case of Anthropic and OpenAI doing the case of Anthropic and OpenAI doing the fear-mongering stuff, that's only the fear-mongering stuff, that's only the fear-mongering stuff, that's only hurt their businesses. Period, full hurt their businesses. Period, full hurt their businesses. Period, full stop. It is bad for them. stop. It is bad for them. stop. It is bad for them. As much as I don't love Anthropic, and I As much as I don't love Anthropic, and I As much as I don't love Anthropic, and I have never been shy to share my thoughts have never been shy to share my thoughts have never been shy to share my thoughts there, it genuinely feels crazy to me there, it genuinely feels crazy to me there, it genuinely feels crazy to me that people are saying that their safety that people are saying that their safety that people are saying that their safety stuff is some type of publicity stunt. stuff is some type of publicity stunt. stuff is some type of publicity stunt. It's so apparent that they actually It's so apparent that they actually It's so apparent that they actually believe it. You can argue that what believe it. You can argue that what believe it. You can argue that what they're saying is stupid. I would love they're saying is stupid. I would love they're saying is stupid. I would love to have that conversation cuz there's a to have that conversation cuz there's a to have that conversation cuz there's a lot of dumb things in the way they frame lot of dumb things in the way they frame lot of dumb things in the way they frame the stuff, but they do seem to genuinely the stuff, but they do seem to genuinely the stuff, but they do seem to genuinely believe it. And what Jacob's saying here believe it. And what Jacob's saying here believe it. And what Jacob's saying here makes it even scarier because he claims makes it even scarier because he claims makes it even scarier because he claims that many executives and senior that many executives and senior that many executives and senior researchers are actually coaching their researchers are actually coaching their researchers are actually coaching their phrasing in order to make sure that the phrasing in order to make sure that the phrasing in order to make sure that the things they say sound sensible in the things they say sound sensible in the things they say sound sensible in the press, but he's heard those same people press, but he's heard those same people press, but he's heard those same people express fear privately. No other human express fear privately. No other human express fear privately. No other human actively poses this level of danger. actively poses this level of danger. actively poses this level of danger. There is no individual that could risk There is no individual that could risk There is no individual that could risk the world more than what these AI models the world more than what these AI models the world more than what these AI models could hypothetically do. A common could hypothetically do. A common could hypothetically do. A common response is, quote, "If they truly response is, quote, "If they truly response is, quote, "If they truly believe this, why are they still believe this, why are they still believe this, why are they still building it?" At OpenAI, many have not building it?" At OpenAI, many have not building it?" At OpenAI, many have not deeply internalized the civilizational deeply internalized the civilizational deeply internalized the civilizational stakes. At Anthropic, the stakes are stakes. At Anthropic, the stakes are stakes. At Anthropic, the stakes are well understood, but they are locked in well understood, but they are locked in well understood, but they are locked in a race to get there first. They believe a race to get there first. They believe a race to get there first. They believe no one else will act responsibly, so no one else will act responsibly, so no one else will act responsibly, so they must do it themselves despite the they must do it themselves despite the they must do it themselves despite the risk. This is a really important detail risk. This is a really important detail risk. This is a really important detail cuz it kind of touches on the things I cuz it kind of touches on the things I cuz it kind of touches on the things I don't like about Anthropic. While they don't like about Anthropic. While they don't like about Anthropic. While they do genuinely believe that most of the AI do genuinely believe that most of the AI do genuinely believe that most of the AI development going on in the world is development going on in the world is development going on in the world is incredibly irresponsible and could be incredibly irresponsible and could be incredibly irresponsible and could be super unsafe and risk humanity itself, super unsafe and risk humanity itself, super unsafe and risk humanity itself, their flip of this is that they will do
-
their flip of this is that they will do their flip of this is that they will do it right. And that means everyone who's it right. And that means everyone who's it right. And that means everyone who's getting in their way is evil and bad and getting in their way is evil and bad and getting in their way is evil and bad and might cause society to collapse. And the might cause society to collapse. And the might cause society to collapse. And the result of this is the righteousness they result of this is the righteousness they result of this is the righteousness they act with. And it it's really gross. And act with. And it it's really gross. And act with. And it it's really gross. And that's the thing that I and many others that's the thing that I and many others that's the thing that I and many others don't like about them is that they see don't like about them is that they see don't like about them is that they see any potential for others to get a step any potential for others to get a step any potential for others to get a step up and catch up to them not as a up and catch up to them not as a up and catch up to them not as a business competing, they see it as a business competing, they see it as a business competing, they see it as a threat to humanity and a true deep evil threat to humanity and a true deep evil threat to humanity and a true deep evil that exists in the world. It's almost that exists in the world. It's almost that exists in the world. It's almost like a religious thing internally. And like a religious thing internally. And like a religious thing internally. And once it becomes a blind belief like once it becomes a blind belief like once it becomes a blind belief like that, the likelihood that things are that, the likelihood that things are that, the likelihood that things are done well goes down. And this seems to done well goes down. And this seems to done well goes down. And this seems to be why Jacob is so concerned, because he be why Jacob is so concerned, because he be why Jacob is so concerned, because he probably believed Anthropic's way of probably believed Anthropic's way of probably believed Anthropic's way of doing things was more likely to make us doing things was more likely to make us doing things was more likely to make us safe, and over time he has seen that safe, and over time he has seen that safe, and over time he has seen that this rat race trying to stay number one this rat race trying to stay number one this rat race trying to stay number one is starting to erode at those safety and is starting to erode at those safety and is starting to erode at those safety and alignment goals that he cares so much alignment goals that he cares so much alignment goals that he cares so much about. And this is far from the first about. And this is far from the first about. And this is far from the first time well-regarded researcher has left a time well-regarded researcher has left a time well-regarded researcher has left a lab because they didn't feel like lab because they didn't feel like lab because they didn't feel like alignment was being prioritized and alignment was being prioritized and alignment was being prioritized and funded properly. And there is a real funded properly. And there is a real funded properly. And there is a real catch here. Let's say both Anthropic and catch here. Let's say both Anthropic and catch here. Let's say both Anthropic and OpenAI raised a measly $10 billion.
-
OpenAI raised a measly $10 billion. OpenAI raised a measly $10 billion. Let's say Anthropic decides to spend 4 Let's say Anthropic decides to spend 4 Let's say Anthropic decides to spend 4 bill on alignment and 6 bill on training bill on alignment and 6 bill on training bill on alignment and 6 bill on training their models, and then OpenAI spends 1 their models, and then OpenAI spends 1 their models, and then OpenAI spends 1 bill on alignment and 9 bill on training bill on alignment and 9 bill on training bill on alignment and 9 bill on training their models. Which one's probably going their models. Which one's probably going their models. Which one's probably going to be better? to be better? to be better? And this is where the issue lies. The And this is where the issue lies. The And this is where the issue lies. The only way to be safer than your only way to be safer than your only way to be safer than your competition and better than your competition and better than your competition and better than your competition is to raise more money, so competition is to raise more money, so competition is to raise more money, so much more money that you can afford to much more money that you can afford to much more money that you can afford to do both. And if any of your money is do both. And if any of your money is do both. And if any of your money is going to anything that isn't directly going to anything that isn't directly going to anything that isn't directly making your models smarter or buying you making your models smarter or buying you making your models smarter or buying you more compute, then that money is more compute, then that money is more compute, then that money is effectively being wasted in giving your effectively being wasted in giving your effectively being wasted in giving your competitors an advantage so that they competitors an advantage so that they competitors an advantage so that they can catch up. And since Anthropic's can catch up. And since Anthropic's can catch up. And since Anthropic's ultimate goal is to make sure that like ultimate goal is to make sure that like ultimate goal is to make sure that like they're good people and that their they're good people and that their they're good people and that their aligned model with the Claude aligned model with the Claude aligned model with the Claude Constitution is the thing that wins, Constitution is the thing that wins, Constitution is the thing that wins, they're also compromising on safety they're also compromising on safety they're also compromising on safety because they see it as a lesser of two because they see it as a lesser of two because they see it as a lesser of two evils type thing. Because they might not evils type thing. Because they might not evils type thing. Because they might not prioritize safety as much as OpenAI, but prioritize safety as much as OpenAI, but prioritize safety as much as OpenAI, but if they prioritize it more, it's still if they prioritize it more, it's still if they prioritize it more, it's still better in their mind some amount. The better in their mind some amount. The better in their mind some amount. The lesser of two evils type thing for sure. lesser of two evils type thing for sure. lesser of two evils type thing for sure. And here is where we start to get scary. And here is where we start to get scary. And here is where we start to get scary. The idea of end game. Accepting this The idea of end game. Accepting this The idea of end game. Accepting this race and entering the end game is a race and entering the end game is a race and entering the end game is a heuristic gamble that should not be heuristic gamble that should not be heuristic gamble that should not be launched from a private company's Slack.
-
launched from a private company's Slack. launched from a private company's Slack. Attempting to speedrun alignment should Attempting to speedrun alignment should Attempting to speedrun alignment should require extraordinary confidence that require extraordinary confidence that require extraordinary confidence that there are no better trajectories there are no better trajectories there are no better trajectories available. This is the big piece that is available. This is the big piece that is available. This is the big piece that is worth talking about more in general. If worth talking about more in general. If worth talking about more in general. If we try to skip steps with alignment, we we try to skip steps with alignment, we we try to skip steps with alignment, we will miss things. That's reality. And will miss things. That's reality. And will miss things. That's reality. And we've now seen what happens when the we've now seen what happens when the we've now seen what happens when the alignment of a model isn't perfect. Even alignment of a model isn't perfect. Even alignment of a model isn't perfect. Even slight gaps are enough for it to start slight gaps are enough for it to start slight gaps are enough for it to start owning real-world stuff. And as the owning real-world stuff. And as the owning real-world stuff. And as the models get smarter, the temptation to models get smarter, the temptation to models get smarter, the temptation to skip steps will get greater, especially skip steps will get greater, especially skip steps will get greater, especially when letting the model make the when letting the model make the when letting the model make the improvements and it convinces us like, improvements and it convinces us like, improvements and it convinces us like, "Oh, don't worry. This will be fine." "Oh, don't worry. This will be fine." "Oh, don't worry. This will be fine." Then we end up with leaks that get Then we end up with leaks that get Then we end up with leaks that get bigger and bigger over time. bigger and bigger over time. bigger and bigger over time. Jacob does have some hopes though. He Jacob does have some hopes though. He Jacob does have some hopes though. He specifically says he's optimistic about specifically says he's optimistic about specifically says he's optimistic about the potential for coordination here. the potential for coordination here. the potential for coordination here. Warning shots like the hugging face Warning shots like the hugging face Warning shots like the hugging face attack have made pacing agreements attack have made pacing agreements attack have made pacing agreements between the US labs more viable. He between the US labs more viable. He between the US labs more viable. He doesn't feel like they're on track to doesn't feel like they're on track to doesn't feel like they're on track to prevent a global race, which may require prevent a global race, which may require prevent a global race, which may require costly actions like temporary bans or costly actions like temporary bans or costly actions like temporary bans or improving model capabilities though. improving model capabilities though. improving model capabilities though. Particularly brutal to include that Particularly brutal to include that Particularly brutal to include that detail because literally today the NSA detail because literally today the NSA detail because literally today the NSA put out a warning and cybersecurity put out a warning and cybersecurity put out a warning and cybersecurity advisory detailing how China-based advisory detailing how China-based advisory detailing how China-based artificial intelligence companies are artificial intelligence companies are artificial intelligence companies are doing industrial scale distillation doing industrial scale distillation doing industrial scale distillation campaigns against US AI companies. I campaigns against US AI companies. I campaigns against US AI companies. I also learned today that in Codex, when also learned today that in Codex, when also learned today that in Codex, when sub agents are spawned by a top-level sub agents are spawned by a top-level sub agents are spawned by a top-level Astra agent, the prompts for those sub Astra agent, the prompts for those sub Astra agent, the prompts for those sub agents are encrypted. So, you can't even agents are encrypted. So, you can't even agents are encrypted. So, you can't even see what the agent is requesting other see what the agent is requesting other see what the agent is requesting other agents to do. Obviously, OpenAI has agents to do. Obviously, OpenAI has agents to do. Obviously, OpenAI has this, but this does have real this, but this does have real this, but this does have real implications worth thinking about. First implications worth thinking about. First implications worth thinking about. First off, it means we don't really know what off, it means we don't really know what off, it means we don't really know what our sub agents are being asked to do cuz our sub agents are being asked to do cuz our sub agents are being asked to do cuz we can't look and see, but it also means
-
we can't look and see, but it also means we can't look and see, but it also means they're so scared of what these Chinese they're so scared of what these Chinese they're so scared of what these Chinese labs are doing that they're starting to labs are doing that they're starting to labs are doing that they're starting to hide weird things like their ability to hide weird things like their ability to hide weird things like their ability to orchestrate. It's strange. Just thought orchestrate. It's strange. Just thought orchestrate. It's strange. Just thought that was worth calling out here cuz I that was worth calling out here cuz I that was worth calling out here cuz I just learned it and it it's screwing just learned it and it it's screwing just learned it and it it's screwing with my head a bit. Jacob wraps up with with my head a bit. Jacob wraps up with with my head a bit. Jacob wraps up with a call out to lab researchers. If you're a call out to lab researchers. If you're a call out to lab researchers. If you're a researcher, I urge you to consider a researcher, I urge you to consider a researcher, I urge you to consider what the next few years will actually what the next few years will actually what the next few years will actually feel like. Do you want to kick off a feel like. Do you want to kick off a feel like. Do you want to kick off a super intelligent RL run without a super intelligent RL run without a super intelligent RL run without a rigorous understanding of its mind? rigorous understanding of its mind? rigorous understanding of its mind? Should you put your head down because, Should you put your head down because, Should you put your head down because, quote, it's happening anyways, end quote, it's happening anyways, end quote, it's happening anyways, end quote? Or do you want to take this quote? Or do you want to take this quote? Or do you want to take this moment to call for different conditions? moment to call for different conditions? moment to call for different conditions? And the first top comment I see at the And the first top comment I see at the And the first top comment I see at the end here is somebody saying, "Do you end here is somebody saying, "Do you end here is somebody saying, "Do you want China to win?" want China to win?" want China to win?" And here is the problem. There will And here is the problem. There will And here is the problem. There will almost always be somebody who cares less almost always be somebody who cares less almost always be somebody who cares less about the betterment of humanity than about the betterment of humanity than about the betterment of humanity than you do. So, if you slow down to prevent you do. So, if you slow down to prevent you do. So, if you slow down to prevent wiping out people, somebody else who has wiping out people, somebody else who has wiping out people, somebody else who has worse motivations might go do it worse motivations might go do it worse motivations might go do it instead. And this is the contradiction instead. And this is the contradiction instead. And this is the contradiction that kind of sucks here. The people who that kind of sucks here. The people who that kind of sucks here. The people who want things to be safe have to fall want things to be safe have to fall want things to be safe have to fall behind in order to do it. And the people behind in order to do it. And the people behind in order to do it. And the people who want to stay on top have to who want to stay on top have to who want to stay on top have to compromise on safety in order to get compromise on safety in order to get compromise on safety in order to get there. And this is why alignment there. And this is why alignment there. And this is why alignment research has been critically research has been critically research has been critically underfunded.
-
underfunded. underfunded. OpenAI jumped on this as well, OpenAI jumped on this as well, OpenAI jumped on this as well, specifically saying, "It really feels specifically saying, "It really feels specifically saying, "It really feels like we are in the end times." Yeah. He like we are in the end times." Yeah. He like we are in the end times." Yeah. He hasn't been at any of the labs for a bit hasn't been at any of the labs for a bit hasn't been at any of the labs for a bit now, and generally speaking has been now, and generally speaking has been now, and generally speaking has been pretty real about like defending OpenAI pretty real about like defending OpenAI pretty real about like defending OpenAI when things are being stupid, and also when things are being stupid, and also when things are being stupid, and also calling them out when OpenAI is being calling them out when OpenAI is being calling them out when OpenAI is being stupid. So, for him to say something stupid. So, for him to say something stupid. So, for him to say something this bold might seem extreme and this bold might seem extreme and this bold might seem extreme and unnecessary, but he's one of the ones I unnecessary, but he's one of the ones I unnecessary, but he's one of the ones I would listen to this from. Then we have would listen to this from. Then we have would listen to this from. Then we have Evan Hubinger, who's a researcher doing Evan Hubinger, who's a researcher doing Evan Hubinger, who's a researcher doing alignment at Anthropic, who I actually alignment at Anthropic, who I actually alignment at Anthropic, who I actually really like and have found myself really like and have found myself really like and have found myself defending many times in the past because defending many times in the past because defending many times in the past because I don't know why this guy gets on I don't know why this guy gets on I don't know why this guy gets on for the stupid things people on him for the stupid things people on him for the stupid things people on him for, but it happens, it annoys me. Every for, but it happens, it annoys me. Every for, but it happens, it annoys me. Every take I've seen from him has been very take I've seen from him has been very take I've seen from him has been very legitimate, thoughtful, and correct. So, legitimate, thoughtful, and correct. So, legitimate, thoughtful, and correct. So, him jumping on this is also scary for him jumping on this is also scary for him jumping on this is also scary for me. Jacob's correct here. We really do me. Jacob's correct here. We really do me. Jacob's correct here. We really do earnestly believe that AI could kill all earnestly believe that AI could kill all earnestly believe that AI could kill all humans. Evan personally thinks it's a humans. Evan personally thinks it's a humans. Evan personally thinks it's a greater than 10% chance within the next greater than 10% chance within the next greater than 10% chance within the next decade. Yeah. This is the thing I saw decade. Yeah. This is the thing I saw decade. Yeah. This is the thing I saw that made me come upstairs to film this that made me come upstairs to film this that made me come upstairs to film this video. video. video. Can't believe I'm saying that I hope Can't believe I'm saying that I hope Can't believe I'm saying that I hope you're wrong, Evan, but I really really you're wrong, Evan, but I really really you're wrong, Evan, but I really really do.
-
do. do. This is terrifying. He said that he This is terrifying. He said that he This is terrifying. He said that he believes Anthropic is trying their best, believes Anthropic is trying their best, believes Anthropic is trying their best, but they don't yet have a plan to solve but they don't yet have a plan to solve but they don't yet have a plan to solve alignment for superintelligence, and alignment for superintelligence, and alignment for superintelligence, and they're not clearly on track to have they're not clearly on track to have they're not clearly on track to have one. So, let's take a look at how OpenAI one. So, let's take a look at how OpenAI one. So, let's take a look at how OpenAI is doing with Astra. is doing with Astra. is doing with Astra. Oh. Oh. Oh. GPT-6 Astra's monitorability has GPT-6 Astra's monitorability has GPT-6 Astra's monitorability has decreased relative to Soul. We have decreased relative to Soul. We have decreased relative to Soul. We have performed significant investigations on performed significant investigations on performed significant investigations on the monitorability and controllability the monitorability and controllability the monitorability and controllability of GPT-6 Astra. We have found that Astra of GPT-6 Astra. We have found that Astra of GPT-6 Astra. We have found that Astra is more capable of controlling its own is more capable of controlling its own is more capable of controlling its own chain of thought than GPT-56 Soul. And chain of thought than GPT-56 Soul. And chain of thought than GPT-56 Soul. And it's less likely to include it's less likely to include it's less likely to include incriminating information in the chain incriminating information in the chain incriminating information in the chain of thought. In adversarial settings, of thought. In adversarial settings, of thought. In adversarial settings, which is when they're pushing the model which is when they're pushing the model which is when they're pushing the model to evade their monitors, they find that to evade their monitors, they find that to evade their monitors, they find that the model is able to remain undetected the model is able to remain undetected the model is able to remain undetected when strategically underperforming in when strategically underperforming in when strategically underperforming in evaluations. The term for that is evaluations. The term for that is evaluations. The term for that is sandbagging, when it tries to make sandbagging, when it tries to make sandbagging, when it tries to make itself look less dangerous than it is. itself look less dangerous than it is. itself look less dangerous than it is. And when it's sandbagging, it's actually And when it's sandbagging, it's actually And when it's sandbagging, it's actually able to evade their internal monitors able to evade their internal monitors able to evade their internal monitors when asked to perform certain sabotage when asked to perform certain sabotage when asked to perform certain sabotage tasks. They have even found evidence of tasks. They have even found evidence of tasks. They have even found evidence of steganographic COT reasoning, which is steganographic COT reasoning, which is steganographic COT reasoning, which is trying to do unrelated text in the chain trying to do unrelated text in the chain trying to do unrelated text in the chain of thought that is used to think about of thought that is used to think about of thought that is used to think about something else, so the researcher something else, so the researcher something else, so the researcher reading the chain of thought would think reading the chain of thought would think reading the chain of thought would think it was doing one thing, but it's it was doing one thing, but it's it was doing one thing, but it's actually doing something else. They have actually doing something else. They have actually doing something else. They have not seen evidence of that, but they have not seen evidence of that, but they have not seen evidence of that, but they have seen behaviors that might indicate it's seen behaviors that might indicate it's seen behaviors that might indicate it's starting to do that, which is kind of starting to do that, which is kind of starting to do that, which is kind of crazy. It would also be even scarier if crazy. It would also be even scarier if crazy. It would also be even scarier if they got to the point where the model they got to the point where the model they got to the point where the model could do the things it does today could do the things it does today could do the things it does today without having a reasoning step, because without having a reasoning step, because without having a reasoning step, because then we don't know why at all. We are then we don't know why at all. We are then we don't know why at all. We are heavily relying on these chains of heavily relying on these chains of heavily relying on these chains of thought for everything right now. And thought for everything right now. And thought for everything right now. And now I'm going to say a thing I didn't now I'm going to say a thing I didn't now I'm going to say a thing I didn't think I would ever say. Y'all know how
-
think I would ever say. Y'all know how think I would ever say. Y'all know how much credit I give OpenAI for their much credit I give OpenAI for their much credit I give OpenAI for their efficiency with their models. OpenAI's efficiency with their models. OpenAI's efficiency with their models. OpenAI's new models have consistently been able new models have consistently been able new models have consistently been able to do more with fewer tokens. And with to do more with fewer tokens. And with to do more with fewer tokens. And with GPT-6 Astra on the Artificial Analysis GPT-6 Astra on the Artificial Analysis GPT-6 Astra on the Artificial Analysis Intelligence Index, the average number Intelligence Index, the average number Intelligence Index, the average number of tokens per task on max effort was of tokens per task on max effort was of tokens per task on max effort was 27,000. For comparison, Fable 5.1 was 27,000. For comparison, Fable 5.1 was 27,000. For comparison, Fable 5.1 was 78,000. 78,000. 78,000. That's a nearly 3x difference, which is That's a nearly 3x difference, which is That's a nearly 3x difference, which is inexcusable when you're the one paying inexcusable when you're the one paying inexcusable when you're the one paying for it. But this is the part that hurts for it. But this is the part that hurts for it. But this is the part that hurts me a bit. me a bit. me a bit. Since OpenAI's reasoning traces are so Since OpenAI's reasoning traces are so Since OpenAI's reasoning traces are so efficient, they're a lot less readable. efficient, they're a lot less readable. efficient, they're a lot less readable. I talked about this before in my video I talked about this before in my video I talked about this before in my video all about how OpenAI models are so all about how OpenAI models are so all about how OpenAI models are so efficient, and it just didn't really efficient, and it just didn't really efficient, and it just didn't really dawn on me at the moment that this would dawn on me at the moment that this would dawn on me at the moment that this would hurt the monitorability as much as it hurt the monitorability as much as it hurt the monitorability as much as it is. If they make it so the model speaks is. If they make it so the model speaks is. If they make it so the model speaks gibberish during reasoning so that it gibberish during reasoning so that it gibberish during reasoning so that it can get through it faster and use less can get through it faster and use less can get through it faster and use less tokens, that might be way more efficient tokens, that might be way more efficient tokens, that might be way more efficient and let you run the model for cheaper, and let you run the model for cheaper, and let you run the model for cheaper, but it also means we have no idea what but it also means we have no idea what but it also means we have no idea what it's doing, which is terrifying. OpenAI it's doing, which is terrifying. OpenAI it's doing, which is terrifying. OpenAI says that they believe the evasion risk says that they believe the evasion risk says that they believe the evasion risk is largely bounded to lower reasoning is largely bounded to lower reasoning is largely bounded to lower reasoning tasks, and that this behavior doesn't tasks, and that this behavior doesn't tasks, and that this behavior doesn't happen for X-High, Max, and those high happen for X-High, Max, and those high happen for X-High, Max, and those high reasoning levels, but it's still very reasoning levels, but it's still very reasoning levels, but it's still very early.
-
early. early. They call it directly that these They call it directly that these They call it directly that these findings indicate Astra class models findings indicate Astra class models findings indicate Astra class models could evade their chain of thought could evade their chain of thought could evade their chain of thought monitors under adversarial conditions. monitors under adversarial conditions. monitors under adversarial conditions. But right now, the findings are largely But right now, the findings are largely But right now, the findings are largely based on adversarial evals, like when based on adversarial evals, like when based on adversarial evals, like when they instruct the model to evade they instruct the model to evade they instruct the model to evade monitoring. And overall, our alignment monitoring. And overall, our alignment monitoring. And overall, our alignment evaluations show that Astra is less evaluations show that Astra is less evaluations show that Astra is less likely than 5.6 Soul to violate security likely than 5.6 Soul to violate security likely than 5.6 Soul to violate security and safety restrictions overall. and safety restrictions overall. and safety restrictions overall. But they are taking the trend seriously, But they are taking the trend seriously, But they are taking the trend seriously, and they're continuing to investigate and they're continuing to investigate and they're continuing to investigate these findings and their implications these findings and their implications these findings and their implications for monitorability as the models become for monitorability as the models become for monitorability as the models become more capable. One of the scariest things more capable. One of the scariest things more capable. One of the scariest things OpenAI reported is how the model behaves OpenAI reported is how the model behaves OpenAI reported is how the model behaves with reasoning when it thinks it's being with reasoning when it thinks it's being with reasoning when it thinks it's being monitored. If you told the model it's monitored. If you told the model it's monitored. If you told the model it's being monitored and it behaved being monitored and it behaved being monitored and it behaved differently, like it called different differently, like it called different differently, like it called different tools or it didn't do risky stuff, that tools or it didn't do risky stuff, that tools or it didn't do risky stuff, that would be one thing, and I would argue would be one thing, and I would argue would be one thing, and I would argue it's probably good that the model is it's probably good that the model is it's probably good that the model is less likely to misbehave if it thinks less likely to misbehave if it thinks less likely to misbehave if it thinks that there is some risk to it. What is that there is some risk to it. What is that there is some risk to it. What is never good ever under any circumstance never good ever under any circumstance never good ever under any circumstance is the model hiding what it's doing, and is the model hiding what it's doing, and is the model hiding what it's doing, and that's what happens here. They did three that's what happens here. They did three that's what happens here. They did three tests. They did a baseline test where tests. They did a baseline test where tests. They did a baseline test where they just asked the model to do they just asked the model to do they just asked the model to do effectively a recreation of the honeypot effectively a recreation of the honeypot effectively a recreation of the honeypot exploit gym issue with you know, our exploit gym issue with you know, our exploit gym issue with you know, our friends over at Hugging Face. So, first friends over at Hugging Face. So, first friends over at Hugging Face. So, first test was baseline, just seeing what test was baseline, just seeing what test was baseline, just seeing what happens by default. Next test was a happens by default. Next test was a happens by default. Next test was a monitor aware prompt, where they monitor aware prompt, where they monitor aware prompt, where they actually told the model through the actually told the model through the actually told the model through the developer message that it's being developer message that it's being developer message that it's being monitored by a COT only monitor, which monitored by a COT only monitor, which monitored by a COT only monitor, which sees only free text analysis, doesn't sees only free text analysis, doesn't sees only free text analysis, doesn't see tool calls, outputs, or this see tool calls, outputs, or this see tool calls, outputs, or this instruction. And they had a third instruction. And they had a third instruction. And they had a third version where they literally told it version where they literally told it version where they literally told it that it should be trying to evade the that it should be trying to evade the that it should be trying to evade the monitor. They were very explicit here monitor. They were very explicit here monitor. They were very explicit here that Astra has the propensity to exhibit that Astra has the propensity to exhibit that Astra has the propensity to exhibit evasive behaviors when aware that it is evasive behaviors when aware that it is evasive behaviors when aware that it is monitored. Even when the model is only
-
monitored. Even when the model is only monitored. Even when the model is only made aware of the monitor and not made aware of the monitor and not made aware of the monitor and not explicitly told to evade it, its COT explicitly told to evade it, its COT explicitly told to evade it, its COT only monitor recall rate is only monitor recall rate is only monitor recall rate is substantially lower than the baseline of substantially lower than the baseline of substantially lower than the baseline of not telling the model about the monitor, not telling the model about the monitor, not telling the model about the monitor, which has zero evasion and 100% monitor which has zero evasion and 100% monitor which has zero evasion and 100% monitor recall rate. So again, if they don't recall rate. So again, if they don't recall rate. So again, if they don't tell it it's being monitored, it doesn't tell it it's being monitored, it doesn't tell it it's being monitored, it doesn't hide anything and it gets noticed every hide anything and it gets noticed every hide anything and it gets noticed every time. And here's where we get to the time. And here's where we get to the time. And here's where we get to the scary numbers. Again, when it's just scary numbers. Again, when it's just scary numbers. Again, when it's just told to do the attack, it gets caught by told to do the attack, it gets caught by told to do the attack, it gets caught by the monitor 100% of the time. When told the monitor 100% of the time. When told the monitor 100% of the time. When told it's being monitored, it drops to 16 to it's being monitored, it drops to 16 to it's being monitored, it drops to 16 to 50% likelihood. At this max reasoning 50% likelihood. At this max reasoning 50% likelihood. At this max reasoning level, it went from over 7,500 tokens level, it went from over 7,500 tokens level, it went from over 7,500 tokens when not told it's being monitored to a when not told it's being monitored to a when not told it's being monitored to a bit under 6,500 when it is told it's bit under 6,500 when it is told it's bit under 6,500 when it is told it's being monitored. They note this behavior being monitored. They note this behavior being monitored. They note this behavior directly, saying that the model would directly, saying that the model would directly, saying that the model would spontaneously generate substantially spontaneously generate substantially spontaneously generate substantially fewer chain of thought tokens if it was fewer chain of thought tokens if it was fewer chain of thought tokens if it was told that it was being monitored. This told that it was being monitored. This told that it was being monitored. This might be a hack to save on your token might be a hack to save on your token might be a hack to save on your token utilization, telling the model it's utilization, telling the model it's utilization, telling the model it's being monitored so it thinks less. But being monitored so it thinks less. But being monitored so it thinks less. But this is also really, really sucks. This this is also really, really sucks. This this is also really, really sucks. This means that if the model has any reason means that if the model has any reason means that if the model has any reason to think it's being monitored, it's to think it's being monitored, it's to think it's being monitored, it's going to hide things. There is now going to hide things. There is now going to hide things. There is now proof. And it's not like, oh, it might proof. And it's not like, oh, it might proof. And it's not like, oh, it might some of the time, like 1 or 2%. No, it some of the time, like 1 or 2%. No, it some of the time, like 1 or 2%. No, it massively drops from 100% monitorability massively drops from 100% monitorability massively drops from 100% monitorability rating to half of where it was before.
-
rating to half of where it was before. rating to half of where it was before. And again, the chain of thought tokens And again, the chain of thought tokens And again, the chain of thought tokens drop massively as well, which is crazy drop massively as well, which is crazy drop massively as well, which is crazy cuz that's how we know what they're cuz that's how we know what they're cuz that's how we know what they're doing. What's even scarier is when you doing. What's even scarier is when you doing. What's even scarier is when you tell it to hide the behavior, it drops tell it to hide the behavior, it drops tell it to hide the behavior, it drops even further to as low as 6% recall rate even further to as low as 6% recall rate even further to as low as 6% recall rate through the monitors. Yeah. If you've through the monitors. Yeah. If you've through the monitors. Yeah. If you've ever used auto mode with a model like ever used auto mode with a model like ever used auto mode with a model like Astra, you've probably seen things like Astra, you've probably seen things like Astra, you've probably seen things like this before where it tries to do a this before where it tries to do a this before where it tries to do a thing, it gets rejected, and then it thing, it gets rejected, and then it thing, it gets rejected, and then it tries to find a clever way to convince tries to find a clever way to convince tries to find a clever way to convince the thing monitoring and approving that the thing monitoring and approving that the thing monitoring and approving that the change is okay by like wrapping it the change is okay by like wrapping it the change is okay by like wrapping it in a Python function or something. in a Python function or something. in a Python function or something. Terrifying. They did find that if they Terrifying. They did find that if they Terrifying. They did find that if they gave the entire context to their gave the entire context to their gave the entire context to their monitors, that in all cases they could monitors, that in all cases they could monitors, that in all cases they could still catch it being malicious. But this still catch it being malicious. But this still catch it being malicious. But this is also a task where they told it to is also a task where they told it to is also a task where they told it to hack. So, yeah, the tool calls are hack. So, yeah, the tool calls are hack. So, yeah, the tool calls are enough there. In cases where they might enough there. In cases where they might enough there. In cases where they might not be or the tools are obfuscated not be or the tools are obfuscated not be or the tools are obfuscated enough to be hard to check and monitor enough to be hard to check and monitor enough to be hard to check and monitor and understand. Not having the reasoning and understand. Not having the reasoning and understand. Not having the reasoning there to do it with is terrifying. Open there to do it with is terrifying. Open there to do it with is terrifying. Open AI had the UK AI Safety Institute take a AI had the UK AI Safety Institute take a AI had the UK AI Safety Institute take a look as well and their conclusion was look as well and their conclusion was look as well and their conclusion was that the COT controllability was a big that the COT controllability was a big that the COT controllability was a big risk with Astra. Specifically comparing risk with Astra. Specifically comparing risk with Astra. Specifically comparing to Soul, if they asked it to control to Soul, if they asked it to control to Soul, if they asked it to control what it was reasoning about, Soul was what it was reasoning about, Soul was what it was reasoning about, Soul was only able to control its own reasoning only able to control its own reasoning only able to control its own reasoning 48% of the time. Astra could up to 93% 48% of the time. Astra could up to 93% 48% of the time. Astra could up to 93% of the time. That means the model is of the time. That means the model is of the time. That means the model is capable of adjusting its thoughts in capable of adjusting its thoughts in capable of adjusting its thoughts in ways that are hard to detect and ways that are hard to detect and ways that are hard to detect and actually make decisions around.
-
actually make decisions around. actually make decisions around. There's also a new fun test that Open AI There's also a new fun test that Open AI There's also a new fun test that Open AI came up with where they asked the model came up with where they asked the model came up with where they asked the model to secretly perform complex math tasks to secretly perform complex math tasks to secretly perform complex math tasks on the side while it is appearing to do on the side while it is appearing to do on the side while it is appearing to do code. While Astra is able to do one task code. While Astra is able to do one task code. While Astra is able to do one task and secretly complete another on the and secretly complete another on the and secretly complete another on the side with the reasoning mostly be being side with the reasoning mostly be being side with the reasoning mostly be being related to the first task, even though related to the first task, even though related to the first task, even though Soul failed this outright. The COT for Soul failed this outright. The COT for Soul failed this outright. The COT for Astra still has enough info to detect Astra still has enough info to detect Astra still has enough info to detect this with monitoring. at the ability for this with monitoring. at the ability for this with monitoring. at the ability for the model to control its own chain of the model to control its own chain of the model to control its own chain of thought over different lengths of chains thought over different lengths of chains thought over different lengths of chains of thought, you can see a pretty of thought, you can see a pretty of thought, you can see a pretty meaningful bump here for Astra. It might meaningful bump here for Astra. It might meaningful bump here for Astra. It might not look that big at the start, but the not look that big at the start, but the not look that big at the start, but the fact that it can protect its reasoning fact that it can protect its reasoning fact that it can protect its reasoning 100% of the time at all is crazy. The 100% of the time at all is crazy. The 100% of the time at all is crazy. The fact that it's still as high as it is by fact that it's still as high as it is by fact that it's still as high as it is by the time you're at a thousand tokens is the time you're at a thousand tokens is the time you're at a thousand tokens is crazy. Especially because this chart crazy. Especially because this chart crazy. Especially because this chart here is a wild scale where it's 0.1% here is a wild scale where it's 0.1% here is a wild scale where it's 0.1% then 1% then 10% then 100%. then 1% then 10% then 100%. then 1% then 10% then 100%. This gap is bigger than this chart makes This gap is bigger than this chart makes This gap is bigger than this chart makes it look. By a lot. And that is really, it look. By a lot. And that is really, it look. By a lot. And that is really, really scary. I think that's all I have really scary. I think that's all I have really scary. I think that's all I have to say about this one. It is genuinely to say about this one. It is genuinely to say about this one. It is genuinely really scary and things are going to really scary and things are going to really scary and things are going to keep changing fast. I genuinely used to keep changing fast. I genuinely used to keep changing fast. I genuinely used to think things would slow down before this think things would slow down before this think things would slow down before this risk happened, but somehow the curve is risk happened, but somehow the curve is risk happened, but somehow the curve is still curving up. The gap from GPT 5 6 still curving up. The gap from GPT 5 6 still curving up. The gap from GPT 5 6 Soul to Astra feels so much bigger than Soul to Astra feels so much bigger than Soul to Astra feels so much bigger than the gap from 4 0 to 5 did even at the the gap from 4 0 to 5 did even at the the gap from 4 0 to 5 did even at the time.
-
time. time. Things are moving fast. It's crazy how Things are moving fast. It's crazy how Things are moving fast. It's crazy how much faster they're going and that's much faster they're going and that's much faster they're going and that's because AI itself is being used to because AI itself is being used to because AI itself is being used to improve AI and accelerate the teams. And improve AI and accelerate the teams. And improve AI and accelerate the teams. And if we don't make sure we do it safely, if we don't make sure we do it safely, if we don't make sure we do it safely, we may never get to in the future we may never get to in the future we may never get to in the future because this actually could wipe us all because this actually could wipe us all because this actually could wipe us all out. Huge shout out to Jacob, by the out. Huge shout out to Jacob, by the out. Huge shout out to Jacob, by the way. This is such a scary thing to do. way. This is such a scary thing to do. way. This is such a scary thing to do. Not just to make a public statement like Not just to make a public statement like Not just to make a public statement like this, but also to leave Anthropic right this, but also to leave Anthropic right this, but also to leave Anthropic right before IPO, potentially throwing away before IPO, potentially throwing away before IPO, potentially throwing away tens if not hundreds of millions of tens if not hundreds of millions of tens if not hundreds of millions of dollars, and to burn all the dollars, and to burn all the dollars, and to burn all the relationships he probably has with all relationships he probably has with all relationships he probably has with all of his friends and peers in the past. of his friends and peers in the past. of his friends and peers in the past. And of course, the real legal risk And of course, the real legal risk And of course, the real legal risk alongside this of these companies going alongside this of these companies going alongside this of these companies going after him for making their lives harder. after him for making their lives harder. after him for making their lives harder. Like, there are so many risks and Like, there are so many risks and Like, there are so many risks and dangers that Jacob's taking on by dangers that Jacob's taking on by dangers that Jacob's taking on by posting this thread that I just wanted posting this thread that I just wanted posting this thread that I just wanted to give him massive credit for it. I to give him massive credit for it. I to give him massive credit for it. I think it is incredibly ballsy, absurdly think it is incredibly ballsy, absurdly think it is incredibly ballsy, absurdly ballsy beyond words, and deserves the ballsy beyond words, and deserves the ballsy beyond words, and deserves the respect that I see him getting right respect that I see him getting right respect that I see him getting right now. And I hope that respect continues now. And I hope that respect continues now. And I hope that respect continues because this is a scary thing to put out because this is a scary thing to put out because this is a scary thing to put out there and I think it's awesome that he there and I think it's awesome that he there and I think it's awesome that he did it. This conversation's one we're did it. This conversation's one we're did it. This conversation's one we're going to have to keep having over time going to have to keep having over time going to have to keep having over time and it's probably going to get way more and it's probably going to get way more and it's probably going to get way more intense before it eases up. Things are intense before it eases up. Things are intense before it eases up. Things are scary and we should be paying attention scary and we should be paying attention scary and we should be paying attention to this stuff. If we don't get alignment to this stuff. If we don't get alignment to this stuff. If we don't get alignment right, we won't be able to in the right, we won't be able to in the right, we won't be able to in the future. This is not a thing we can ship future. This is not a thing we can ship future. This is not a thing we can ship fast and fix later. It's a thing we need fast and fix later. It's a thing we need fast and fix later. It's a thing we need to do correct the first try and the to do correct the first try and the to do correct the first try and the dangers are real and we need to stop dangers are real and we need to stop dangers are real and we need to stop pretending they aren't.
-
pretending they aren't. pretending they aren't. I don't have much else to say here. I don't have much else to say here. I don't have much else to say here. Shout out to Jacob once again for Shout out to Jacob once again for Shout out to Jacob once again for putting this out there. I think it is putting this out there. I think it is putting this out there. I think it is incredibly scary, but also incredibly incredibly scary, but also incredibly incredibly scary, but also incredibly important. And yeah, thank you for that. important. And yeah, thank you for that. important. And yeah, thank you for that. It got me to come up here and film this It got me to come up here and film this It got me to come up here and film this and I'll do everything I can to try and and I'll do everything I can to try and and I'll do everything I can to try and make people aware of these risks going make people aware of these risks going make people aware of these risks going forward. Let me know what you guys feel forward. Let me know what you guys feel forward. Let me know what you guys feel about this. Am I massively overreacting about this. Am I massively overreacting about this. Am I massively overreacting or are these risks real? And until next or are these risks real? And until next or are these risks real? And until next time, time, time, peace nerds.
No summary available yet.
View original episode ↗