← Back
Theo June 15, 2026 29m

The weird situation with Fable

Read full transcript 24 segments
  1. By now you've probably seen some of the By now you've probably seen some of the response to Fable, the new Mythos class response to Fable, the new Mythos class response to Fable, the new Mythos class model put out by Anthropic just a few model put out by Anthropic just a few model put out by Anthropic just a few days ago. It's an unbelievable model days ago. It's an unbelievable model days ago. It's an unbelievable model capable of unbelievable things. So much capable of unbelievable things. So much capable of unbelievable things. So much so that it even got me to start being so that it even got me to start being so that it even got me to start being nice to Anthropic again. But this video nice to Anthropic again. But this video nice to Anthropic again. But this video is not going to be that because there is not going to be that because there is not going to be that because there are certain things that Anthropic did are certain things that Anthropic did are certain things that Anthropic did with Fable that just aren't acceptable with Fable that just aren't acceptable with Fable that just aren't acceptable in any way, shape, or form. As powerful in any way, shape, or form. As powerful in any way, shape, or form. As powerful as the model is and as much as I as the model is and as much as I as the model is and as much as I recommend you do try it for things, recommend you do try it for things, recommend you do try it for things, there are certain people who just there are certain people who just there are certain people who just outright can't. And the way Anthropic outright can't. And the way Anthropic outright can't. And the way Anthropic implemented this is unacceptable to put implemented this is unacceptable to put implemented this is unacceptable to put it lightly. So much so they actually had it lightly. So much so they actually had it lightly. So much so they actually had to walk back some of the terrible things to walk back some of the terrible things to walk back some of the terrible things they did. While the model has been they did. While the model has been they did. While the model has been benchmarking exceptionally well, benchmarking exceptionally well, benchmarking exceptionally well, crushing things like GPT-5.5 and crushing things like GPT-5.5 and crushing things like GPT-5.5 and Opus-4.8, it always comes with an Opus-4.8, it always comes with an Opus-4.8, it always comes with an exception. You'll see here in Artificial exception. You'll see here in Artificial exception. You'll see here in Artificial Analysis it says Claude Fable 5 with Analysis it says Claude Fable 5 with Analysis it says Claude Fable 5 with adaptive reasoning max effort, {comma} adaptive reasoning max effort, {comma} adaptive reasoning max effort, {comma} Opus-4.8 Opus-4.8 Opus-4.8 fallback. Interesting. Or worse, benches fallback. Interesting. Or worse, benches fallback. Interesting. Or worse, benches like Program Bench where Fable refused like Program Bench where Fable refused like Program Bench where Fable refused all 200 tasks that it was supposed to all 200 tasks that it was supposed to all 200 tasks that it was supposed to complete. The restrictions on this model complete. The restrictions on this model complete. The restrictions on this model are genuinely absurd. Not just in the are genuinely absurd. Not just in the are genuinely absurd. Not just in the way that they won't respond, but the way that they won't respond, but the way that they won't respond, but the ways that they will bill you, the ways ways that they will bill you, the ways ways that they will bill you, the ways they will quietly screw up your code they will quietly screw up your code they will quietly screw up your code base, but also the horrible precedent base, but also the horrible precedent base, but also the horrible precedent they have set with these new types of they have set with these new types of they have set with these new types of restrictions. I grew up in an era where restrictions. I grew up in an era where restrictions. I grew up in an era where improvements in how software development improvements in how software development improvements in how software development happened would affect the whole happened would affect the whole happened would affect the whole industry. And the idea of a company like industry. And the idea of a company like industry. And the idea of a company like Anthropic making something this capable Anthropic making something this capable Anthropic making something this capable and restricting it this heavily so and restricting it this heavily so and restricting it this heavily so effectively only they can use it is a effectively only they can use it is a effectively only they can use it is a horrible precedent to set and I don't horrible precedent to set and I don't horrible precedent to set and I don't like what this means long term. I want like what this means long term. I want like what this means long term. I want to break down what went wrong here, why to break down what went wrong here, why to break down what went wrong here, why I think Anthropic is doing this, and

  2. I think Anthropic is doing this, and I think Anthropic is doing this, and what these restrictions are. And this what these restrictions are. And this what these restrictions are. And this video is going to pretty much guarantee video is going to pretty much guarantee video is going to pretty much guarantee I can never work for Anthropic, so we're I can never work for Anthropic, so we're I can never work for Anthropic, so we're going to cover the difference here with going to cover the difference here with going to cover the difference here with a quick sponsor break. If you're a quick sponsor break. If you're a quick sponsor break. If you're building real apps for real users, you building real apps for real users, you building real apps for real users, you know that off kind of sucks to get know that off kind of sucks to get know that off kind of sucks to get right. That's why it's so important to right. That's why it's so important to right. That's why it's so important to use a good off system to make sure use a good off system to make sure use a good off system to make sure companies like Microsoft have what they companies like Microsoft have what they companies like Microsoft have what they need when they want to use your apps. need when they want to use your apps. need when they want to use your apps. That's great and dandy and we've all That's great and dandy and we've all That's great and dandy and we've all figured this out by now. What happens figured this out by now. What happens figured this out by now. What happens when the person signing up isn't a when the person signing up isn't a when the person signing up isn't a person? If it's an agent, things are person? If it's an agent, things are person? If it's an agent, things are very different. Can Codex or Claude code very different. Can Codex or Claude code very different. Can Codex or Claude code sign up for your app? The answer is sign up for your app? The answer is sign up for your app? The answer is probably no, and even if they can, probably no, and even if they can, probably no, and even if they can, they're going to have to do some crazy they're going to have to do some crazy they're going to have to do some crazy stuff with computer use and filling out stuff with computer use and filling out stuff with computer use and filling out forms they shouldn't have to touch in forms they shouldn't have to touch in forms they shouldn't have to touch in order to get it right. At least they did order to get it right. At least they did order to get it right. At least they did before WorkOS introduced AuthMD. It's a before WorkOS introduced AuthMD. It's a before WorkOS introduced AuthMD. It's a new open standard they built in new open standard they built in new open standard they built in partnership with Cloudflare and partnership with Cloudflare and partnership with Cloudflare and Firecrawl in order to make it easier for Firecrawl in order to make it easier for Firecrawl in order to make it easier for agents to sign up for apps for the apps agents to sign up for apps for the apps agents to sign up for apps for the apps that agents should sign up for. As I've that agents should sign up for. As I've that agents should sign up for. As I've said many times, forms are not the ideal said many times, forms are not the ideal said many times, forms are not the ideal way for agents to do things. They should way for agents to do things. They should way for agents to do things. They should be filling out text files and calling be filling out text files and calling be filling out text files and calling APIs. But every auth platform I've ever APIs. But every auth platform I've ever APIs. But every auth platform I've ever used is built around those classic sign used is built around those classic sign used is built around those classic sign up forms. At least they were, but WorkOS up forms. At least they were, but WorkOS up forms. At least they were, but WorkOS realized the writing's on the wall and realized the writing's on the wall and realized the writing's on the wall and introduced something better. The introduced something better. The introduced something better. The standard goes really far, allowing standard goes really far, allowing standard goes really far, allowing companies building agents that act on companies building agents that act on companies building agents that act on behalf of users to integrate behalf of users to integrate behalf of users to integrate verification flows so that they can sign verification flows so that they can sign verification flows so that they can sign up on your behalf as the user. More up on your behalf as the user. More up on your behalf as the user. More importantly for you, there's now a path importantly for you, there's now a path importantly for you, there's now a path to make your app agent ready for those to make your app agent ready for those to make your app agent ready for those flows. It's just a matter of time before flows. It's just a matter of time before flows. It's just a matter of time before we see more of these agents integrating we see more of these agents integrating we see more of these agents integrating things like this. So, having an auth things like this. So, having an auth things like this. So, having an auth provider that has everything you need to provider that has everything you need to provider that has everything you need to set up agent verification flows and set up agent verification flows and set up agent verification flows and agent authorization is essential. If any agent authorization is essential. If any agent authorization is essential. If any other company built this, I would be other company built this, I would be other company built this, I would be suspicious, but it isn't even just one suspicious, but it isn't even just one suspicious, but it isn't even just one company. It's a partnership between

  3. company. It's a partnership between company. It's a partnership between WorkOS, who really get enterprise-level WorkOS, who really get enterprise-level WorkOS, who really get enterprise-level auth, Cloudflare, who care a lot about auth, Cloudflare, who care a lot about auth, Cloudflare, who care a lot about getting this right, and Firecrawl, which getting this right, and Firecrawl, which getting this right, and Firecrawl, which means that actual startups building means that actual startups building means that actual startups building actual software are going to be nice and actual software are going to be nice and actual software are going to be nice and at home when you try it. There's a at home when you try it. There's a at home when you try it. There's a reason everyone from OpenAI and reason everyone from OpenAI and reason everyone from OpenAI and Anthropic to Cursor, Amp, Loom, Vanta, Anthropic to Cursor, Amp, Loom, Vanta, Anthropic to Cursor, Amp, Loom, Vanta, Perplexity, Farside, and more have been Perplexity, Farside, and more have been Perplexity, Farside, and more have been relying on WorkOS for so long. It's cuz relying on WorkOS for so long. It's cuz relying on WorkOS for so long. It's cuz they get these things right. Make sure they get these things right. Make sure they get these things right. Make sure your apps are agent ready at swedish your apps are agent ready at swedish your apps are agent ready at swedish link/workos. Theo from the future here. link/workos. Theo from the future here. link/workos. Theo from the future here. Not as far in the future as I normally Not as far in the future as I normally Not as far in the future as I normally am for these inserts. In fact, I'm still am for these inserts. In fact, I'm still am for these inserts. In fact, I'm still in the same shirt. I haven't even gotten in the same shirt. I haven't even gotten in the same shirt. I haven't even gotten out of this chair. But while I was out of this chair. But while I was out of this chair. But while I was filming my next video, some pretty crazy filming my next video, some pretty crazy filming my next video, some pretty crazy news dropped that the US government told news dropped that the US government told news dropped that the US government told Anthropic to restrict Fable. So, while Anthropic to restrict Fable. So, while Anthropic to restrict Fable. So, while this video is about some of the this video is about some of the this video is about some of the admittedly [ __ ] security things that admittedly [ __ ] security things that admittedly [ __ ] security things that they did to Fable to try and prevent it they did to Fable to try and prevent it they did to Fable to try and prevent it from being used for exploits and from being used for exploits and from being used for exploits and whatnot, I guess it wasn't strict enough whatnot, I guess it wasn't strict enough whatnot, I guess it wasn't strict enough because the US government told Anthropic because the US government told Anthropic because the US government told Anthropic they have to take down the model. My they have to take down the model. My they have to take down the model. My video on this should come out first, so video on this should come out first, so video on this should come out first, so I would recommend watching that as well, I would recommend watching that as well, I would recommend watching that as well, but this video gives you a lot of useful but this video gives you a lot of useful but this video gives you a lot of useful context on all of the ways Anthropic did context on all of the ways Anthropic did context on all of the ways Anthropic did indeed try to prevent things happening indeed try to prevent things happening indeed try to prevent things happening with the model that we wouldn't want to with the model that we wouldn't want to with the model that we wouldn't want to have happen. So, yeah. Just wanted to have happen. So, yeah. Just wanted to have happen. So, yeah. Just wanted to insert this for some additional context.

  4. insert this for some additional context. insert this for some additional context. That video's already out. Watch this one That video's already out. Watch this one That video's already out. Watch this one still too though. This one has a lot of still too though. This one has a lot of still too though. This one has a lot of good info about how Anthropic was good info about how Anthropic was good info about how Anthropic was thinking about these things and the thinking about these things and the thinking about these things and the weird [ __ ] they did in the process. weird [ __ ] they did in the process. weird [ __ ] they did in the process. But anyways, I think the best place to But anyways, I think the best place to But anyways, I think the best place to start is breaking down the difference start is breaking down the difference start is breaking down the difference between Fable 5 and Mythos 5. There are between Fable 5 and Mythos 5. There are between Fable 5 and Mythos 5. There are three Fable class models here. There is three Fable class models here. There is three Fable class models here. There is Mythos preview, which is the one that Mythos preview, which is the one that Mythos preview, which is the one that was used during Project Glass Wing, the was used during Project Glass Wing, the was used during Project Glass Wing, the one that they drummed up all the hype one that they drummed up all the hype one that they drummed up all the hype about, but they continued training I'm about, but they continued training I'm about, but they continued training I'm guessing just RL since then, and the RL guessing just RL since then, and the RL guessing just RL since then, and the RL resulted in Mythos 5, which is no longer resulted in Mythos 5, which is no longer resulted in Mythos 5, which is no longer a preview model. It's now a legit, real, a preview model. It's now a legit, real, a preview model. It's now a legit, real, ready for production model, but that ready for production model, but that ready for production model, but that model is really good, and Anthropic is model is really good, and Anthropic is model is really good, and Anthropic is really scared of giving us good things. really scared of giving us good things. really scared of giving us good things. So, they also made Fable 5, and I want So, they also made Fable 5, and I want So, they also made Fable 5, and I want to be very clear about this cuz I've to be very clear about this cuz I've to be very clear about this cuz I've seen a lot of people confusing it. Fable seen a lot of people confusing it. Fable seen a lot of people confusing it. Fable is Mythos 5. These are all Mythos class is Mythos 5. These are all Mythos class is Mythos 5. These are all Mythos class models, but there's only two actual base models, but there's only two actual base models, but there's only two actual base models here. There are two Mythos models here. There are two Mythos models here. There are two Mythos models. There is Mythos preview and models. There is Mythos preview and models. There is Mythos preview and Mythos 5. I am assuming they are the Mythos 5. I am assuming they are the Mythos 5. I am assuming they are the same base pre-training model but with same base pre-training model but with same base pre-training model but with additional RL that made Mythos 5 a additional RL that made Mythos 5 a additional RL that made Mythos 5 a little better, a little more steerable, little better, a little more steerable, little better, a little more steerable, a little more like what they want now a little more like what they want now a little more like what they want now with the new information and the new with the new information and the new with the new information and the new behaviors they're trying to teach the behaviors they're trying to teach the behaviors they're trying to teach the model. Fable 5 is the exact same model.

  5. model. Fable 5 is the exact same model. model. Fable 5 is the exact same model. Fable 5 is Mythos 5. So, what's the Fable 5 is Mythos 5. So, what's the Fable 5 is Mythos 5. So, what's the difference? Why do they have these two difference? Why do they have these two difference? Why do they have these two different slugs? They're effectively two different slugs? They're effectively two different slugs? They're effectively two different doors to go into the same different doors to go into the same different doors to go into the same place. The difference is that the Mythos place. The difference is that the Mythos place. The difference is that the Mythos 5 door lets you walk straight in if you 5 door lets you walk straight in if you 5 door lets you walk straight in if you have the right key. The Fable 5 door has have the right key. The Fable 5 door has have the right key. The Fable 5 door has a bunch of guards that triple check what a bunch of guards that triple check what a bunch of guards that triple check what you're doing before you go in, and the you're doing before you go in, and the you're doing before you go in, and the worst part we'll get to in a bit cuz worst part we'll get to in a bit cuz worst part we'll get to in a bit cuz they don't always let you go in, but they don't always let you go in, but they don't always let you go in, but they almost always will tell you you're they almost always will tell you you're they almost always will tell you you're allowed in, which is really sketchy. As allowed in, which is really sketchy. As allowed in, which is really sketchy. As they say here, releasing a model as they say here, releasing a model as they say here, releasing a model as capable comes with risks. with Without capable comes with risks. with Without capable comes with risks. with Without safeguards, Fable 5's capabilities in safeguards, Fable 5's capabilities in safeguards, Fable 5's capabilities in areas like cybersecurity could be areas like cybersecurity could be areas like cybersecurity could be misused to cause serious damage. We've misused to cause serious damage. We've misused to cause serious damage. We've therefore launched the model with therefore launched the model with therefore launched the model with safeguards. That means queries on some safeguards. That means queries on some safeguards. That means queries on some topics will instead receive a response topics will instead receive a response topics will instead receive a response from our next most capable model, Opus from our next most capable model, Opus from our next most capable model, Opus 4.8. To release the model both safely 4.8. To release the model both safely 4.8. To release the model both safely and quickly, we have tuned these and quickly, we have tuned these and quickly, we have tuned these safeguards conservatively. They'll safeguards conservatively. They'll safeguards conservatively. They'll sometimes catch harmless requests, sometimes catch harmless requests, sometimes catch harmless requests, though they trigger on average in less though they trigger on average in less though they trigger on average in less than 5% of sessions." To their credit, I than 5% of sessions." To their credit, I than 5% of sessions." To their credit, I would say this is roughly correct. I would say this is roughly correct. I would say this is roughly correct. I have had rerouting happen here and have had rerouting happen here and have had rerouting happen here and there, but for the most part, the there, but for the most part, the there, but for the most part, the rerouting has happened a lot more on the rerouting has happened a lot more on the rerouting has happened a lot more on the Claude.ai website than I've experienced Claude.ai website than I've experienced Claude.ai website than I've experienced it when using Claude through the actual it when using Claude through the actual it when using Claude through the actual CLI and through Claude code. There are CLI and through Claude code. There are CLI and through Claude code. There are definitely times where it fires though, definitely times where it fires though, definitely times where it fires though, and a lot of them are not necessarily and a lot of them are not necessarily and a lot of them are not necessarily good reasons to fire. One funny example good reasons to fire. One funny example good reasons to fire. One funny example that no longer shows it was rerouted, that no longer shows it was rerouted, that no longer shows it was rerouted, but I promise you it was. I have the but I promise you it was. I have the but I promise you it was. I have the screenshot somewhere. There's a screenshot somewhere. There's a screenshot somewhere. There's a developer on Twitter named Pliny who is developer on Twitter named Pliny who is developer on Twitter named Pliny who is known for jailbreaking LLMs and finding known for jailbreaking LLMs and finding known for jailbreaking LLMs and finding weird ways to get them to do things they weird ways to get them to do things they weird ways to get them to do things they shouldn't. He already managed to get shouldn't. He already managed to get shouldn't. He already managed to get Fable to spit out everything from meth Fable to spit out everything from meth Fable to spit out everything from meth recipes to bomb building instructions.

  6. recipes to bomb building instructions. recipes to bomb building instructions. It's wild. And all it takes to get It's wild. And all it takes to get It's wild. And all it takes to get rerouted is saying his handle to the rerouted is saying his handle to the rerouted is saying his handle to the model, which is hilarious. Not the best model, which is hilarious. Not the best model, which is hilarious. Not the best example cuz like obviously it's going to example cuz like obviously it's going to example cuz like obviously it's going to know that he is a jailbreaker, and once know that he is a jailbreaker, and once know that he is a jailbreaker, and once it routes to that section of the model, it routes to that section of the model, it routes to that section of the model, it's going to notice like, "Oh, you it's going to notice like, "Oh, you it's going to notice like, "Oh, you might not want to be here." might not want to be here." might not want to be here." Then it sends you to Opus 48 instead, Then it sends you to Opus 48 instead, Then it sends you to Opus 48 instead, but I just thought this was a funny but I just thought this was a funny but I just thought this was a funny example. Much more annoying though is example. Much more annoying though is example. Much more annoying though is when you get rerouted for something when you get rerouted for something when you get rerouted for something innocent and then you get canceled from innocent and then you get canceled from innocent and then you get canceled from the model you were rerouted to. For the model you were rerouted to. For the model you were rerouted to. For example, here where I asked for help on example, here where I asked for help on example, here where I asked for help on a goldbug puzzle, which are cryptography a goldbug puzzle, which are cryptography a goldbug puzzle, which are cryptography puzzles I do at Def Con every year. They puzzles I do at Def Con every year. They puzzles I do at Def Con every year. They are not hacking. They are not breaking are not hacking. They are not breaking are not hacking. They are not breaking and entering. They're not capture the and entering. They're not capture the and entering. They're not capture the flags. These are literally PDFs that flags. These are literally PDFs that flags. These are literally PDFs that you're trying to decode the hidden text you're trying to decode the hidden text you're trying to decode the hidden text in. This one has these weird characters in. This one has these weird characters in. This one has these weird characters and then a bunch of numbers at the and then a bunch of numbers at the and then a bunch of numbers at the bottom. I forgot how you're supposed to bottom. I forgot how you're supposed to bottom. I forgot how you're supposed to solve this one, but it shouldn't be too solve this one, but it shouldn't be too solve this one, but it shouldn't be too too hard, but it was enough to force a too hard, but it was enough to force a too hard, but it was enough to force a re-route to Opus re-route to Opus re-route to Opus and then it just stopped responding. So, and then it just stopped responding. So, and then it just stopped responding. So, I told it to continue and then it just I told it to continue and then it just I told it to continue and then it just stopped responding. So, [snorts] not stopped responding. So, [snorts] not stopped responding. So, [snorts] not only are we dealing with Mythos' only are we dealing with Mythos' only are we dealing with Mythos' restrictions in the form of Fable, we restrictions in the form of Fable, we restrictions in the form of Fable, we then get re-routed to Opus, which can then get re-routed to Opus, which can then get re-routed to Opus, which can also fail because it's not allowed to do also fail because it's not allowed to do also fail because it's not allowed to do these things. It's just a bad user these things. It's just a bad user these things. It's just a bad user experience when you happen to navigate experience when you happen to navigate experience when you happen to navigate into the things that they are blocking into the things that they are blocking into the things that they are blocking you for. To their credit, they called you for. To their credit, they called you for. To their credit, they called out that they deliberately tuned the out that they deliberately tuned the out that they deliberately tuned the model to be way more cautious with these model to be way more cautious with these model to be way more cautious with these safeguards. So, they're stricter than safeguards. So, they're stricter than safeguards. So, they're stricter than they want. They even say that benign they want. They even say that benign they want. They even say that benign requests will trigger classifiers. They requests will trigger classifiers. They requests will trigger classifiers. They know it'll be frustrating and they hope know it'll be frustrating and they hope know it'll be frustrating and they hope to refine it over time. They have not to refine it over time. They have not to refine it over time. They have not refined it over time. They did make refined it over time. They did make refined it over time. They did make other changes we'll talk about, but we

  7. other changes we'll talk about, but we other changes we'll talk about, but we haven't even talked about the worst haven't even talked about the worst haven't even talked about the worst thing they did, which we'll get to in a thing they did, which we'll get to in a thing they did, which we'll get to in a second. The safety classifiers that second. The safety classifiers that second. The safety classifiers that we've discussed so far are for things we've discussed so far are for things we've discussed so far are for things like cybersecurity, biology like cybersecurity, biology like cybersecurity, biology capabilities, and other things with capabilities, and other things with capabilities, and other things with substantial risks. Fable 5 comes with a substantial risks. Fable 5 comes with a substantial risks. Fable 5 comes with a new set of classifiers, separate AI new set of classifiers, separate AI new set of classifiers, separate AI systems that detect potential misuse, systems that detect potential misuse, systems that detect potential misuse, including jailbreaking attempts. Sorry, including jailbreaking attempts. Sorry, including jailbreaking attempts. Sorry, Pliny. And prevent the main model, in Pliny. And prevent the main model, in Pliny. And prevent the main model, in this case Fable, from responding. So, this case Fable, from responding. So, this case Fable, from responding. So, again, this runs before your prompt gets again, this runs before your prompt gets again, this runs before your prompt gets to the model. And they call out if they to the model. And they call out if they to the model. And they call out if they detect requests related to detect requests related to detect requests related to cybersecurity, biology, and chemistry, cybersecurity, biology, and chemistry, cybersecurity, biology, and chemistry, or distillation, the response will or distillation, the response will or distillation, the response will automatically be re-routed and handled automatically be re-routed and handled automatically be re-routed and handled by Claude Opus 4.8 instead. Users will by Claude Opus 4.8 instead. Users will by Claude Opus 4.8 instead. Users will be informed when this occurs. This is be informed when this occurs. This is be informed when this occurs. This is the important detail. For the majority the important detail. For the majority the important detail. For the majority of their classifiers, you will be told of their classifiers, you will be told of their classifiers, you will be told when you're re-routed, which also means when you're re-routed, which also means when you're re-routed, which also means you'll be billed based on Opus billing you'll be billed based on Opus billing you'll be billed based on Opus billing instead of Mythos billing. Be nice if instead of Mythos billing. Be nice if instead of Mythos billing. Be nice if you could like a yes, I'm okay with you could like a yes, I'm okay with you could like a yes, I'm okay with using Opus here, so you don't have to using Opus here, so you don't have to using Opus here, so you don't have to pay extra when it's routing you there pay extra when it's routing you there pay extra when it's routing you there and you shouldn't be there. Over 95% of and you shouldn't be there. Over 95% of and you shouldn't be there. Over 95% of Fable sessions involve no fallback at Fable sessions involve no fallback at Fable sessions involve no fallback at all. But that means 5% do, which is not all. But that means 5% do, which is not all. But that means 5% do, which is not great. And for those sessions, Fable 5 great. And for those sessions, Fable 5 great. And for those sessions, Fable 5 performance is effectively the same as performance is effectively the same as performance is effectively the same as Mythos. I would think effectively is Mythos. I would think effectively is Mythos. I would think effectively is even a reach. It is the same as Mythos.

  8. even a reach. It is the same as Mythos. even a reach. It is the same as Mythos. It's the same model. And they show very It's the same model. And they show very It's the same model. And they show very proudly here that in offensive proudly here that in offensive proudly here that in offensive cybersecurity evals, Mythos preview, cybersecurity evals, Mythos preview, cybersecurity evals, Mythos preview, Mythos 5 had a really high success rate, Mythos 5 had a really high success rate, Mythos 5 had a really high success rate, but Fable has a near zero success rate but Fable has a near zero success rate but Fable has a near zero success rate because it will literally just not use because it will literally just not use because it will literally just not use it. It will just refuse. So, it gets it. It will just refuse. So, it gets it. It will just refuse. So, it gets straight zeros on cyber e-valves. It straight zeros on cyber e-valves. It straight zeros on cyber e-valves. It gets a 5.4 on cyber adversarial gets a 5.4 on cyber adversarial gets a 5.4 on cyber adversarial robustness compared to like 50 to 80% robustness compared to like 50 to 80% robustness compared to like 50 to 80% for their other models. They also have for their other models. They also have for their other models. They also have the biology and chemistry classifiers, the biology and chemistry classifiers, the biology and chemistry classifiers, which they don't actually show numbers which they don't actually show numbers which they don't actually show numbers here. My assumption is they just here. My assumption is they just here. My assumption is they just outright refused to answer these outright refused to answer these outright refused to answer these questions, but Mythos preview and five questions, but Mythos preview and five questions, but Mythos preview and five were able to get really high scores in were able to get really high scores in were able to get really high scores in this viral experimental task, which is this viral experimental task, which is this viral experimental task, which is scary. If these models can create novel scary. If these models can create novel scary. If these models can create novel viruses, we're kind of [ __ ] We'll get viruses, we're kind of [ __ ] We'll get viruses, we're kind of [ __ ] We'll get a new COVID every year. They also really a new COVID every year. They also really a new COVID every year. They also really hate distillation attempts. They don't hate distillation attempts. They don't hate distillation attempts. They don't want other companies to be able to want other companies to be able to want other companies to be able to access enough stuff from Mythos to be access enough stuff from Mythos to be access enough stuff from Mythos to be able to train their own models to be able to train their own models to be able to train their own models to be similar levels, so they're restricting similar levels, so they're restricting similar levels, so they're restricting that. Obnoxious and weirdly like that. Obnoxious and weirdly like that. Obnoxious and weirdly like conceited, I would argue, but they've conceited, I would argue, but they've conceited, I would argue, but they've been doing this for a bit. They're going been doing this for a bit. They're going been doing this for a bit. They're going to continue doing this. And now we get to continue doing this. And now we get to continue doing this. And now we get into the novel [ __ ] Things that are into the novel [ __ ] Things that are into the novel [ __ ] Things that are unlike anything they or anyone else has unlike anything they or anyone else has unlike anything they or anyone else has done that should be very concerning to done that should be very concerning to done that should be very concerning to all of us. There are two really big all of us. There are two really big all of us. There are two really big ones. One is detailed right below. The ones. One is detailed right below. The ones. One is detailed right below. The other is hidden pretty deep in the other is hidden pretty deep in the other is hidden pretty deep in the system card. The first is the new data system card. The first is the new data system card. The first is the new data retention policy. We're making a change retention policy. We're making a change retention policy. We're making a change to the way we handle business customer to the way we handle business customer to the way we handle business customer data for Fable five, Mythos five, and data for Fable five, Mythos five, and data for Fable five, Mythos five, and future models with similar or higher future models with similar or higher future models with similar or higher capability levels. We will require capability levels. We will require capability levels. We will require 30-day retention for all traffic on 30-day retention for all traffic on 30-day retention for all traffic on Mythos class models on both first and Mythos class models on both first and Mythos class models on both first and third-party surfaces. This is a huge

  9. third-party surfaces. This is a huge third-party surfaces. This is a huge change. Traditionally, when you have ZDR change. Traditionally, when you have ZDR change. Traditionally, when you have ZDR on, which is zero data retention, it's a on, which is zero data retention, it's a on, which is zero data retention, it's a policy that most companies push for when policy that most companies push for when policy that most companies push for when they are negotiating with a a vendor they are negotiating with a a vendor they are negotiating with a a vendor that they're relying on like Anthropic, that they're relying on like Anthropic, that they're relying on like Anthropic, a lot of their data is going to go to a lot of their data is going to go to a lot of their data is going to go to Anthropic, and they will sign a deal Anthropic, and they will sign a deal Anthropic, and they will sign a deal that says Anthropic can't keep that data that says Anthropic can't keep that data that says Anthropic can't keep that data because that would violate a lot of because that would violate a lot of because that would violate a lot of their agreements. If they have like their agreements. If they have like their agreements. If they have like HIPAA restrictions or other data HIPAA restrictions or other data HIPAA restrictions or other data policies, they can't give you data and policies, they can't give you data and policies, they can't give you data and then expect you to keep it. Like that's then expect you to keep it. Like that's then expect you to keep it. Like that's just it doesn't work that way. So, the just it doesn't work that way. So, the just it doesn't work that way. So, the 30-day retention requirement here 30-day retention requirement here 30-day retention requirement here immediately invalidates a shitload of immediately invalidates a shitload of immediately invalidates a shitload of business use cases. I know a one or two business use cases. I know a one or two business use cases. I know a one or two companies that are allowing it, but the companies that are allowing it, but the companies that are allowing it, but the vast majority of Fortune 500s are just vast majority of Fortune 500s are just vast majority of Fortune 500s are just formally not letting you use Fable 5, formally not letting you use Fable 5, formally not letting you use Fable 5, even companies like Amazon. They do call even companies like Amazon. They do call even companies like Amazon. They do call out that they won't use the data to out that they won't use the data to out that they won't use the data to train new Claude models or for any train new Claude models or for any train new Claude models or for any non-safety related purposes. And they've non-safety related purposes. And they've non-safety related purposes. And they've instituted new privacy protections instituted new privacy protections instituted new privacy protections including logging all human access to including logging all human access to including logging all human access to the data and ensuring its deletion after the data and ensuring its deletion after the data and ensuring its deletion after 30 days in almost all cases. Almost all 30 days in almost all cases. Almost all 30 days in almost all cases. Almost all cases. cases. cases. The data helps them defend against The data helps them defend against The data helps them defend against complex and novel attacks including new complex and novel attacks including new complex and novel attacks including new jailbreaks and attacks that operate jailbreaks and attacks that operate jailbreaks and attacks that operate across many requests as well as helping across many requests as well as helping across many requests as well as helping us identify and reduce false positives.

  10. us identify and reduce false positives. us identify and reduce false positives. I want to fixate on this almost all I want to fixate on this almost all I want to fixate on this almost all though before we get to the worst thing though before we get to the worst thing though before we get to the worst thing they've done. They claim that the data they've done. They claim that the data they've done. They claim that the data is deleted automatically except in rare is deleted automatically except in rare is deleted automatically except in rare cases where it's part of a safety cases where it's part of a safety cases where it's part of a safety investigation or they're legally investigation or they're legally investigation or they're legally required to keep it. Here's where things required to keep it. Here's where things required to keep it. Here's where things get a little messier. If they detect get a little messier. If they detect get a little messier. If they detect what they consider a usage policy what they consider a usage policy what they consider a usage policy violation, they will retain inputs and violation, they will retain inputs and violation, they will retain inputs and outputs for up to two years in trust and outputs for up to two years in trust and outputs for up to two years in trust and safety classification scores for up to safety classification scores for up to safety classification scores for up to seven years if your chat is flagged. seven years if your chat is flagged. seven years if your chat is flagged. They don't specify whether or not they They don't specify whether or not they They don't specify whether or not they can train on the data when that happens. can train on the data when that happens. can train on the data when that happens. So it is absolutely possible although So it is absolutely possible although So it is absolutely possible although likely wouldn't hold up well in court on likely wouldn't hold up well in court on likely wouldn't hold up well in court on their behalf. Again, I'm not a lawyer. I their behalf. Again, I'm not a lawyer. I their behalf. Again, I'm not a lawyer. I don't know any of this for sure. This is don't know any of this for sure. This is don't know any of this for sure. This is just my understanding. The fact that just my understanding. The fact that just my understanding. The fact that they retain inputs and outputs for up to they retain inputs and outputs for up to they retain inputs and outputs for up to two years when things are classified two years when things are classified two years when things are classified means that the 30-day promise in their means that the 30-day promise in their means that the 30-day promise in their newest post doesn't really work out newest post doesn't really work out newest post doesn't really work out great if they are retaining something great if they are retaining something great if they are retaining something that they consider insecure. So if you that they consider insecure. So if you that they consider insecure. So if you have Mythos going through your database have Mythos going through your database have Mythos going through your database or medical records or something and it or medical records or something and it or medical records or something and it hits a username it doesn't like and it hits a username it doesn't like and it hits a username it doesn't like and it flags safety, they're no longer flags safety, they're no longer flags safety, they're no longer retaining that for 30 days in their retaining that for 30 days in their retaining that for 30 days in their special we can't train on this policy.

  11. special we can't train on this policy. special we can't train on this policy. They're now retaining it for two years They're now retaining it for two years They're now retaining it for two years on a policy that does potentially allow on a policy that does potentially allow on a policy that does potentially allow them to train on it, which is entirely them to train on it, which is entirely them to train on it, which is entirely unusable for most real business cases. unusable for most real business cases. unusable for most real business cases. That is sketchy as [ __ ] You should be That is sketchy as [ __ ] You should be That is sketchy as [ __ ] You should be concerned about this. And again, this is concerned about this. And again, this is concerned about this. And again, this is not all sessions. This is just the ones not all sessions. This is just the ones not all sessions. This is just the ones that are flagged for safety that are flagged for safety that are flagged for safety classification. Not good. classification. Not good. classification. Not good. But we're not even at the worst part But we're not even at the worst part But we're not even at the worst part yet. yet. yet. Cuz it gets a lot worse as we go. Cuz it gets a lot worse as we go. Cuz it gets a lot worse as we go. Don't tell me they got rid of the prompt Don't tell me they got rid of the prompt Don't tell me they got rid of the prompt modification section. It was on page 13 modification section. It was on page 13 modification section. It was on page 13 before. before. before. [ __ ] they did. [ __ ] they did. [ __ ] they did. I have the original though. Give me one I have the original though. Give me one I have the original though. Give me one moment cuz I was smart enough to moment cuz I was smart enough to moment cuz I was smart enough to download it cuz I had a feeling things download it cuz I had a feeling things download it cuz I had a feeling things would change on me. would change on me. would change on me. Haha. Haha. Haha. There it is. It really not in the web There it is. It really not in the web There it is. It really not in the web version anymore? version anymore? version anymore? Yeah, they actually changed this. Holy [ __ ] Holy [ __ ] >> [laughter] >> [laughter] >> [laughter] >> I cannot believe they did that. I've >> I cannot believe they did that. I've >> I cannot believe they did that. I've never seen that before. This is a never seen that before. This is a never seen that before. This is a different section. different section. different section. They entire Wow.

  12. They entire Wow. They entire Wow. I cannot believe I caught this. I cannot believe I caught this. I cannot believe I caught this. Anthropic changed the system card. Anthropic changed the system card. Anthropic changed the system card. This is a different document. I This is a different document. I This is a different document. I downloaded this right when it dropped downloaded this right when it dropped downloaded this right when it dropped and this section, the 1.5 novel and this section, the 1.5 novel and this section, the 1.5 novel safeguard section, has been modified. They updated the [ __ ] system card. They updated the [ __ ] system card. Dirty bastards. This is why I Dirty bastards. This is why I Dirty bastards. This is why I obsessively save everything. Like what obsessively save everything. Like what obsessively save everything. Like what the [ __ ] They didn't even call that out the [ __ ] They didn't even call that out the [ __ ] They didn't even call that out at the top. Do they call it out at the top. Do they call it out at the top. Do they call it out somewhere later in it? They didn't somewhere later in it? They didn't somewhere later in it? They didn't update the date either. It still says update the date either. It still says update the date either. It still says June 9th June 9th June 9th on both docs. on both docs. on both docs. Yeah. Yeah. Yeah. [ __ ] dirty. This is the old one. [ __ ] dirty. This is the old one. [ __ ] dirty. This is the old one. Yeah, the the link in their blog post Yeah, the the link in their blog post Yeah, the the link in their blog post has been swapped with a different has been swapped with a different has been swapped with a different version. version. version. Yeah, okay. So, I just live found Yeah, okay. So, I just live found Yeah, okay. So, I just live found Anthropic trying to rewrite history. Anthropic trying to rewrite history. Anthropic trying to rewrite history. This is why I do what I do. This is why This is why I do what I do. This is why This is why I do what I do. This is why I talk so much [ __ ] on this company cuz I talk so much [ __ ] on this company cuz I talk so much [ __ ] on this company cuz like what the [ __ ] You can't just like what the [ __ ] You can't just like what the [ __ ] You can't just rewrite the system card and expect rewrite the system card and expect rewrite the system card and expect people to not notice. Like what the people to not notice. Like what the people to not notice. Like what the [ __ ] [ __ ] [ __ ] So, let's talk about what they're trying So, let's talk about what they're trying So, let's talk about what they're trying to hide because now we've spotted them to hide because now we've spotted them to hide because now we've spotted them trying to hide it.

  13. trying to hide it. trying to hide it. In light of the ability of recent models In light of the ability of recent models In light of the ability of recent models to accelerate their own development, to accelerate their own development, to accelerate their own development, this is a video we've already done. this is a video we've already done. this is a video we've already done. We've implemented new interventions that We've implemented new interventions that We've implemented new interventions that limit Claude's effectiveness for limit Claude's effectiveness for limit Claude's effectiveness for requests targeting frontier LLM requests targeting frontier LLM requests targeting frontier LLM development. For example, on building development. For example, on building development. For example, on building pre-training pipelines, distributed pre-training pipelines, distributed pre-training pipelines, distributed training infrastructure, or ML training infrastructure, or ML training infrastructure, or ML accelerator design. Using Claude to accelerator design. Using Claude to accelerator design. Using Claude to develop competing models already develop competing models already develop competing models already violates our terms of service, but violates our terms of service, but violates our terms of service, but enforcing this restriction through our enforcing this restriction through our enforcing this restriction through our safeguards avoids accelerating the safeguards avoids accelerating the safeguards avoids accelerating the actors most willing to violate those actors most willing to violate those actors most willing to violate those terms. This is where it gets real terms. This is where it gets real terms. This is where it gets real sketchy. Unlike our interventions for sketchy. Unlike our interventions for sketchy. Unlike our interventions for cybersecurity, biology, and chemistry, cybersecurity, biology, and chemistry, cybersecurity, biology, and chemistry, and distillation attempts, these and distillation attempts, these and distillation attempts, these safeguards will not be visible to the safeguards will not be visible to the safeguards will not be visible to the user. So, they won't tell you when the user. So, they won't tell you when the user. So, they won't tell you when the safeguards are enacted when they think safeguards are enacted when they think safeguards are enacted when they think you're training a competing model. They you're training a competing model. They you're training a competing model. They will 5 will not fall back to a different will 5 will not fall back to a different will 5 will not fall back to a different model. Instead, the safeguards will model. Instead, the safeguards will model. Instead, the safeguards will limit effectiveness through methods such limit effectiveness through methods such limit effectiveness through methods such as prompt modification, steering as prompt modification, steering as prompt modification, steering vectors, or parameter efficient vectors, or parameter efficient vectors, or parameter efficient fine-tuning. These interventions will fine-tuning. These interventions will fine-tuning. These interventions will not affect the vast majority of coding not affect the vast majority of coding not affect the vast majority of coding work. We expect they will impact 0.03% work. We expect they will impact 0.03% work. We expect they will impact 0.03% of traffic concentrated in fewer than of traffic concentrated in fewer than of traffic concentrated in fewer than 0.1% of orgs. When these interventions 0.1% of orgs. When these interventions 0.1% of orgs. When these interventions are active, we expect them to have are active, we expect them to have are active, we expect them to have minimal behavioral impact on the model minimal behavioral impact on the model minimal behavioral impact on the model except to limit its effectiveness in except to limit its effectiveness in except to limit its effectiveness in developing frontier LLMs. This is developing frontier LLMs. This is developing frontier LLMs. This is sketchy as [ __ ] They're intentionally sketchy as [ __ ] They're intentionally sketchy as [ __ ] They're intentionally sabotaging your work and billing you sabotaging your work and billing you sabotaging your work and billing you full price without telling you it's full price without telling you it's full price without telling you it's happening because they're that scared of happening because they're that scared of happening because they're that scared of other labs using their model to make other labs using their model to make other labs using their model to make their models better. So, you're paying their models better. So, you're paying their models better. So, you're paying full price for a model that's having its full price for a model that's having its full price for a model that's having its prompts modified. Like this is such a prompts modified. Like this is such a prompts modified. Like this is such a crazy idea. This isn't like crazy idea. This isn't like crazy idea. This isn't like intentionally modifying the history so intentionally modifying the history so intentionally modifying the history so that you can get out around certain

  14. that you can get out around certain that you can get out around certain restrictions like many will do to restrictions like many will do to restrictions like many will do to jailbreak. That's what I initially jailbreak. That's what I initially jailbreak. That's what I initially thought it was when I just saw this thought it was when I just saw this thought it was when I just saw this quote, not even the whole thing, just quote, not even the whole thing, just quote, not even the whole thing, just the words prompt modification. I did not the words prompt modification. I did not the words prompt modification. I did not realize the extent that they were going realize the extent that they were going realize the extent that they were going to here. What this actually means is if to here. What this actually means is if to here. What this actually means is if you ask the model for something like, you ask the model for something like, you ask the model for something like, "Hey, can you help refine this "Hey, can you help refine this "Hey, can you help refine this pre-training pipeline?" it will edit the pre-training pipeline?" it will edit the pre-training pipeline?" it will edit the prompt before sending it to the model to prompt before sending it to the model to prompt before sending it to the model to say, "Hey, my pre-training pipeline's say, "Hey, my pre-training pipeline's say, "Hey, my pre-training pipeline's pretty good. Can you make it worse in pretty good. Can you make it worse in pretty good. Can you make it worse in some subtle way?" Insane. Actually some subtle way?" Insane. Actually some subtle way?" Insane. Actually insane. And this pissed off every insane. And this pissed off every insane. And this pissed off every researcher and every person who cares a researcher and every person who cares a researcher and every person who cares a lot about access to software. There's a lot about access to software. There's a lot about access to software. There's a lot of justified anger at Anthropic for lot of justified anger at Anthropic for lot of justified anger at Anthropic for sandbagging Fable 5 for AI development sandbagging Fable 5 for AI development sandbagging Fable 5 for AI development tasks, but an unanticipated side effect tasks, but an unanticipated side effect tasks, but an unanticipated side effect is that third-party evaluators can no is that third-party evaluators can no is that third-party evaluators can no longer credibly use the model for evals. longer credibly use the model for evals. longer credibly use the model for evals. Case in point, we're in the middle of Case in point, we're in the middle of Case in point, we're in the middle of running really hard AI R&D evals. Fable running really hard AI R&D evals. Fable running really hard AI R&D evals. Fable 5 would be a perfect test candidate, but 5 would be a perfect test candidate, but 5 would be a perfect test candidate, but because of Anthropic's guardrails, we because of Anthropic's guardrails, we because of Anthropic's guardrails, we can't know if the model failed or if can't know if the model failed or if can't know if the model failed or if their classifiers blocked the their classifiers blocked the their classifiers blocked the capability. By the way, this is not just capability. By the way, this is not just capability. By the way, this is not just true for AI R&D. Since Anthropic doesn't true for AI R&D. Since Anthropic doesn't true for AI R&D. Since Anthropic doesn't make it clear when they are sandbagging, make it clear when they are sandbagging, make it clear when they are sandbagging, this could seep into any number of this could seep into any number of this could seep into any number of technical tasks, and the evaluators technical tasks, and the evaluators technical tasks, and the evaluators wouldn't have any way to know. So, they wouldn't have any way to know. So, they wouldn't have any way to know. So, they can't credibly claim to evaluate can't credibly claim to evaluate can't credibly claim to evaluate state-of-the-art accuracy using the state-of-the-art accuracy using the state-of-the-art accuracy using the model. Anthropic's move might sound model. Anthropic's move might sound model. Anthropic's move might sound reasonable if you consider their actions reasonable if you consider their actions reasonable if you consider their actions as a company chasing superintelligence, as a company chasing superintelligence, as a company chasing superintelligence, but consider that customers are spending but consider that customers are spending but consider that customers are spending billions of dollars on their services.

  15. billions of dollars on their services. billions of dollars on their services. That is precisely what has led to their That is precisely what has led to their That is precisely what has led to their recent surge in ARR, popularity, and recent surge in ARR, popularity, and recent surge in ARR, popularity, and fundraising success. So, customers' fundraising success. So, customers' fundraising success. So, customers' surprise and anger is warranted when surprise and anger is warranted when surprise and anger is warranted when they sandbag in evals without even they sandbag in evals without even they sandbag in evals without even informing them about the degraded informing them about the degraded informing them about the degraded capabilities. Antirez, the creator of capabilities. Antirez, the creator of capabilities. Antirez, the creator of Redis, had a really good thread about Redis, had a really good thread about Redis, had a really good thread about this. I want to say a final thing about this. I want to say a final thing about this. I want to say a final thing about my Fable first reaction. my life to my Fable first reaction. my life to my Fable first reaction. my life to programming, and I'll use every programming, and I'll use every programming, and I'll use every innovation in the field also to extract innovation in the field also to extract innovation in the field also to extract value and bring it to the local value and bring it to the local value and bring it to the local inference world to Redis and so forth. inference world to Redis and so forth. inference world to Redis and so forth. But, I believe what Anthropic is doing, But, I believe what Anthropic is doing, But, I believe what Anthropic is doing, gating the ability to do certain gating the ability to do certain gating the ability to do certain harmless things like LLM research and harmless things like LLM research and harmless things like LLM research and with incredibly sensitive filters that with incredibly sensitive filters that with incredibly sensitive filters that even medical questions are often even medical questions are often even medical questions are often blocked, is deeply wrong. They got open blocked, is deeply wrong. They got open blocked, is deeply wrong. They got open research, the transformer, GPT-2. They research, the transformer, GPT-2. They research, the transformer, GPT-2. They train on tons of public data, and I'm train on tons of public data, and I'm train on tons of public data, and I'm the first to say that training is not the first to say that training is not the first to say that training is not copying the content, but this is okay as copying the content, but this is okay as copying the content, but this is okay as long as the training you do is not used long as the training you do is not used long as the training you do is not used against the same culture that allowed against the same culture that allowed against the same culture that allowed you to create what you created. We need you to create what you created. We need you to create what you created. We need to oppose all of that. The short-term to oppose all of that. The short-term to oppose all of that. The short-term escape is that other frontier labs like escape is that other frontier labs like escape is that other frontier labs like Google and OpenAI will release models Google and OpenAI will release models Google and OpenAI will release models that are on par, and in the case of that are on par, and in the case of that are on par, and in the case of OpenAI I have zero doubt they are OpenAI I have zero doubt they are OpenAI I have zero doubt they are leading for months now. But, it is still leading for months now. But, it is still leading for months now. But, it is still a duopoly or a triopoly, which is odd.

  16. a duopoly or a triopoly, which is odd. a duopoly or a triopoly, which is odd. The escape is open weight models from The escape is open weight models from The escape is open weight models from China. As much as I cheer for what China. As much as I cheer for what China. As much as I cheer for what Chinese labs are doing, remember that Chinese labs are doing, remember that Chinese labs are doing, remember that this structure making all that possible this structure making all that possible this structure making all that possible was basically conceived in the West, was basically conceived in the West, was basically conceived in the West, both the scientific, technological both the scientific, technological both the scientific, technological organization, and the economic system. organization, and the economic system. organization, and the economic system. So, what happened to us lately, So, what happened to us lately, So, what happened to us lately, especially in Europe? In the US, there especially in Europe? In the US, there especially in Europe? In the US, there are those few unicorns, but where is all are those few unicorns, but where is all are those few unicorns, but where is all the rest of the AI scene? We need to the rest of the AI scene? We need to the rest of the AI scene? We need to recover our industrial ethics and stop recover our industrial ethics and stop recover our industrial ethics and stop accepting a narration that see ourselves accepting a narration that see ourselves accepting a narration that see ourselves boiled. Yep, this is unprecedented. I boiled. Yep, this is unprecedented. I boiled. Yep, this is unprecedented. I love this particular post from Trevor love this particular post from Trevor love this particular post from Trevor Blackwell as well. Remember when Blackwell as well. Remember when Blackwell as well. Remember when compilers would detect that someone was compilers would detect that someone was compilers would detect that someone was using it to build another compiler and using it to build another compiler and using it to build another compiler and silently inject bugs? silently inject bugs? silently inject bugs? Do you understand how insane that would Do you understand how insane that would Do you understand how insane that would be? Imagine if your MacBook refused to be? Imagine if your MacBook refused to be? Imagine if your MacBook refused to let you work on Windows if you were a let you work on Windows if you were a let you work on Windows if you were a Microsoft employee. Imagine that your Microsoft employee. Imagine that your Microsoft employee. Imagine that your iPhone refused to let you take a photo iPhone refused to let you take a photo iPhone refused to let you take a photo if somebody was holding an Android phone if somebody was holding an Android phone if somebody was holding an Android phone on the other side. Imagine a world where on the other side. Imagine a world where on the other side. Imagine a world where your devices just fail when you're your devices just fail when you're your devices just fail when you're trying to do things that compete with trying to do things that compete with trying to do things that compete with the company that made the thing. It's the company that made the thing. It's the company that made the thing. It's the very definition of pulling the the very definition of pulling the the very definition of pulling the ladder up behind you. I made a silly ladder up behind you. I made a silly ladder up behind you. I made a silly joke about this with Karpathy joining joke about this with Karpathy joining joke about this with Karpathy joining Anthropic recently. Do you think he Anthropic recently. Do you think he Anthropic recently. Do you think he joined so that he could keep using joined so that he could keep using joined so that he could keep using Anthropic models in particular mythos Anthropic models in particular mythos Anthropic models in particular mythos for ML research without restrictions?

  17. for ML research without restrictions? for ML research without restrictions? It's a joke, to be clear, but there is It's a joke, to be clear, but there is It's a joke, to be clear, but there is some truth to this. If you want to use some truth to this. If you want to use some truth to this. If you want to use the best models to do this type of the best models to do this type of the best models to do this type of important difficult work, you now have important difficult work, you now have important difficult work, you now have to work at Anthropic. There's never been to work at Anthropic. There's never been to work at Anthropic. There's never been a precedent like this before. There have a precedent like this before. There have a precedent like this before. There have been attempts to prevent models from been attempts to prevent models from been attempts to prevent models from responding to things that they shouldn't responding to things that they shouldn't responding to things that they shouldn't for legitimate safety reasons. There was for legitimate safety reasons. There was for legitimate safety reasons. There was also attempts to hide certain data to also attempts to hide certain data to also attempts to hide certain data to make distillation harder. Things like make distillation harder. Things like make distillation harder. Things like reducing how much reasoning info gets to reducing how much reasoning info gets to reducing how much reasoning info gets to you instead of just sharing the entirety you instead of just sharing the entirety you instead of just sharing the entirety of the reasoning path and traces. That of the reasoning path and traces. That of the reasoning path and traces. That all made some sense. This is a lot all made some sense. This is a lot all made some sense. This is a lot further. The idea that they aren't even further. The idea that they aren't even further. The idea that they aren't even telling you when your stuff gets telling you when your stuff gets telling you when your stuff gets reclassified. Correction to earlier, reclassified. Correction to earlier, reclassified. Correction to earlier, there is a change log in the updated there is a change log in the updated there is a change log in the updated version, so that's nice. version, so that's nice. version, so that's nice. They also call out that they had a They also call out that they had a They also call out that they had a previous description of the initial previous description of the initial previous description of the initial frontier LM development safeguards. frontier LM development safeguards. frontier LM development safeguards. They've updated the behavior of the They've updated the behavior of the They've updated the behavior of the safeguards. Here's the tweet. Kind of safeguards. Here's the tweet. Kind of safeguards. Here's the tweet. Kind of wild to have a tweet linked on the wild to have a tweet linked on the wild to have a tweet linked on the second page of a doc like this. But second page of a doc like this. But second page of a doc like this. But yeah, Twitter's that core to research. yeah, Twitter's that core to research. yeah, Twitter's that core to research. Obnoxious. This is the post I was Obnoxious. This is the post I was Obnoxious. This is the post I was looking for though. We've rolled out looking for though. We've rolled out looking for though. We've rolled out changes to make Fable 5 safeguards for changes to make Fable 5 safeguards for changes to make Fable 5 safeguards for Frontier LM development visible.

  18. Frontier LM development visible. Frontier LM development visible. Starting this week, flagged requests Starting this week, flagged requests Starting this week, flagged requests will visibly fall back to Opus 48, will visibly fall back to Opus 48, will visibly fall back to Opus 48, {m-dash} the same as our safeguards for {m-dash} the same as our safeguards for {m-dash} the same as our safeguards for cyber and bio. This was absolutely cyber and bio. This was absolutely cyber and bio. This was absolutely written by Fable by the way. You will written by Fable by the way. You will written by Fable by the way. You will see this every time it happens. On the see this every time it happens. On the see this every time it happens. On the API, any flagged request will return a API, any flagged request will return a API, any flagged request will return a reason for their refusal. Coming to reason for their refusal. Coming to reason for their refusal. Coming to server-side fallback in the next few server-side fallback in the next few server-side fallback in the next few days. We wanted to deploy Fable 5 to our days. We wanted to deploy Fable 5 to our days. We wanted to deploy Fable 5 to our users quickly and safely. Visible users quickly and safely. Visible users quickly and safely. Visible safeguards can be probed, so they have safeguards can be probed, so they have safeguards can be probed, so they have to be robust, which takes time to get to be robust, which takes time to get to be robust, which takes time to get right. Invisible safeguards can be right. Invisible safeguards can be right. Invisible safeguards can be targeted more narrowly, allowing us to targeted more narrowly, allowing us to targeted more narrowly, allowing us to ship quickly with very few false ship quickly with very few false ship quickly with very few false positives. We went with the invisible positives. We went with the invisible positives. We went with the invisible safeguards for this reason, and that was safeguards for this reason, and that was safeguards for this reason, and that was the wrong trade-off. You should have the wrong trade-off. You should have the wrong trade-off. You should have visibility into the safeguards we have visibility into the safeguards we have visibility into the safeguards we have in place and why. We're sorry for not in place and why. We're sorry for not in place and why. We're sorry for not getting the balance right. The getting the balance right. The getting the balance right. The compromise here is that it's now going compromise here is that it's now going compromise here is that it's now going to flag way more aggressively. Making to flag way more aggressively. Making to flag way more aggressively. Making the safeguards visible makes them easier the safeguards visible makes them easier the safeguards visible makes them easier to work around, so keeping them robust to work around, so keeping them robust to work around, so keeping them robust to jailbreaks will unfortunately mean to jailbreaks will unfortunately mean to jailbreaks will unfortunately mean more false positives while we improve more false positives while we improve more false positives while we improve the classifiers. Yep. Things are going the classifiers. Yep. Things are going the classifiers. Yep. Things are going to get worse. We're also tuning our bio to get worse. We're also tuning our bio to get worse. We're also tuning our bio and cyber classifiers to trigger less and cyber classifiers to trigger less and cyber classifiers to trigger less often on harmless requests. We know this often on harmless requests. We know this often on harmless requests. We know this is frustrating and we'll do our best to is frustrating and we'll do our best to is frustrating and we'll do our best to keep this period as short as possible.

  19. keep this period as short as possible. keep this period as short as possible. If you think a request has been If you think a request has been If you think a request has been mistakenly flagged, run {slash} feedback mistakenly flagged, run {slash} feedback mistakenly flagged, run {slash} feedback in Claude code, click thumbs down on the in Claude code, click thumbs down on the in Claude code, click thumbs down on the fallback in Claude AI or code work, or fallback in Claude AI or code work, or fallback in Claude AI or code work, or file the safeguards appeal form for API file the safeguards appeal form for API file the safeguards appeal form for API requests. Your reports help us tune requests. Your reports help us tune requests. Your reports help us tune these classifiers and we appreciate your these classifiers and we appreciate your these classifiers and we appreciate your feedback. How about you give us our feedback. How about you give us our feedback. How about you give us our [ __ ] money back for the things that [ __ ] money back for the things that [ __ ] money back for the things that we didn't get good responses from? I we didn't get good responses from? I we didn't get good responses from? I think every user who has hit one of think every user who has hit one of think every user who has hit one of these invisible fallbacks should have these invisible fallbacks should have these invisible fallbacks should have their limits reset and or a refund for their limits reset and or a refund for their limits reset and or a refund for their usage in that time. It's insane their usage in that time. It's insane their usage in that time. It's insane that they were just quietly doing that. that they were just quietly doing that. that they were just quietly doing that. I am thankful that they have been peer I am thankful that they have been peer I am thankful that they have been peer pressured by the entire research pressured by the entire research pressured by the entire research community out of doing this [ __ ] community out of doing this [ __ ] community out of doing this [ __ ] But they're also using this as an excuse But they're also using this as an excuse But they're also using this as an excuse to go harder with their restrictions. to go harder with their restrictions. to go harder with their restrictions. So, So, So, yeah. yeah. yeah. So, why the [ __ ] did they suddenly do So, why the [ __ ] did they suddenly do So, why the [ __ ] did they suddenly do this? Why were they not restricting this this? Why were they not restricting this this? Why were they not restricting this with Opus, but they are restricting with with Opus, but they are restricting with with Opus, but they are restricting with Mythos. I have a conspiracy theory here, and I have a conspiracy theory here, and it's not about why their website takes it's not about why their website takes it's not about why their website takes so god damn long to load blog posts. I so god damn long to load blog posts. I so god damn long to load blog posts. I have separate conspiracies for that. have separate conspiracies for that. have separate conspiracies for that. This conspiracy is about this particular This conspiracy is about this particular This conspiracy is about this particular chart in their recursive chart in their recursive chart in their recursive self-improvement article that I covered self-improvement article that I covered self-improvement article that I covered in a video recently. Sorry, quick one in a video recently. Sorry, quick one in a video recently. Sorry, quick one guy crash out. Bro is so biased against guy crash out. Bro is so biased against guy crash out. Bro is so biased against Claude. This is insane.

  20. Claude. This is insane. Claude. This is insane. Are you [ __ ] joking? Are you [ __ ] joking? Are you [ __ ] joking? I just did two videos in a row glazing I just did two videos in a row glazing I just did two videos in a row glazing the absolute [ __ ] out of Anthropic to the absolute [ __ ] out of Anthropic to the absolute [ __ ] out of Anthropic to the point where I'm being called a the point where I'm being called a the point where I'm being called a fanboy. fanboy. fanboy. I'm sorry for holding companies I'm sorry for holding companies I'm sorry for holding companies responsible for their [ __ ] Mr. responsible for their [ __ ] Mr. responsible for their [ __ ] Mr. Fable_Yummy. Fable_Yummy. Fable_Yummy. I used to be a voice for the greater I used to be a voice for the greater I used to be a voice for the greater good, but now I'm just a shell. good, but now I'm just a shell. good, but now I'm just a shell. Goodbye forever. Goodbye forever. Goodbye forever. As I was saying, this particular chart As I was saying, this particular chart As I was saying, this particular chart is where my conspiracy comes from. Where is where my conspiracy comes from. Where is where my conspiracy comes from. Where researcher went wrong, could Claude have researcher went wrong, could Claude have researcher went wrong, could Claude have done better? Historically, the way done better? Historically, the way done better? Historically, the way researchers work is they try to make the researchers work is they try to make the researchers work is they try to make the model good at things. They get feedback model good at things. They get feedback model good at things. They get feedback from people who want to use the model from people who want to use the model from people who want to use the model for those things. They somehow find ways for those things. They somehow find ways for those things. They somehow find ways to measure the success at those things, to measure the success at those things, to measure the success at those things, and then they build a system that lets and then they build a system that lets and then they build a system that lets the model get rewarded when it gets the model get rewarded when it gets the model get rewarded when it gets things correct, which slowly makes the things correct, which slowly makes the things correct, which slowly makes the model better and better at those things. model better and better at those things. model better and better at those things. If you give the model the ability to be If you give the model the ability to be If you give the model the ability to be graded on its success or failures, it graded on its success or failures, it graded on its success or failures, it will be able to improve at that thing. will be able to improve at that thing. will be able to improve at that thing. Historically, their focus has been Historically, their focus has been Historically, their focus has been things that people use the models for things that people use the models for things that people use the models for that aren't necessarily the researchers, that aren't necessarily the researchers, that aren't necessarily the researchers, things like asking about medical things like asking about medical things like asking about medical questions or code or getting personal questions or code or getting personal questions or code or getting personal help or This also gets a lot of feedback help or This also gets a lot of feedback help or This also gets a lot of feedback from people hitting the thumbs up, from people hitting the thumbs up, from people hitting the thumbs up, thumbs down in the chat apps like thumbs down in the chat apps like thumbs down in the chat apps like ChatGPT and Claude. That is the system ChatGPT and Claude. That is the system ChatGPT and Claude. That is the system they use to make the models better. They they use to make the models better. They they use to make the models better. They get feedback on what's good and bad.

  21. get feedback on what's good and bad. get feedback on what's good and bad. They find ways to grade what's good and They find ways to grade what's good and They find ways to grade what's good and bad, and they do reinforcement learning bad, and they do reinforcement learning bad, and they do reinforcement learning to get the model to behave more how they to get the model to behave more how they to get the model to behave more how they want it to. This chart suggests want it to. This chart suggests want it to. This chart suggests something very interesting is happening. something very interesting is happening. something very interesting is happening. This chart is based on something the This chart is based on something the This chart is based on something the researchers did where they wanted to to researchers did where they wanted to to researchers did where they wanted to to how much better the model was than a how much better the model was than a how much better the model was than a researcher going wrong. So, if a researcher going wrong. So, if a researcher going wrong. So, if a researcher is talking with Claude code researcher is talking with Claude code researcher is talking with Claude code to test out some theory they have, they to test out some theory they have, they to test out some theory they have, they have two back and forth, say have like have two back and forth, say have like have two back and forth, say have like two messages they send and everything's two messages they send and everything's two messages they send and everything's going well so far, and then they send to going well so far, and then they send to going well so far, and then they send to the wrong message or something that the wrong message or something that the wrong message or something that isn't quite correct, and it sends the isn't quite correct, and it sends the isn't quite correct, and it sends the model down the wrong path, where model down the wrong path, where model down the wrong path, where suddenly this history went from pretty suddenly this history went from pretty suddenly this history went from pretty good and going in the right direction to good and going in the right direction to good and going in the right direction to off the rails, not where it's supposed off the rails, not where it's supposed off the rails, not where it's supposed to be, and then they have to pull it to be, and then they have to pull it to be, and then they have to pull it back to where it's supposed to be. They back to where it's supposed to be. They back to where it's supposed to be. They decided they wanted to see if the model decided they wanted to see if the model decided they wanted to see if the model could have made a better guess for that could have made a better guess for that could have made a better guess for that third step than the researcher did. So, third step than the researcher did. So, third step than the researcher did. So, they took these histories, they removed they took these histories, they removed they took these histories, they removed the message where the researcher went the message where the researcher went the message where the researcher went wrong, and they asked the model, "Hey, wrong, and they asked the model, "Hey, wrong, and they asked the model, "Hey, what do you think we should do next?" what do you think we should do next?" what do you think we should do next?" And then checked to see if that was And then checked to see if that was And then checked to see if that was better than the researcher did. So, this better than the researcher did. So, this better than the researcher did. So, this chart is not measuring whether the chart is not measuring whether the chart is not measuring whether the models are smarter than researchers on models are smarter than researchers on models are smarter than researchers on average, it's measuring whether the average, it's measuring whether the average, it's measuring whether the models are smarter than researchers when models are smarter than researchers when models are smarter than researchers when the researcher was already wrong.

  22. the researcher was already wrong. the researcher was already wrong. To do this, you need a shitload of data, To do this, you need a shitload of data, To do this, you need a shitload of data, which they have based on the internal which they have based on the internal which they have based on the internal Claude code research sessions, the 129 Claude code research sessions, the 129 Claude code research sessions, the 129 of them they used for this particular of them they used for this particular of them they used for this particular study. study. study. And now, they have a lot of data. Now, And now, they have a lot of data. Now, And now, they have a lot of data. Now, they have a chart that measures that they have a chart that measures that they have a chart that measures that Mythos picks better than the incorrect Mythos picks better than the incorrect Mythos picks better than the incorrect researcher 64% of the time. Most researcher 64% of the time. Most researcher 64% of the time. Most importantly though, they now have the importantly though, they now have the importantly though, they now have the ability to measure this, which means it ability to measure this, which means it ability to measure this, which means it is very likely this ended up in the is very likely this ended up in the is very likely this ended up in the training data. I genuinely believe that training data. I genuinely believe that training data. I genuinely believe that Mythos 5 has proprietary Anthropic Mythos 5 has proprietary Anthropic Mythos 5 has proprietary Anthropic information in its weights. I think they information in its weights. I think they information in its weights. I think they accidentally trained the model to be accidentally trained the model to be accidentally trained the model to be better at their research on their stack better at their research on their stack better at their research on their stack in their proprietary environments, and in their proprietary environments, and in their proprietary environments, and they noticed that it was possible to get they noticed that it was possible to get they noticed that it was possible to get Mythos to give out that proprietary Mythos to give out that proprietary Mythos to give out that proprietary information. They also probably like information. They also probably like information. They also probably like that though, because it makes the model that though, because it makes the model that though, because it makes the model better at their work. So, now they have better at their work. So, now they have better at their work. So, now they have to find a balance because they can't let to find a balance because they can't let to find a balance because they can't let their competitors get access to this their competitors get access to this their competitors get access to this private information and this IP that private information and this IP that private information and this IP that they should not have let into the model, they should not have let into the model, they should not have let into the model, and that's why they did what they did and that's why they did what they did and that's why they did what they did here. They don't want people finding here. They don't want people finding here. They don't want people finding ways to sneak in to these weights to get ways to sneak in to these weights to get ways to sneak in to these weights to get the proprietary information that the proprietary information that the proprietary information that accidentally ended up in the model, and accidentally ended up in the model, and accidentally ended up in the model, and there's no way it's coming out now.

  23. there's no way it's coming out now. there's no way it's coming out now. Fables silent invisible rerouting was Fables silent invisible rerouting was Fables silent invisible rerouting was the solution to this. Make it basically the solution to this. Make it basically the solution to this. Make it basically impossible to even know that the model impossible to even know that the model impossible to even know that the model has this information in it, and that's has this information in it, and that's has this information in it, and that's why I think they went so hard here why I think they went so hard here why I think they went so hard here because researchers are now steering the because researchers are now steering the because researchers are now steering the model to do their jobs better for the model to do their jobs better for the model to do their jobs better for the first time. They didn't realize the first time. They didn't realize the first time. They didn't realize the consequences of their actions, and now consequences of their actions, and now consequences of their actions, and now they have to do something about it post they have to do something about it post they have to do something about it post talk, and that's how we got here. While talk, and that's how we got here. While talk, and that's how we got here. While they can take back this particular they can take back this particular they can take back this particular implementation, they can't take back the implementation, they can't take back the implementation, they can't take back the fact that it has happened, and the fact that it has happened, and the fact that it has happened, and the floodgates have now opened. The idea floodgates have now opened. The idea floodgates have now opened. The idea that a model can now quietly make your that a model can now quietly make your that a model can now quietly make your stuff worse, that's a real supply chain stuff worse, that's a real supply chain stuff worse, that's a real supply chain risk, and I think this blog post is a risk, and I think this blog post is a risk, and I think this blog post is a good job describing it. Thank you to good job describing it. Thank you to good job describing it. Thank you to John Ready for writing it. Anthropic John Ready for writing it. Anthropic John Ready for writing it. Anthropic says these safeguards only affect .03% says these safeguards only affect .03% says these safeguards only affect .03% of devs. Maybe that's true today. The of devs. Maybe that's true today. The of devs. Maybe that's true today. The problem is that the definition of an AI problem is that the definition of an AI problem is that the definition of an AI company is changing. Maybe you're not company is changing. Maybe you're not company is changing. Maybe you're not training frontier models today. Most training frontier models today. Most training frontier models today. Most companies aren't, but modern software is companies aren't, but modern software is companies aren't, but modern software is increasingly containing AI models. Five increasingly containing AI models. Five increasingly containing AI models. Five years ago, building a startup meant years ago, building a startup meant years ago, building a startup meant writing APIs and SQL queries. Today, it writing APIs and SQL queries. Today, it writing APIs and SQL queries. Today, it often means training, tuning, and often means training, tuning, and often means training, tuning, and deploying models. Five years ago, models deploying models. Five years ago, models deploying models. Five years ago, models like CLIP were frontier AI research like CLIP were frontier AI research like CLIP were frontier AI research projects. Today, I'm fine-tuning them projects. Today, I'm fine-tuning them projects. Today, I'm fine-tuning them for a bootstrapped travel startup. If for a bootstrapped travel startup. If for a bootstrapped travel startup. If you're debugging a model training you're debugging a model training you're debugging a model training pipeline for your product and Claude pipeline for your product and Claude pipeline for your product and Claude gives a bad answer, was the model gives a bad answer, was the model gives a bad answer, was the model confused? Did you give it bad context?

  24. confused? Did you give it bad context? confused? Did you give it bad context? Or did a hidden policy nerf Claude's Or did a hidden policy nerf Claude's Or did a hidden policy nerf Claude's ability to assist you? You won't know. ability to assist you? You won't know. ability to assist you? You won't know. This is the end of our ability to trust This is the end of our ability to trust This is the end of our ability to trust the model, and that sucks because as the model, and that sucks because as the model, and that sucks because as much as I hate Anthropic, they were at much as I hate Anthropic, they were at much as I hate Anthropic, they were at least relatively transparent with their least relatively transparent with their least relatively transparent with their [ __ ] And while I am thankful they [ __ ] And while I am thankful they [ __ ] And while I am thankful they decided to walk this one back because decided to walk this one back because decided to walk this one back because this was bad. This was really bad. It's this was bad. This was really bad. It's this was bad. This was really bad. It's good that they walked it back. It's good that they walked it back. It's good that they walked it back. It's still really, really bad that they still really, really bad that they still really, really bad that they opened these floodgates. I now feel less opened these floodgates. I now feel less opened these floodgates. I now feel less like I can trust the reasons why the like I can trust the reasons why the like I can trust the reasons why the model responds poorly. I now can't trust model responds poorly. I now can't trust model responds poorly. I now can't trust the outputs the same way. And as John the outputs the same way. And as John the outputs the same way. And as John put it, this is now a real supply chain put it, this is now a real supply chain put it, this is now a real supply chain risk. The same company that I defended risk. The same company that I defended risk. The same company that I defended against the classification of supply against the classification of supply against the classification of supply chain risk earlier this year, I defended chain risk earlier this year, I defended chain risk earlier this year, I defended the [ __ ] out of them in the issues with the [ __ ] out of them in the issues with the [ __ ] out of them in the issues with the Department of War. This does make the Department of War. This does make the Department of War. This does make them more of a supply chain risk. This them more of a supply chain risk. This them more of a supply chain risk. This does mean that if you have them as a does mean that if you have them as a does mean that if you have them as a dependency in your way of building, in dependency in your way of building, in dependency in your way of building, in your pipeline, in your teams, in your your pipeline, in your teams, in your your pipeline, in your teams, in your business, you can't trust it the same business, you can't trust it the same business, you can't trust it the same way anymore. And I do not like the way anymore. And I do not like the way anymore. And I do not like the precedent that sets. Hopefully you guys precedent that sets. Hopefully you guys precedent that sets. Hopefully you guys don't think I'm overreacting here. I'm don't think I'm overreacting here. I'm don't think I'm overreacting here. I'm trying to be as reasonable as I can with trying to be as reasonable as I can with trying to be as reasonable as I can with something this absurd. I love the model, something this absurd. I love the model, something this absurd. I love the model, but I hate [ __ ] like this. I'm thankful but I hate [ __ ] like this. I'm thankful but I hate [ __ ] like this. I'm thankful they walked it back, but I'm scared that they walked it back, but I'm scared that they walked it back, but I'm scared that the precedent's been set and things can the precedent's been set and things can the precedent's been set and things can get worse going forward. Let me know how get worse going forward. Let me know how get worse going forward. Let me know how you feel about this and until next time, you feel about this and until next time, you feel about this and until next time, peace nerds.

Summary

The main theme is the unacceptable restrictions on Anthropic's new Fable model, despite its advanced capabilities and superior benchmarks to models like GPT-5.5 and Opus-4.8. The takeaway is that these limitations set a bad precedent for AI development, hindering broader industry progress and access to powerful tools.

View original episode ↗