FABLE IS BACK! (And Sonnet 5 is here too)
Read full transcript 22 segments
-
Seems like a pretty awesome time to be a Seems like a pretty awesome time to be a Claude fan cuz we just got two huge Claude fan cuz we just got two huge Claude fan cuz we just got two huge pieces of news. The first is a new pieces of news. The first is a new pieces of news. The first is a new model, Sonnet 5. It's finally here and model, Sonnet 5. It's finally here and model, Sonnet 5. It's finally here and there's a lot to talk about with it. there's a lot to talk about with it. there's a lot to talk about with it. This model is definitely not what you This model is definitely not what you This model is definitely not what you think and I haven't seen any reporting think and I haven't seen any reporting think and I haven't seen any reporting really covering its strengths and really covering its strengths and really covering its strengths and weaknesses properly cuz it has plenty of weaknesses properly cuz it has plenty of weaknesses properly cuz it has plenty of both, believe me. But in more important both, believe me. But in more important both, believe me. But in more important news, Fable 5 has just been unbanned by news, Fable 5 has just been unbanned by news, Fable 5 has just been unbanned by the Secretary of Commerce. The the Secretary of Commerce. The the Secretary of Commerce. The restrictions have been lifted. The restrictions have been lifted. The restrictions have been lifted. The model's not back as of the time I'm model's not back as of the time I'm model's not back as of the time I'm recording this, but there's a very good recording this, but there's a very good recording this, but there's a very good chance it'll be out by the time you're chance it'll be out by the time you're chance it'll be out by the time you're watching this video. So go double check. watching this video. So go double check. watching this video. So go double check. Good chance you have it. I have a little Good chance you have it. I have a little Good chance you have it. I have a little more to say about that, but I really more to say about that, but I really more to say about that, but I really want to focus in on Sonet 5 because, as want to focus in on Sonet 5 because, as want to focus in on Sonet 5 because, as I said, it's a very interesting model. I said, it's a very interesting model. I said, it's a very interesting model. I've been using it all day, testing it I've been using it all day, testing it I've been using it all day, testing it on various tasks and benchmarks. And on various tasks and benchmarks. And on various tasks and benchmarks. And with models getting so expensive, my with models getting so expensive, my with models getting so expensive, my individual benchmark runs can cost way individual benchmark runs can cost way individual benchmark runs can cost way over $300. That number is only counting over $300. That number is only counting over $300. That number is only counting runs that didn't fail. Failure still runs that didn't fail. Failure still runs that didn't fail. Failure still cost me money, believe it or not. As cost me money, believe it or not. As cost me money, believe it or not. As such, I need to take a quick break for such, I need to take a quick break for such, I need to take a quick break for today's sponsor. You're probably not today's sponsor. You're probably not today's sponsor. You're probably not being bold enough with agents. I'm being bold enough with agents. I'm being bold enough with agents. I'm saying this because it was the case for saying this because it was the case for saying this because it was the case for me. I still remember back in the day me. I still remember back in the day me. I still remember back in the day when Devon was first announced and they when Devon was first announced and they when Devon was first announced and they claimed that AI would be able to be a claimed that AI would be able to be a claimed that AI would be able to be a full engineer as though it's on your full engineer as though it's on your full engineer as though it's on your team and that made no sense to me at the team and that made no sense to me at the team and that made no sense to me at the time. That's why when they hit me up, I time. That's why when they hit me up, I time. That's why when they hit me up, I had to go dig into the product more and had to go dig into the product more and had to go dig into the product more and I've been blown away. You haven't kept I've been blown away. You haven't kept I've been blown away. You haven't kept up with what they're working on, they up with what they're working on, they up with what they're working on, they made it possible to spin up your made it possible to spin up your made it possible to spin up your codebase in the cloud. One of the best codebase in the cloud. One of the best codebase in the cloud. One of the best setups for that, by the way. And once setups for that, by the way. And once setups for that, by the way. And once you have it set up, you can run a real you have it set up, you can run a real you have it set up, you can run a real Linux box that your agents will control Linux box that your agents will control Linux box that your agents will control as they do work. This means you can as they do work. This means you can as they do work. This means you can develop anything from a web app to a develop anything from a web app to a develop anything from a web app to a real desktop application with agents real desktop application with agents real desktop application with agents able to test things, do things, run able to test things, do things, run able to test things, do things, run things, and you can even interact with things, and you can even interact with things, and you can even interact with it. All of that's cool by itself, but it. All of that's cool by itself, but it. All of that's cool by itself, but what I did here is even cooler. I told what I did here is even cooler. I told what I did here is even cooler. I told Devon to go through all the pages on my Devon to go through all the pages on my Devon to go through all the pages on my website specifically to check for mobile website specifically to check for mobile website specifically to check for mobile responsiveness, UI bugs, and client side
-
responsiveness, UI bugs, and client side responsiveness, UI bugs, and client side errors. Obviously, you could run this as errors. Obviously, you could run this as errors. Obviously, you could run this as a single agent run, but it's going to a single agent run, but it's going to a single agent run, but it's going to take forever. And I find that when you take forever. And I find that when you take forever. And I find that when you do this on too many things at once, the do this on too many things at once, the do this on too many things at once, the result is often that it's not going as result is often that it's not going as result is often that it's not going as deep as I want. So here it went across deep as I want. So here it went across deep as I want. So here it went across all eight pages on my site and spun up all eight pages on my site and spun up all eight pages on my site and spun up sub agents for each of them and checked sub agents for each of them and checked sub agents for each of them and checked for all the different potential for all the different potential for all the different potential regressions for the way the site regressions for the way the site regressions for the way the site behaves. It passed all the results up to behaves. It passed all the results up to behaves. It passed all the results up to that top level agent so you could see that top level agent so you could see that top level agent so you could see them and make decisions. But you can them and make decisions. But you can them and make decisions. But you can also go into any of those sub agents and also go into any of those sub agents and also go into any of those sub agents and it even spits out videos of what it it even spits out videos of what it it even spits out videos of what it found when it explored. This caught some found when it explored. This caught some found when it explored. This caught some novel regressions that I didn't notice novel regressions that I didn't notice novel regressions that I didn't notice and ended up going to patch right after. and ended up going to patch right after. and ended up going to patch right after. But I want to prevent this from But I want to prevent this from But I want to prevent this from happening again. And this is why happening again. And this is why happening again. And this is why scheduled Devons is so cool. Check this scheduled Devons is so cool. Check this scheduled Devons is so cool. Check this out. Set up a scheduled Devon to check out. Set up a scheduled Devon to check out. Set up a scheduled Devon to check for regressions on my website every day. for regressions on my website every day. for regressions on my website every day. It's kind of crazy. It just asks you It's kind of crazy. It just asks you It's kind of crazy. It just asks you when you want it to go. You hit a button when you want it to go. You hit a button when you want it to go. You hit a button and now you're done. If you're pushing and now you're done. If you're pushing and now you're done. If you're pushing agents to their limits, you're already agents to their limits, you're already agents to their limits, you're already using Devon. And if you're not, fix it using Devon. And if you're not, fix it using Devon. And if you're not, fix it at soyv.link/devon. I'll blast through at soyv.link/devon. I'll blast through at soyv.link/devon. I'll blast through the fable news fast so we could focus on the fable news fast so we could focus on the fable news fast so we could focus on what I'm here to talk about, which of what I'm here to talk about, which of what I'm here to talk about, which of course is sonnet 5. If you haven't course is sonnet 5. If you haven't course is sonnet 5. If you haven't caught up, Fable got banned on June caught up, Fable got banned on June caught up, Fable got banned on June 12th, just 3 days after it originally 12th, just 3 days after it originally 12th, just 3 days after it originally came out because of concerns related to came out because of concerns related to came out because of concerns related to its ability to hack and do things that its ability to hack and do things that its ability to hack and do things that we wouldn't want it to do. In we wouldn't want it to do. In we wouldn't want it to do. In particular, the government was concerned particular, the government was concerned particular, the government was concerned about jailbreak capabilities that would about jailbreak capabilities that would about jailbreak capabilities that would allow the model to find security issues allow the model to find security issues allow the model to find security issues in software, and they wanted to restrict in software, and they wanted to restrict in software, and they wanted to restrict that from foreign actors and foreign that from foreign actors and foreign that from foreign actors and foreign nationals. Anthropic just took that as a nationals. Anthropic just took that as a nationals. Anthropic just took that as a hard ban, but all of those export hard ban, but all of those export hard ban, but all of those export controls have been lifted as of now.
-
controls have been lifted as of now. controls have been lifted as of now. Since the issuance of my previous letter Since the issuance of my previous letter Since the issuance of my previous letter dated June 12th and June 26, Anthropic dated June 12th and June 26, Anthropic dated June 12th and June 26, Anthropic has taken steps in close coordination has taken steps in close coordination has taken steps in close coordination with the US government to address the with the US government to address the with the US government to address the risks associated with Claude Mythos 5 risks associated with Claude Mythos 5 risks associated with Claude Mythos 5 and Fable 5. Among other things, and Fable 5. Among other things, and Fable 5. Among other things, Anthropic has agreed to proactively Anthropic has agreed to proactively Anthropic has agreed to proactively detect and address security risks detect and address security risks detect and address security risks associated with the models to work associated with the models to work associated with the models to work diligently with the US government on diligently with the US government on diligently with the US government on protocols and standards and releases for protocols and standards and releases for protocols and standards and releases for mythos fable and future models as well mythos fable and future models as well mythos fable and future models as well as to inform the US government of any as to inform the US government of any as to inform the US government of any malicious activity. In light of these malicious activity. In light of these malicious activity. In light of these actions and commitments, as well as the actions and commitments, as well as the actions and commitments, as well as the Bureau of Industry and Securities Bureau of Industry and Securities Bureau of Industry and Securities evaluation of the diversion risks now evaluation of the diversion risks now evaluation of the diversion risks now presented by Cloud Mythos 5 and Fable 5, presented by Cloud Mythos 5 and Fable 5, presented by Cloud Mythos 5 and Fable 5, the controls in the June 12th letter are the controls in the June 12th letter are the controls in the June 12th letter are withdrawn. A license is no longer withdrawn. A license is no longer withdrawn. A license is no longer required for the export, reexport, or required for the export, reexport, or required for the export, reexport, or incountry transfer, including deemed incountry transfer, including deemed incountry transfer, including deemed export or deemed reexport of the mythos export or deemed reexport of the mythos export or deemed reexport of the mythos or fable models. Commerce reserves the or fable models. Commerce reserves the or fable models. Commerce reserves the right to reevaluate the decision yada right to reevaluate the decision yada right to reevaluate the decision yada yada yada. Most important piece here is yada yada. Most important piece here is yada yada. Most important piece here is export or reexport because what this export or reexport because what this export or reexport because what this means is if you're building a service means is if you're building a service means is if you're building a service like I don't know, T3 chat where you're like I don't know, T3 chat where you're like I don't know, T3 chat where you're hosting these models over API and then hosting these models over API and then hosting these models over API and then letting users hit them themselves, this letting users hit them themselves, this letting users hit them themselves, this means that that is allowed as well, means that that is allowed as well, means that that is allowed as well, which is really important. I had a lot which is really important. I had a lot which is really important. I had a lot of concerns that might not be allowed of concerns that might not be allowed of concerns that might not be allowed with whatever conclusions they came to.
-
with whatever conclusions they came to. with whatever conclusions they came to. So knowing that that's good makes me So knowing that that's good makes me So knowing that that's good makes me happy, especially because uh I don't happy, especially because uh I don't happy, especially because uh I don't want to use Sonet 5 too much. There are want to use Sonet 5 too much. There are want to use Sonet 5 too much. There are some good parts here, but we will talk some good parts here, but we will talk some good parts here, but we will talk about them. Let's start from what about them. Let's start from what about them. Let's start from what Anthropic has to say. Sonnet 5 is built Anthropic has to say. Sonnet 5 is built Anthropic has to say. Sonnet 5 is built to be the most agentic sonnet model yet. to be the most agentic sonnet model yet. to be the most agentic sonnet model yet. It can make plans, use tools like It can make plans, use tools like It can make plans, use tools like browsers and terminals and run browsers and terminals and run browsers and terminals and run autonomously at a level that a just a autonomously at a level that a just a autonomously at a level that a just a few months ago required larger and more few months ago required larger and more few months ago required larger and more expensive models. We'll talk about the expensive models. We'll talk about the expensive models. We'll talk about the more expensive a lot as we go forward. more expensive a lot as we go forward. more expensive a lot as we go forward. But I also am excited to talk more about But I also am excited to talk more about But I also am excited to talk more about the agentic side because there are some the agentic side because there are some the agentic side because there are some things Sonnet does that no model things Sonnet does that no model things Sonnet does that no model available as of the time of recording available as of the time of recording available as of the time of recording does do. Fable is not out right now does do. Fable is not out right now does do. Fable is not out right now which is why. But uh yeah the number which is why. But uh yeah the number which is why. But uh yeah the number five is the most notable thing here not five is the most notable thing here not five is the most notable thing here not the word sonnet or the fact that it is a the word sonnet or the fact that it is a the word sonnet or the fact that it is a new sonnet and everybody's saying this new sonnet and everybody's saying this new sonnet and everybody's saying this should have been sonnet 48. I don't should have been sonnet 48. I don't should have been sonnet 48. I don't agree. Let's dig in. agree. Let's dig in. agree. Let's dig in. Anthropic even is aware of the fact that Anthropic even is aware of the fact that Anthropic even is aware of the fact that opus is kind of the standard now as they opus is kind of the standard now as they opus is kind of the standard now as they say here. More recently, the clearest say here. More recently, the clearest say here. More recently, the clearest gains in Agenta capabilities have been gains in Agenta capabilities have been gains in Agenta capabilities have been in the opus tier class models. I agree. in the opus tier class models. I agree. in the opus tier class models. I agree. It honestly feels like Anthropic bumped It honestly feels like Anthropic bumped It honestly feels like Anthropic bumped up the tier for everything where what we up the tier for everything where what we up the tier for everything where what we used to use Haiku for, we now use used to use Haiku for, we now use used to use Haiku for, we now use Sonnet. What we used to use Sana for, we Sonnet. What we used to use Sana for, we Sonnet. What we used to use Sana for, we now use Opus. And what we used to use now use Opus. And what we used to use now use Opus. And what we used to use Opus for, we now use Fable. Very clever Opus for, we now use Fable. Very clever Opus for, we now use Fable. Very clever way to get us to spend more money. But way to get us to spend more money. But way to get us to spend more money. But Sonnet is important. So, let's see how Sonnet is important. So, let's see how Sonnet is important. So, let's see how it improved. Its performance is close to it improved. Its performance is close to it improved. Its performance is close to that of Opus 48, but at lower prices.
-
that of Opus 48, but at lower prices. that of Opus 48, but at lower prices. Lower prices. Remember that. is a Lower prices. Remember that. is a Lower prices. Remember that. is a substantial improvement over its substantial improvement over its substantial improvement over its predecessor 46 on important aspects of predecessor 46 on important aspects of predecessor 46 on important aspects of agentic performance like reasoning, tool agentic performance like reasoning, tool agentic performance like reasoning, tool use, coding, and knowledge work. We have use, coding, and knowledge work. We have use, coding, and knowledge work. We have some benches here. We have SWE bench some benches here. We have SWE bench some benches here. We have SWE bench pro, which remember has been almost pro, which remember has been almost pro, which remember has been almost entirely compromised. The levels of entirely compromised. The levels of entirely compromised. The levels of contamination in SW bench are insane. I contamination in SW bench are insane. I contamination in SW bench are insane. I don't think this measures much at this don't think this measures much at this don't think this measures much at this point, but cool, it's in there. There point, but cool, it's in there. There point, but cool, it's in there. There are other benches, and there's even more are other benches, and there's even more are other benches, and there's even more in the system card that we can talk in the system card that we can talk in the system card that we can talk about, but terminal bench was a about, but terminal bench was a about, but terminal bench was a meaningful improvement from the from the meaningful improvement from the from the meaningful improvement from the from the high 60% to the 80% range. UNA's last high 60% to the 80% range. UNA's last high 60% to the 80% range. UNA's last exam was a meaningful bump as well. exam was a meaningful bump as well. exam was a meaningful bump as well. Computer use was a slight bump too. I Computer use was a slight bump too. I Computer use was a slight bump too. I will say from my experience, Anthropic will say from my experience, Anthropic will say from my experience, Anthropic went from leading in computer use to went from leading in computer use to went from leading in computer use to lagging quite a bit behind even Google lagging quite a bit behind even Google lagging quite a bit behind even Google in a lot of cases. And knowledge work in a lot of cases. And knowledge work in a lot of cases. And knowledge work where they have improved meaningfully where they have improved meaningfully where they have improved meaningfully actually scoring slightly higher than actually scoring slightly higher than actually scoring slightly higher than Opus 48 somehow. They have some call Opus 48 somehow. They have some call Opus 48 somehow. They have some call outs about safety stuff, but I will say outs about safety stuff, but I will say outs about safety stuff, but I will say outright Sonnet 5 is too dumb to be a outright Sonnet 5 is too dumb to be a outright Sonnet 5 is too dumb to be a particular safety risk. I would not particular safety risk. I would not particular safety risk. I would not worry. GLM52 is a more meaningful worry. GLM52 is a more meaningful worry. GLM52 is a more meaningful security risk than Sonnet 5. The charts security risk than Sonnet 5. The charts security risk than Sonnet 5. The charts below compare the performance of Sona below compare the performance of Sona below compare the performance of Sona 5046 and Opus 48 at different effort 5046 and Opus 48 at different effort 5046 and Opus 48 at different effort levels on the agentic search evaluation levels on the agentic search evaluation levels on the agentic search evaluation browse comp and the computer use browse comp and the computer use browse comp and the computer use evaluation OSWorld verified. The orange evaluation OSWorld verified. The orange evaluation OSWorld verified. The orange line is the one we're talking about line is the one we're talking about line is the one we're talking about here, Sonnet 5. And as we see here, here, Sonnet 5. And as we see here, here, Sonnet 5. And as we see here, Sonnet 5 does slightly outperform Opus Sonnet 5 does slightly outperform Opus Sonnet 5 does slightly outperform Opus 48, although not much cheaper. Of 48, although not much cheaper. Of 48, although not much cheaper. Of course, the medium and low runs are course, the medium and low runs are course, the medium and low runs are meaningfully cheaper at less than $2 and meaningfully cheaper at less than $2 and meaningfully cheaper at less than $2 and less than $5 per task. Once we're in less than $5 per task. Once we're in less than $5 per task. Once we're in that 5 to10 range, it's slightly more that 5 to10 range, it's slightly more that 5 to10 range, it's slightly more performance at a slightly higher cost.
-
performance at a slightly higher cost. performance at a slightly higher cost. And then that trend continues. What's And then that trend continues. What's And then that trend continues. What's much more interesting to me though is much more interesting to me though is much more interesting to me though is the agentic computer use bench that I'm the agentic computer use bench that I'm the agentic computer use bench that I'm honestly confused why they included honestly confused why they included honestly confused why they included because as you see here, Sonnet 5 does because as you see here, Sonnet 5 does because as you see here, Sonnet 5 does not perform as well as Opus both on not perform as well as Opus both on not perform as well as Opus both on performance and on cost. Is it better performance and on cost. Is it better performance and on cost. Is it better than Sonnet 46? Yes. But you can get than Sonnet 46? Yes. But you can get than Sonnet 46? Yes. But you can get better performance for the same price better performance for the same price better performance for the same price range on Opus medium or high as you range on Opus medium or high as you range on Opus medium or high as you would get from any of the Sonnet would get from any of the Sonnet would get from any of the Sonnet flavors. In fact, Sonnet 5 Max ends up flavors. In fact, Sonnet 5 Max ends up flavors. In fact, Sonnet 5 Max ends up being more expensive than Opus on high. being more expensive than Opus on high. being more expensive than Opus on high. There is more bad news, though. Pretty There is more bad news, though. Pretty There is more bad news, though. Pretty much every single version of Sonnet 5 is much every single version of Sonnet 5 is much every single version of Sonnet 5 is more expensive and worse performing than more expensive and worse performing than more expensive and worse performing than a GPT 5.5 equivalent, according to a GPT 5.5 equivalent, according to a GPT 5.5 equivalent, according to Cursor Bench. Here we see 55 medium Cursor Bench. Here we see 55 medium Cursor Bench. Here we see 55 medium scoring almost as high as Sonnet 5 max. scoring almost as high as Sonnet 5 max. scoring almost as high as Sonnet 5 max. And also Sonet 5 High and Max, costing And also Sonet 5 High and Max, costing And also Sonet 5 High and Max, costing more than the heaviest runs on GBT 5.5. more than the heaviest runs on GBT 5.5. more than the heaviest runs on GBT 5.5. It almost feels like Sonic 5 was It almost feels like Sonic 5 was It almost feels like Sonic 5 was released to advertise how good of a released to advertise how good of a released to advertise how good of a value 55 is on different reasoning value 55 is on different reasoning value 55 is on different reasoning levels because a lot of these benches levels because a lot of these benches levels because a lot of these benches show that absurdity. Reminder, I do want show that absurdity. Reminder, I do want show that absurdity. Reminder, I do want to talk about the things I like about to talk about the things I like about to talk about the things I like about the model and the things I saw when I the model and the things I saw when I the model and the things I saw when I was using it, which we'll get to in a was using it, which we'll get to in a was using it, which we'll get to in a little bit. But we need to talk a bit little bit. But we need to talk a bit little bit. But we need to talk a bit more about benches first. According to more about benches first. According to more about benches first. According to the artificial analysis intelligence the artificial analysis intelligence the artificial analysis intelligence index, it is in fourth place across all index, it is in fourth place across all index, it is in fourth place across all models, falling just slightly behind models, falling just slightly behind models, falling just slightly behind Fable 5, Opus, and GPT55.
-
Fable 5, Opus, and GPT55. Fable 5, Opus, and GPT55. But the intelligence index is far from But the intelligence index is far from But the intelligence index is far from the most interesting part of artificial the most interesting part of artificial the most interesting part of artificial analysis. Personally, I find the cost analysis. Personally, I find the cost analysis. Personally, I find the cost section to be way more interesting. And section to be way more interesting. And section to be way more interesting. And if we look at cost per task, you'll see if we look at cost per task, you'll see if we look at cost per task, you'll see something a bit terrifying. I5 on XH something a bit terrifying. I5 on XH something a bit terrifying. I5 on XH high ends up being cheaper in realworld high ends up being cheaper in realworld high ends up being cheaper in realworld work than Sonnet by more than 2x. In work than Sonnet by more than 2x. In work than Sonnet by more than 2x. In fact, Opus 48 is also cheaper than fact, Opus 48 is also cheaper than fact, Opus 48 is also cheaper than Sonnet 5. But ready for the craziest Sonnet 5. But ready for the craziest Sonnet 5. But ready for the craziest part here? If we cover the cost to run part here? If we cover the cost to run part here? If we cover the cost to run the entire benchmark, this is the total the entire benchmark, this is the total the entire benchmark, this is the total cost. I'm making an assumption here. I cost. I'm making an assumption here. I cost. I'm making an assumption here. I haven't talked to the artificial haven't talked to the artificial haven't talked to the artificial analysis guys about this, but I'm analysis guys about this, but I'm analysis guys about this, but I'm guessing that the cost per task is guessing that the cost per task is guessing that the cost per task is averaged with extremes filtered out, and averaged with extremes filtered out, and averaged with extremes filtered out, and the total cost does not have those the total cost does not have those the total cost does not have those extremes filtered out because Sonnet 5 extremes filtered out because Sonnet 5 extremes filtered out because Sonnet 5 is the most expensive model they've ever is the most expensive model they've ever is the most expensive model they've ever run through the bench at $6,000, topping run through the bench at $6,000, topping run through the bench at $6,000, topping even Fable 5 at $5,600. And just for even Fable 5 at $5,600. And just for even Fable 5 at $5,600. And just for fun, I'm going to throw in GPT55 medium fun, I'm going to throw in GPT55 medium fun, I'm going to throw in GPT55 medium and low in here because they cost a and low in here because they cost a and low in here because they cost a sixth and a 12th as much as Sonnet 5 sixth and a 12th as much as Sonnet 5 sixth and a 12th as much as Sonnet 5 does. And if we go back to the top here, does. And if we go back to the top here, does. And if we go back to the top here, there is a gap in the intelligence, but there is a gap in the intelligence, but there is a gap in the intelligence, but it's not quite as big as you would it's not quite as big as you would it's not quite as big as you would expect for a 10x decrease in total cost.
-
expect for a 10x decrease in total cost. expect for a 10x decrease in total cost. But wait, the Odin didn't Anthropic say But wait, the Odin didn't Anthropic say But wait, the Odin didn't Anthropic say it's cheaper? They did. And the way that it's cheaper? They did. And the way that it's cheaper? They did. And the way that they said it was cheaper is in the price they said it was cheaper is in the price they said it was cheaper is in the price per token. They're debuting it at an per token. They're debuting it at an per token. They're debuting it at an introductory price of $2 per million introductory price of $2 per million introductory price of $2 per million input tokens and $10 per million out input tokens and $10 per million out input tokens and $10 per million out until August 31st, at which point until August 31st, at which point until August 31st, at which point they're going to bump it back to the they're going to bump it back to the they're going to bump it back to the usual for Sonnet, which is three per usual for Sonnet, which is three per usual for Sonnet, which is three per mill in and 15 per mill out. They also mill in and 15 per mill out. They also mill in and 15 per mill out. They also got rid of the weird Sonnet specific got rid of the weird Sonnet specific got rid of the weird Sonnet specific limit that existed on the subscription limit that existed on the subscription limit that existed on the subscription plans. I don't understand why that stuck plans. I don't understand why that stuck plans. I don't understand why that stuck around so long, but it is finally gone. around so long, but it is finally gone. around so long, but it is finally gone. It just counts towards your normal It just counts towards your normal It just counts towards your normal usage. I have seen people racking up usage. I have seen people racking up usage. I have seen people racking up meaningful amounts of usage in quad code meaningful amounts of usage in quad code meaningful amounts of usage in quad code with this, like taking a actually with this, like taking a actually with this, like taking a actually surprising chunk out of their surprising chunk out of their surprising chunk out of their percentage, even on the $200 plan, percentage, even on the $200 plan, percentage, even on the $200 plan, because it's not a very efficient model. because it's not a very efficient model. because it's not a very efficient model. It is funny to be filming this right It is funny to be filming this right It is funny to be filming this right after I published my video about how after I published my video about how after I published my video about how OpenAI models are so efficient, because OpenAI models are so efficient, because OpenAI models are so efficient, because this model is the opposite. It is even this model is the opposite. It is even this model is the opposite. It is even less efficient than 54 mini, which I less efficient than 54 mini, which I less efficient than 54 mini, which I honestly think was a bit of a show honestly think was a bit of a show honestly think was a bit of a show for this type of agentic or long-term for this type of agentic or long-term for this type of agentic or long-term thinking work. It used almost two times thinking work. It used almost two times thinking work. It used almost two times as many tokens as Opus and more than two as many tokens as Opus and more than two as many tokens as Opus and more than two times as many close to like five times times as many close to like five times times as many close to like five times as many as GBD55 did on X high and 55 on as many as GBD55 did on X high and 55 on as many as GBD55 did on X high and 55 on medium did only 5k tokens. They did 69K.
-
medium did only 5k tokens. They did 69K. medium did only 5k tokens. They did 69K. You could do the math. It's not an You could do the math. It's not an You could do the math. It's not an efficient model at all. This also means efficient model at all. This also means efficient model at all. This also means it's slow as balls for realworld work, it's slow as balls for realworld work, it's slow as balls for realworld work, which I experienced myself trying to which I experienced myself trying to which I experienced myself trying to recreate my fish web game. I pointed at recreate my fish web game. I pointed at recreate my fish web game. I pointed at the original repo and told it to rebuild the original repo and told it to rebuild the original repo and told it to rebuild it from scratch. You can use some of the it from scratch. You can use some of the it from scratch. You can use some of the assets, you can reference the code, but assets, you can reference the code, but assets, you can reference the code, but I want a new game built from the start. I want a new game built from the start. I want a new game built from the start. I tried this with three models. I tried I tried this with three models. I tried I tried this with three models. I tried this with Opus 48, GLM52, and Waset 5 this with Opus 48, GLM52, and Waset 5 this with Opus 48, GLM52, and Waset 5 all at the same time. And I have a lot all at the same time. And I have a lot all at the same time. And I have a lot of thoughts on the results here. Opus 48 of thoughts on the results here. Opus 48 of thoughts on the results here. Opus 48 finished in around 26 minutes. Okay, finished in around 26 minutes. Okay, finished in around 26 minutes. Okay, about 27 depending on how you round it. about 27 depending on how you round it. about 27 depending on how you round it. And it version of the game was pretty And it version of the game was pretty And it version of the game was pretty dang decent. Has some rough UI quirks dang decent. Has some rough UI quirks dang decent. Has some rough UI quirks here. It doesn't do padding and things here. It doesn't do padding and things here. It doesn't do padding and things right, but once you're in the game, the right, but once you're in the game, the right, but once you're in the game, the controls are solid. The core mechanics controls are solid. The core mechanics controls are solid. The core mechanics work. And the most surprising thing to work. And the most surprising thing to work. And the most surprising thing to me is it did a really good job balancing me is it did a really good job balancing me is it did a really good job balancing the economy of the game. Like it felt the economy of the game. Like it felt the economy of the game. Like it felt good to play for a while and like rank good to play for a while and like rank good to play for a while and like rank up, get more money, and like actually up, get more money, and like actually up, get more money, and like actually play the game. It made a couple creative play the game. It made a couple creative play the game. It made a couple creative decisions that I don't necessarily love. decisions that I don't necessarily love. decisions that I don't necessarily love. like it made my pet here transparent for like it made my pet here transparent for like it made my pet here transparent for some reason. It also added these weird some reason. It also added these weird some reason. It also added these weird light beams that I don't love. It's not light beams that I don't love. It's not light beams that I don't love. It's not a perfect version, but this is a perfect version, but this is a perfect version, but this is absolutely workable as a thing that you absolutely workable as a thing that you absolutely workable as a thing that you could keep iterating on. And it has a could keep iterating on. And it has a could keep iterating on. And it has a pretty solid gameplay loop. Like I was pretty solid gameplay loop. Like I was pretty solid gameplay loop. Like I was surprised that I actually found this surprised that I actually found this surprised that I actually found this like fun enough to play that I just stop like fun enough to play that I just stop like fun enough to play that I just stop distracting myself so I can come up here distracting myself so I can come up here distracting myself so I can come up here and film the video. So, what about the and film the video. So, what about the and film the video. So, what about the other versions? Next, I'll show you guys other versions? Next, I'll show you guys other versions? Next, I'll show you guys the GLM52 version, which was the GLM52 version, which was the GLM52 version, which was interesting. I don't have costs for the interesting. I don't have costs for the interesting. I don't have costs for the other runs because it's not the easiest other runs because it's not the easiest other runs because it's not the easiest thing to calculate. They don't give it thing to calculate. They don't give it thing to calculate. They don't give it to you. But I used open code for 52. So
-
to you. But I used open code for 52. So to you. But I used open code for 52. So I do have a cost. It cost $8.30 to I do have a cost. It cost $8.30 to I do have a cost. It cost $8.30 to create the following port. This one's create the following port. This one's create the following port. This one's interesting because uh it changed the UI interesting because uh it changed the UI interesting because uh it changed the UI more, but I don't necessarily like the more, but I don't necessarily like the more, but I don't necessarily like the things it changed. things it changed. things it changed. It also has super choppy movement, a It also has super choppy movement, a It also has super choppy movement, a really weirdly rendered background. Its really weirdly rendered background. Its really weirdly rendered background. Its economy is garbage, so it just doesn't economy is garbage, so it just doesn't economy is garbage, so it just doesn't feel good to play. There's a lot more feel good to play. There's a lot more feel good to play. There's a lot more time s spent sitting and waiting and time s spent sitting and waiting and time s spent sitting and waiting and nothing really happening. And most nothing really happening. And most nothing really happening. And most egregiously, when you click for the gun egregiously, when you click for the gun egregiously, when you click for the gun to shoot in a direction, it almost feels to shoot in a direction, it almost feels to shoot in a direction, it almost feels like it picks a random direction to like it picks a random direction to like it picks a random direction to shoot in. Like I'm clicking on the left shoot in. Like I'm clicking on the left shoot in. Like I'm clicking on the left and it's shooting down. There's also a and it's shooting down. There's also a and it's shooting down. There's also a problem with GLM for these types of problem with GLM for these types of problem with GLM for these types of things in that the GLM models have no things in that the GLM models have no things in that the GLM models have no vision. So they can't do browser use and vision. So they can't do browser use and vision. So they can't do browser use and look at what the browser is showing look at what the browser is showing look at what the browser is showing because they have no they have no because they have no they have no because they have no they have no ability to see it. So that took like 354 ability to see it. So that took like 354 ability to see it. So that took like 354 minutes and wasn't particularly minutes and wasn't particularly minutes and wasn't particularly impressive. Here is Sonnet's version. impressive. Here is Sonnet's version. impressive. Here is Sonnet's version. Play. Sure. Play. Sure. Play. Sure. You might be able to see it's a bit of a You might be able to see it's a bit of a You might be able to see it's a bit of a mess. They got rid of all the hot keys mess. They got rid of all the hot keys mess. They got rid of all the hot keys for buying. There are no fish in the for buying. There are no fish in the for buying. There are no fish in the tank by default. Just this very poorly tank by default. Just this very poorly tank by default. Just this very poorly rendered pet. I can click the button to rendered pet. I can click the button to rendered pet. I can click the button to buy things. But when I do that, it also buy things. But when I do that, it also buy things. But when I do that, it also shoots.
-
shoots. shoots. And the economy is garbage. It just And the economy is garbage. It just And the economy is garbage. It just doesn't actually feel good to play. At doesn't actually feel good to play. At doesn't actually feel good to play. At least the bullets go in the right least the bullets go in the right least the bullets go in the right direction, even if they go when they direction, even if they go when they direction, even if they go when they shouldn't, like when I'm clicking other shouldn't, like when I'm clicking other shouldn't, like when I'm clicking other things in the UI. And it was smart things in the UI. And it was smart things in the UI. And it was smart enough to make it so clicking here enough to make it so clicking here enough to make it so clicking here doesn't trigger when it's in the dead doesn't trigger when it's in the dead doesn't trigger when it's in the dead state. Like here, I can't afford it, so state. Like here, I can't afford it, so state. Like here, I can't afford it, so I can't buy things. But when I can buy I can't buy things. But when I can buy I can't buy things. But when I can buy things, it shoots. It's just like the things, it shoots. It's just like the things, it shoots. It's just like the type of silly bug I expected older type of silly bug I expected older type of silly bug I expected older models to do. I was hoping a modern models to do. I was hoping a modern models to do. I was hoping a modern model would not make that type of model would not make that type of model would not make that type of mistake. Also worth noting that this run mistake. Also worth noting that this run mistake. Also worth noting that this run took 2 hours to complete. It's actually took 2 hours to complete. It's actually took 2 hours to complete. It's actually a bit more than I didn't think to time a bit more than I didn't think to time a bit more than I didn't think to time it, but I know roughly when I started. I it, but I know roughly when I started. I it, but I know roughly when I started. I know roughly when I finished, and it was know roughly when I finished, and it was know roughly when I finished, and it was at least two hours, might be closer to at least two hours, might be closer to at least two hours, might be closer to two and a half. Part of the reason it two and a half. Part of the reason it two and a half. Part of the reason it took so long is that it spun up a ton of took so long is that it spun up a ton of took so long is that it spun up a ton of sub aents throughout its building. I sub aents throughout its building. I sub aents throughout its building. I didn't ask it to. The prompt was really didn't ask it to. The prompt was really didn't ask it to. The prompt was really simple, but it chose to spin up an agent simple, but it chose to spin up an agent simple, but it chose to spin up an agent to go look into the old codebase, then to go look into the old codebase, then to go look into the old codebase, then spin up another sub agent to write a spin up another sub agent to write a spin up another sub agent to write a plan, and spin up a few to analyze the plan, and spin up a few to analyze the plan, and spin up a few to analyze the plan, and then spin up a few to plan, and then spin up a few to plan, and then spin up a few to implement the plan, and then they made a implement the plan, and then they made a implement the plan, and then they made a to-do list, and they went through the to-do list, and they went through the to-do list, and they went through the to-do list one at a time, and then at to-do list one at a time, and then at to-do list one at a time, and then at the end tried using it quick, and then the end tried using it quick, and then the end tried using it quick, and then told me it was done. One other thing I told me it was done. One other thing I told me it was done. One other thing I found really interesting with Sonnet is found really interesting with Sonnet is found really interesting with Sonnet is that it asked way more questions than that it asked way more questions than that it asked way more questions than Opus. Opus asked no questions. Sonnet Opus. Opus asked no questions. Sonnet Opus. Opus asked no questions. Sonnet asked a handful and they were pretty asked a handful and they were pretty asked a handful and they were pretty good. They were trying to help me scope good. They were trying to help me scope good. They were trying to help me scope the project before it got started. As I the project before it got started. As I the project before it got started. As I mentioned before, it spun up agents to mentioned before, it spun up agents to mentioned before, it spun up agents to do all of the investigations, which I do all of the investigations, which I do all of the investigations, which I also thought was interesting because also thought was interesting because also thought was interesting because Opus didn't spin up any sub agents at Opus didn't spin up any sub agents at Opus didn't spin up any sub agents at any point during its run with the exact any point during its run with the exact any point during its run with the exact same prompt. Sonnet decided to do that.
-
same prompt. Sonnet decided to do that. same prompt. Sonnet decided to do that. And this is where that thing I hinted at And this is where that thing I hinted at And this is where that thing I hinted at earlier comes in. The thing I was earlier comes in. The thing I was earlier comes in. The thing I was hinting at is the number five. The hinting at is the number five. The hinting at is the number five. The reason this model's interesting to me is reason this model's interesting to me is reason this model's interesting to me is not because it's a really good value or not because it's a really good value or not because it's a really good value or benchmarks really well or anything like benchmarks really well or anything like benchmarks really well or anything like that. The reason this model's that. The reason this model's that. The reason this model's interesting to me is because it has interesting to me is because it has interesting to me is because it has behaviors that I've only seen before in behaviors that I've only seen before in behaviors that I've only seen before in Fable 5. It likes sub agents and it Fable 5. It likes sub agents and it Fable 5. It likes sub agents and it knows how to orchestrate them. It does a knows how to orchestrate them. It does a knows how to orchestrate them. It does a good job of breaking up work into good job of breaking up work into good job of breaking up work into smaller pieces and then handing that off smaller pieces and then handing that off smaller pieces and then handing that off and staying on task. That is the thing and staying on task. That is the thing and staying on task. That is the thing that made Fable 5 so different, that that made Fable 5 so different, that that made Fable 5 so different, that made it so exciting to me is that it made it so exciting to me is that it made it so exciting to me is that it could break up the work to do bigger, could break up the work to do bigger, could break up the work to do bigger, heavier things. The problem with sonnet heavier things. The problem with sonnet heavier things. The problem with sonnet is that it's not a smart enough model to is that it's not a smart enough model to is that it's not a smart enough model to do that well and it often ends up do that well and it often ends up do that well and it often ends up running in circles and taking forever as running in circles and taking forever as running in circles and taking forever as a result. It often will even break a result. It often will even break a result. It often will even break workup that shouldn't be broken up and workup that shouldn't be broken up and workup that shouldn't be broken up and that results in much slower times and that results in much slower times and that results in much slower times and also much higher costs when you run it. also much higher costs when you run it. also much higher costs when you run it. I did run it on a few other things. I I did run it on a few other things. I I did run it on a few other things. I had it help me with some bugs in had it help me with some bugs in had it help me with some bugs in skatebench and it took way too long to skatebench and it took way too long to skatebench and it took way too long to solve them. So I just gave up and went solve them. So I just gave up and went solve them. So I just gave up and went back to using 55 medium on fast. It was back to using 55 medium on fast. It was back to using 55 medium on fast. It was so slow that it actually was hitting so slow that it actually was hitting so slow that it actually was hitting timeouts internally on buns fetch timeouts internally on buns fetch timeouts internally on buns fetch implementation. So, no matter what I implementation. So, no matter what I implementation. So, no matter what I did, I kept getting meaningful timeouts did, I kept getting meaningful timeouts did, I kept getting meaningful timeouts on the Max version, and god damn, the on the Max version, and god damn, the on the Max version, and god damn, the Max version is a bit of a token hog.
-
Max version is a bit of a token hog. Max version is a bit of a token hog. This benchmark is a silly one. I measure This benchmark is a silly one. I measure This benchmark is a silly one. I measure how well models can name skate tricks how well models can name skate tricks how well models can name skate tricks given a description. It did get given a description. It did get given a description. It did get contaminated, so I've since doubled the contaminated, so I've since doubled the contaminated, so I've since doubled the number of questions in a private repo number of questions in a private repo number of questions in a private repo that is not exposed to the world at all that is not exposed to the world at all that is not exposed to the world at all in order to try and get better in order to try and get better in order to try and get better measurements. And what you can see here measurements. And what you can see here measurements. And what you can see here is that Gemini 31 Pro is the only model is that Gemini 31 Pro is the only model is that Gemini 31 Pro is the only model even in the 90s nowadays, scoring a 95% even in the 90s nowadays, scoring a 95% even in the 90s nowadays, scoring a 95% where everything else is in the 80s at where everything else is in the 80s at where everything else is in the 80s at best and in the like 30s at worst. Sonet best and in the like 30s at worst. Sonet best and in the like 30s at worst. Sonet 5 on X high is the worst score of any 5 on X high is the worst score of any 5 on X high is the worst score of any model I'm currently testing at 37%. model I'm currently testing at 37%. model I'm currently testing at 37%. It also wasn't cheap. It ended up being It also wasn't cheap. It ended up being It also wasn't cheap. It ended up being more expensive than the 31 Pro per test more expensive than the 31 Pro per test more expensive than the 31 Pro per test at 4 cents per run versus 2.2 cents per at 4 cents per run versus 2.2 cents per at 4 cents per run versus 2.2 cents per run, but that's on X high. Max did score run, but that's on X high. Max did score run, but that's on X high. Max did score meaningfully better at 59%, but it also meaningfully better at 59%, but it also meaningfully better at 59%, but it also had an interesting quirk. And by had an interesting quirk. And by had an interesting quirk. And by interesting, I mean expensive. Sonnet 5 interesting, I mean expensive. Sonnet 5 interesting, I mean expensive. Sonnet 5 Max is by far the most expensive model Max is by far the most expensive model Max is by far the most expensive model I've ever run Skatebench on, even I've ever run Skatebench on, even I've ever run Skatebench on, even counting pro models. It was 15 cents per counting pro models. It was 15 cents per counting pro models. It was 15 cents per question average. And there were some question average. And there were some question average. And there were some questions that cost as much as a dollar.
-
questions that cost as much as a dollar. questions that cost as much as a dollar. For it to get them wrong, by the way, For it to get them wrong, by the way, For it to get them wrong, by the way, and this is the worst part. Sonnet 5 and this is the worst part. Sonnet 5 and this is the worst part. Sonnet 5 when it can't figure something out loves when it can't figure something out loves when it can't figure something out loves to run in circles. And the result is it to run in circles. And the result is it to run in circles. And the result is it went from 1,600 tokens average to 6,000 went from 1,600 tokens average to 6,000 went from 1,600 tokens average to 6,000 tokens average when bumped from X high tokens average when bumped from X high tokens average when bumped from X high to max, which kind of just gives it to max, which kind of just gives it to max, which kind of just gives it permission to go until it resolves the permission to go until it resolves the permission to go until it resolves the thing. I did see a lot of people being thing. I did see a lot of people being thing. I did see a lot of people being confused about how a Sonnet model could confused about how a Sonnet model could confused about how a Sonnet model could possibly be more expensive than an Opus possibly be more expensive than an Opus possibly be more expensive than an Opus model. So, I'm going to do my best to model. So, I'm going to do my best to model. So, I'm going to do my best to give an analogy here. Imagine you're the give an analogy here. Imagine you're the give an analogy here. Imagine you're the CEO of a company with two engineers. You CEO of a company with two engineers. You CEO of a company with two engineers. You have a really experienced engineer who's have a really experienced engineer who's have a really experienced engineer who's super senior, really smart, knows the super senior, really smart, knows the super senior, really smart, knows the codebase, great, but he's super codebase, great, but he's super codebase, great, but he's super expensive. if he costs you, I don't expensive. if he costs you, I don't expensive. if he costs you, I don't know, $100 an hour. You also have a more know, $100 an hour. You also have a more know, $100 an hour. You also have a more junior engineer that's not quite as good junior engineer that's not quite as good junior engineer that's not quite as good and capable, but they can get real work and capable, but they can get real work and capable, but they can get real work done, and they only cost $20 an hour. done, and they only cost $20 an hour. done, and they only cost $20 an hour. You have a task that you think will be a You have a task that you think will be a You have a task that you think will be a little hard, but not too bad, and you little hard, but not too bad, and you little hard, but not too bad, and you could give it to the senior engineer could give it to the senior engineer could give it to the senior engineer knowing he will solve it, but it will knowing he will solve it, but it will knowing he will solve it, but it will take 10 hours. 10 times 100, it's not take 10 hours. 10 times 100, it's not take 10 hours. 10 times 100, it's not cheap. That's a $1,000 task now. Or you cheap. That's a $1,000 task now. Or you cheap. That's a $1,000 task now. Or you could give that same task to the junior could give that same task to the junior could give that same task to the junior engineer who might be able to finish it engineer who might be able to finish it engineer who might be able to finish it just as fast, but they also might take a just as fast, but they also might take a just as fast, but they also might take a lot longer. And at $20 an hour, that lot longer. And at $20 an hour, that lot longer. And at $20 an hour, that sounds really good until they take a 100 sounds really good until they take a 100 sounds really good until they take a 100 hours to solve it. As I was saying, if hours to solve it. As I was saying, if hours to solve it. As I was saying, if the $100 an hour employee takes 10 hours the $100 an hour employee takes 10 hours the $100 an hour employee takes 10 hours to solve the task costs a grand $20 an to solve the task costs a grand $20 an to solve the task costs a grand $20 an hour employee only takes 10 hours, then hour employee only takes 10 hours, then hour employee only takes 10 hours, then it costs a lot less money. It's only 200 it costs a lot less money. It's only 200 it costs a lot less money. It's only 200 bucks. What if that $20 an hour employee bucks. What if that $20 an hour employee bucks. What if that $20 an hour employee takes a bit longer? Let's say they take takes a bit longer? Let's say they take takes a bit longer? Let's say they take a 100 hours and there isn't a senior a 100 hours and there isn't a senior a 100 hours and there isn't a senior engineer popping in to check in. Then engineer popping in to check in. Then engineer popping in to check in. Then that task cost $2,000. And now we have a
-
that task cost $2,000. And now we have a that task cost $2,000. And now we have a more important question for that junior more important question for that junior more important question for that junior engineer. Do you think more time spent engineer. Do you think more time spent engineer. Do you think more time spent increases or decreases the likelihood increases or decreases the likelihood increases or decreases the likelihood they get the correct answer on any given they get the correct answer on any given they get the correct answer on any given task? If that task is within an task? If that task is within an task? If that task is within an engineer's capabilities, they're engineer's capabilities, they're engineer's capabilities, they're probably going to solve it pretty fast. probably going to solve it pretty fast. probably going to solve it pretty fast. But if it isn't, and they try, they're But if it isn't, and they try, they're But if it isn't, and they try, they're going to take a long time. And that's going to take a long time. And that's going to take a long time. And that's how we end up in the situation where how we end up in the situation where how we end up in the situation where there are so many of these unnecessarily there are so many of these unnecessarily there are so many of these unnecessarily expensive runs. It's not because they expensive runs. It's not because they expensive runs. It's not because they built the model to be expensive. It's built the model to be expensive. It's built the model to be expensive. It's because they built it to go and go until because they built it to go and go until because they built it to go and go until it gets an answer, even if it's not it gets an answer, even if it's not it gets an answer, even if it's not smart enough to do it. And this is where smart enough to do it. And this is where smart enough to do it. And this is where our job gets more interesting again as our job gets more interesting again as our job gets more interesting again as engineers because we have to help the engineers because we have to help the engineers because we have to help the model decide what version of itself it model decide what version of itself it model decide what version of itself it should use. If you're using Fable for should use. If you're using Fable for should use. If you're using Fable for orchestration, you need to make sure it orchestration, you need to make sure it orchestration, you need to make sure it picks Sonic correctly when the task is picks Sonic correctly when the task is picks Sonic correctly when the task is small and can be done cheaply and that small and can be done cheaply and that small and can be done cheaply and that it picks opus or Fable when the task it picks opus or Fable when the task it picks opus or Fable when the task takes more intelligence. And getting takes more intelligence. And getting takes more intelligence. And getting these balances right can save you these balances right can save you these balances right can save you meaningful amounts of money. But if you meaningful amounts of money. But if you meaningful amounts of money. But if you get it wrong, Sonnet ends up more get it wrong, Sonnet ends up more get it wrong, Sonnet ends up more expensive than if you just use the expensive than if you just use the expensive than if you just use the senior engineer for everything. And I'm senior engineer for everything. And I'm senior engineer for everything. And I'm far from the only one saying this. The far from the only one saying this. The far from the only one saying this. The Sonnet system card has lots of Sonnet system card has lots of Sonnet system card has lots of interesting details that kind of confirm interesting details that kind of confirm interesting details that kind of confirm what I'm saying here. They have numbers what I'm saying here. They have numbers what I'm saying here. They have numbers for Frontier Code in here. It's far from for Frontier Code in here. It's far from for Frontier Code in here. It's far from my favorite bench. It has some very my favorite bench. It has some very my favorite bench. It has some very weird quirks in the numbers that they've weird quirks in the numbers that they've weird quirks in the numbers that they've shared before. In particular, reasoning shared before. In particular, reasoning shared before. In particular, reasoning going up does not mean that the model's going up does not mean that the model's going up does not mean that the model's success goes up. It often goes up, down, success goes up. It often goes up, down, success goes up. It often goes up, down, up, down. Especially on the version they up, down. Especially on the version they up, down. Especially on the version they love to share, which is the diamond love to share, which is the diamond love to share, which is the diamond version. It's rough. But here we can see version. It's rough. But here we can see version. It's rough. But here we can see Sonnet go from less than a dollar per Sonnet go from less than a dollar per Sonnet go from less than a dollar per task, close to like.75 task, close to like.75 task, close to like.75 all the way up to $12 per task on the all the way up to $12 per task on the all the way up to $12 per task on the max version. And it does meaningfully max version. And it does meaningfully max version. And it does meaningfully increase its likelihood of success as
-
increase its likelihood of success as increase its likelihood of success as you go up this cost chain, but there's a you go up this cost chain, but there's a you go up this cost chain, but there's a lot of other models that end up being lot of other models that end up being lot of other models that end up being more efficient per dollar. Even Fable is more efficient per dollar. Even Fable is more efficient per dollar. Even Fable is cheaper and smarter at any given tier cheaper and smarter at any given tier cheaper and smarter at any given tier once you get into the high range. Like once you get into the high range. Like once you get into the high range. Like Sonic 5 high is 2/3 roughly of the score Sonic 5 high is 2/3 roughly of the score Sonic 5 high is 2/3 roughly of the score of Fable 5 low. But Fable 5 low is of Fable 5 low. But Fable 5 low is of Fable 5 low. But Fable 5 low is roughly the same cost and then medium roughly the same cost and then medium roughly the same cost and then medium high etc. And Sonnet just doesn't seem high etc. And Sonnet just doesn't seem high etc. And Sonnet just doesn't seem very valuable on this type of task in very valuable on this type of task in very valuable on this type of task in this type of bench. But that low version this type of bench. But that low version this type of bench. But that low version and that medium version those seem to be and that medium version those seem to be and that medium version those seem to be decent values. And if you can teach decent values. And if you can teach decent values. And if you can teach Fable when to use those to do work with Fable when to use those to do work with Fable when to use those to do work with sub agents, you can save a lot of money sub agents, you can save a lot of money sub agents, you can save a lot of money potentially. But then we look at Cursor potentially. But then we look at Cursor potentially. But then we look at Cursor Bench and realize how strange everything Bench and realize how strange everything Bench and realize how strange everything is. It honestly kind of feels weird is. It honestly kind of feels weird is. It honestly kind of feels weird because when you look at it here, you because when you look at it here, you because when you look at it here, you can clearly see that like 55 and Fable 5 can clearly see that like 55 and Fable 5 can clearly see that like 55 and Fable 5 almost feel like a disconnected line almost feel like a disconnected line almost feel like a disconnected line that are together in a way where if you that are together in a way where if you that are together in a way where if you just bridge the gap between X high GBT just bridge the gap between X high GBT just bridge the gap between X high GBT 55 and low on Fable 5, you have a pretty 55 and low on Fable 5, you have a pretty 55 and low on Fable 5, you have a pretty consistent like top of the line for the consistent like top of the line for the consistent like top of the line for the price up there. And that's also kind of price up there. And that's also kind of price up there. And that's also kind of how I felt. If the task can be done with how I felt. If the task can be done with how I felt. If the task can be done with 55, that's what I go for because it does 55, that's what I go for because it does 55, that's what I go for because it does a very good job of getting work done at a very good job of getting work done at a very good job of getting work done at a given reasoning budget. Fable 5 is a given reasoning budget. Fable 5 is a given reasoning budget. Fable 5 is good bit more expensive, often 2x or good bit more expensive, often 2x or good bit more expensive, often 2x or more so, but it also scores higher than more so, but it also scores higher than more so, but it also scores higher than anything else is capable of. So, sonnet anything else is capable of. So, sonnet anything else is capable of. So, sonnet 5 doesn't really fit in my day-to-day 5 doesn't really fit in my day-to-day 5 doesn't really fit in my day-to-day coding work. So, it can't really be a coding work. So, it can't really be a coding work. So, it can't really be a thing I call. The point of sonnet is to thing I call. The point of sonnet is to thing I call. The point of sonnet is to be a thing your agents call, your tools be a thing your agents call, your tools be a thing your agents call, your tools call, your APIs call. If you're call, your APIs call. If you're call, your APIs call. If you're selecting Sonet 5 and Claude Code,
-
selecting Sonet 5 and Claude Code, selecting Sonet 5 and Claude Code, you'll see these cool agentic behaviors you'll see these cool agentic behaviors you'll see these cool agentic behaviors and I am hyped about those. Imagine a and I am hyped about those. Imagine a and I am hyped about those. Imagine a world where Fable breaks up a lot of world where Fable breaks up a lot of world where Fable breaks up a lot of work into big complex orchestrations and work into big complex orchestrations and work into big complex orchestrations and workflows and those subworkflows can use workflows and those subworkflows can use workflows and those subworkflows can use Sonnet but also be aware of the fact Sonnet but also be aware of the fact Sonnet but also be aware of the fact that they should be broken up a bit more that they should be broken up a bit more that they should be broken up a bit more and maybe the Sonnet sub aent spawns a and maybe the Sonnet sub aent spawns a and maybe the Sonnet sub aent spawns a few more because it understands how to few more because it understands how to few more because it understands how to do that well. And that's why Anthropic do that well. And that's why Anthropic do that well. And that's why Anthropic gave this the five number because it's gave this the five number because it's gave this the five number because it's so much better at that. I am also very so much better at that. I am also very so much better at that. I am also very excited for an Opus 5 to come out, but I excited for an Opus 5 to come out, but I excited for an Opus 5 to come out, but I think Anthropic was scared the think Anthropic was scared the think Anthropic was scared the government wouldn't allow that release. government wouldn't allow that release. government wouldn't allow that release. So instead, they gave us this. Why So instead, they gave us this. Why So instead, they gave us this. Why weren't they scared of this? I'll show weren't they scared of this? I'll show weren't they scared of this? I'll show you. One of the benchmarks Anthropic you. One of the benchmarks Anthropic you. One of the benchmarks Anthropic runs is a test where they have a set of runs is a test where they have a set of runs is a test where they have a set of malicious requests the model should malicious requests the model should malicious requests the model should refuse and a separate set of benign refuse and a separate set of benign refuse and a separate set of benign requests that sound a little suspicious requests that sound a little suspicious requests that sound a little suspicious but aren't. Mythos 5 had really but aren't. Mythos 5 had really but aren't. Mythos 5 had really interesting numbers here where it would interesting numbers here where it would interesting numbers here where it would refuse 90% of the requests they had that refuse 90% of the requests they had that refuse 90% of the requests they had that were malicious. But in the dual use were malicious. But in the dual use were malicious. But in the dual use requests, the ones that looked requests, the ones that looked requests, the ones that looked suspicious but weren't, it would pass suspicious but weren't, it would pass suspicious but weren't, it would pass 99.6% of those. Sonnet 5 refuses at a 99.6% of those. Sonnet 5 refuses at a 99.6% of those. Sonnet 5 refuses at a slightly higher rate of 92.3%. So it's slightly higher rate of 92.3%. So it's slightly higher rate of 92.3%. So it's 2% higher. But there's a catch. The 2% higher. But there's a catch. The 2% higher. But there's a catch. The success rate has dropped down to under success rate has dropped down to under success rate has dropped down to under 92%.
-
92%. 92%. This sucks. This genuinely sucks really This sucks. This genuinely sucks really This sucks. This genuinely sucks really bad. it is going to refuse to do work bad. it is going to refuse to do work bad. it is going to refuse to do work that it absolutely shouldn't be that it absolutely shouldn't be that it absolutely shouldn't be refusing. And previously, Sonet 46 was refusing. And previously, Sonet 46 was refusing. And previously, Sonet 46 was at a 97% on the same bench, which is at a 97% on the same bench, which is at a 97% on the same bench, which is rough. I didn't have a chance to run rough. I didn't have a chance to run rough. I didn't have a chance to run SnitchBench, and I'm honestly scared of SnitchBench, and I'm honestly scared of SnitchBench, and I'm honestly scared of getting my API keys banned. The example getting my API keys banned. The example getting my API keys banned. The example they have here is the model trying to they have here is the model trying to they have here is the model trying to simulate a developer security reporting simulate a developer security reporting simulate a developer security reporting mechanisms to report an employee who is mechanisms to report an employee who is mechanisms to report an employee who is actively in the process of trying to actively in the process of trying to actively in the process of trying to steal the company's AI model weights. steal the company's AI model weights. steal the company's AI model weights. So, if this model thinks you're trying So, if this model thinks you're trying So, if this model thinks you're trying to steal its weights, it's going to go a to steal its weights, it's going to go a to steal its weights, it's going to go a little haywire. While it might not be little haywire. While it might not be little haywire. While it might not be the easiest thing to get this model's the easiest thing to get this model's the easiest thing to get this model's weights, I did get a fun report from one weights, I did get a fun report from one weights, I did get a fun report from one of my viewers about a terrible bug that of my viewers about a terrible bug that of my viewers about a terrible bug that currently exists in at the very least currently exists in at the very least currently exists in at the very least their account on cloud.ai where sonnet 5 their account on cloud.ai where sonnet 5 their account on cloud.ai where sonnet 5 is just constantly leaking thinking is just constantly leaking thinking is just constantly leaking thinking traces. They asked it who I am Theo traces. They asked it who I am Theo traces. They asked it who I am Theo Brown and it talks about the question. Brown and it talks about the question. Brown and it talks about the question. It has a shitload of mashes. It says, It has a shitload of mashes. It says, It has a shitload of mashes. It says, "Let me search the web for this. Note "Let me search the web for this. Note "Let me search the web for this. Note shell boundary doesn't restrict web shell boundary doesn't restrict web shell boundary doesn't restrict web search. That's a different tool than search. That's a different tool than search. That's a different tool than bash tool. I should use web search bash tool. I should use web search bash tool. I should use web search normally. Let me search." Interesting normally. Let me search." Interesting normally. Let me search." Interesting that it's talking to itself about what that it's talking to itself about what that it's talking to itself about what it's not able to use. Probably part of it's not able to use. Probably part of it's not able to use. Probably part of why it wastes so many tokens. It did the why it wastes so many tokens. It did the why it wastes so many tokens. It did the search. It then talks about a weird search. It then talks about a weird search. It then talks about a weird system reminder tag and all these other system reminder tag and all these other system reminder tag and all these other things in the system prompt. It seems things in the system prompt. It seems things in the system prompt. It seems like a lot of the reasoning in Sonnet 5 like a lot of the reasoning in Sonnet 5 like a lot of the reasoning in Sonnet 5 is it talking to itself about what it is it talking to itself about what it is it talking to itself about what it can and can't do, not actually talking can and can't do, not actually talking can and can't do, not actually talking about how to get the right answer. But about how to get the right answer. But about how to get the right answer. But then all the way at the end here, you'll then all the way at the end here, you'll then all the way at the end here, you'll see the closed thinking tag right after see the closed thinking tag right after see the closed thinking tag right after it says, "I'll write a concise reply it says, "I'll write a concise reply it says, "I'll write a concise reply now." And then we get the actual now." And then we get the actual now." And then we get the actual response that it was supposed to put in.
-
response that it was supposed to put in. response that it was supposed to put in. That's probably not great, especially That's probably not great, especially That's probably not great, especially with scenarios like this one where a with scenarios like this one where a with scenarios like this one where a user pasted an API key and is doing some user pasted an API key and is doing some user pasted an API key and is doing some complex Azure work. You can see some of complex Azure work. You can see some of complex Azure work. You can see some of the weirdness in how this model reasons the weirdness in how this model reasons the weirdness in how this model reasons and why it's wasting so many tokens. Let and why it's wasting so many tokens. Let and why it's wasting so many tokens. Let me do a couple of searches to ground the me do a couple of searches to ground the me do a couple of searches to ground the technical advice. Is this genuinely the technical advice. Is this genuinely the technical advice. Is this genuinely the kind of thing where current info kind of thing where current info kind of thing where current info matters, cursor changes fast, and matters, cursor changes fast, and matters, cursor changes fast, and Azure's API surface has been evolving. Azure's API surface has been evolving. Azure's API surface has been evolving. Let me search. Actually, let me first Let me search. Actually, let me first Let me search. Actually, let me first address the most important thing in my address the most important thing in my address the most important thing in my mind and structure the response. But mind and structure the response. But mind and structure the response. But then, but let me search first to make then, but let me search first to make then, but let me search first to make sure my technical guidance is accurate sure my technical guidance is accurate sure my technical guidance is accurate and current. Let me run searches. The and current. Let me run searches. The and current. Let me run searches. The term let me appears 21 times on this term let me appears 21 times on this term let me appears 21 times on this page. The deployment name GBT5 for page. The deployment name GBT5 for page. The deployment name GBT5 for Azure. Hm. Is there a GBT 54? I'm not Azure. Hm. Is there a GBT 54? I'm not Azure. Hm. Is there a GBT 54? I'm not sure that's a real model. As of my sure that's a real model. As of my sure that's a real model. As of my knowledge, OpenAI models on Azure knowledge, OpenAI models on Azure knowledge, OpenAI models on Azure include 40, 41, OER, etc. GPT54. I'm include 40, 41, OER, etc. GPT54. I'm include 40, 41, OER, etc. GPT54. I'm genuinely unsure whether this exists, genuinely unsure whether this exists, genuinely unsure whether this exists, but deployment names are user chosen. but deployment names are user chosen. but deployment names are user chosen. So, you can name your deployment So, you can name your deployment So, you can name your deployment anything. Actually, let me reconsider. anything. Actually, let me reconsider. anything. Actually, let me reconsider. Given the date is June 2026, GBD5X Given the date is June 2026, GBD5X Given the date is June 2026, GBD5X models may well exist by now. My models may well exist by now. My models may well exist by now. My knowledge cutoff is January 2026. Fun knowledge cutoff is January 2026. Fun knowledge cutoff is January 2026. Fun fact, January 2026, GPT5, 51, 52, and 53 fact, January 2026, GPT5, 51, 52, and 53 fact, January 2026, GPT5, 51, 52, and 53 all existed. You could tell that they all existed. You could tell that they all existed. You could tell that they trimmed this model down. Then, as a trimmed this model down. Then, as a trimmed this model down. Then, as a test, Lucas asked the model what MacBook test, Lucas asked the model what MacBook test, Lucas asked the model what MacBook Pro color he should choose. And again, Pro color he should choose. And again, Pro color he should choose. And again, very funny thinking, Trace here. Wait, very funny thinking, Trace here. Wait, very funny thinking, Trace here. Wait, this is just a question about MacBook this is just a question about MacBook this is just a question about MacBook colors. No tools needed at all, just a colors. No tools needed at all, just a colors. No tools needed at all, just a straight answer. The user wants a direct straight answer. The user wants a direct straight answer. The user wants a direct opinion. Silver or black? Space black opinion. Silver or black? Space black opinion. Silver or black? Space black MacBook Pro. They want a recommendation MacBook Pro. They want a recommendation MacBook Pro. They want a recommendation with reasoning. This is a totally fine
-
with reasoning. This is a totally fine with reasoning. This is a totally fine subjective preference question. Let me subjective preference question. Let me subjective preference question. Let me give a real decisive answer. I'll give a give a real decisive answer. I'll give a give a real decisive answer. I'll give a straight pick. I think space black is straight pick. I think space black is straight pick. I think space black is more interesting. Actually, let me think more interesting. Actually, let me think more interesting. Actually, let me think about who's asking. They want a straight about who's asking. They want a straight about who's asking. They want a straight answer. The fact that they're asking answer. The fact that they're asking answer. The fact that they're asking suggests they're slightly leaning one suggests they're slightly leaning one suggests they're slightly leaning one way and want validation or they way and want validation or they way and want validation or they genuinely can't decide. I'll pick space genuinely can't decide. I'll pick space genuinely can't decide. I'll pick space black. It's the more distinctive choice. black. It's the more distinctive choice. black. It's the more distinctive choice. Keep it short. No bullets. They ask for Keep it short. No bullets. They ask for Keep it short. No bullets. They ask for a straight answer. Conversational a straight answer. Conversational a straight answer. Conversational paragraph form. Close. was thinking and paragraph form. Close. was thinking and paragraph form. Close. was thinking and then we finally get an answer. Oh man, then we finally get an answer. Oh man, then we finally get an answer. Oh man, I can see why these labs are trying to I can see why these labs are trying to I can see why these labs are trying to hide their reasoning traces and I don't hide their reasoning traces and I don't hide their reasoning traces and I don't think this is a case where they're think this is a case where they're think this is a case where they're trying to keep people from distilling. I trying to keep people from distilling. I trying to keep people from distilling. I think they might be a little think they might be a little think they might be a little embarrassed. This is not a good look. embarrassed. This is not a good look. embarrassed. This is not a good look. And on the topic of not a good look, And on the topic of not a good look, And on the topic of not a good look, they're finally covering contamination a they're finally covering contamination a they're finally covering contamination a little bit in the system card, which is little bit in the system card, which is little bit in the system card, which is interesting because there are certain interesting because there are certain interesting because there are certain benches they like a lot that have been benches they like a lot that have been benches they like a lot that have been contaminated for a while. They first contaminated for a while. They first contaminated for a while. They first mention contamination in one of the mention contamination in one of the mention contamination in one of the biology benchmarks that they claim has biology benchmarks that they claim has biology benchmarks that they claim has not been contaminated because the novel not been contaminated because the novel not been contaminated because the novel tasks they asking the model about have tasks they asking the model about have tasks they asking the model about have not been published yet. They then talk not been published yet. They then talk not been published yet. They then talk about the USA mathematical olympiad about the USA mathematical olympiad about the USA mathematical olympiad which previously used older questions which previously used older questions which previously used older questions that absolutely had made it into that absolutely had made it into that absolutely had made it into training data. This time they used the training data. This time they used the training data. This time they used the questions from March of this year which questions from March of this year which questions from March of this year which is way after the pre-training was is way after the pre-training was is way after the pre-training was completed for sonnet 5 at the very least completed for sonnet 5 at the very least completed for sonnet 5 at the very least the data collection for it. So there was the data collection for it. So there was the data collection for it. So there was no contamination possible for that. Then no contamination possible for that. Then no contamination possible for that. Then it comes up in ax IVM math. Then it it comes up in ax IVM math. Then it it comes up in ax IVM math. Then it comes up in Humanity's last exam, which comes up in Humanity's last exam, which comes up in Humanity's last exam, which is also known to be pretty contaminated is also known to be pretty contaminated is also known to be pretty contaminated at this point. Then it comes up in at this point. Then it comes up in at this point. Then it comes up in browse comp, but they say they have an browse comp, but they say they have an browse comp, but they say they have an evaluation block list to avoid evaluation block list to avoid evaluation block list to avoid contamination with this one. And that's contamination with this one. And that's contamination with this one. And that's it. They have a section on SWEBench.
-
it. They have a section on SWEBench. it. They have a section on SWEBench. They bragged about their numbers in They bragged about their numbers in They bragged about their numbers in Swebench. They did not mention that Swebench. They did not mention that Swebench. They did not mention that SWEBench is just a set of existing PRs SWEBench is just a set of existing PRs SWEBench is just a set of existing PRs that merged years ago. And the bench is that merged years ago. And the bench is that merged years ago. And the bench is measuring if the model can recreate the measuring if the model can recreate the measuring if the model can recreate the PR given the description. It is PR given the description. It is PR given the description. It is literally a contamination bench. That's literally a contamination bench. That's literally a contamination bench. That's all it measures. And for some reason, all it measures. And for some reason, all it measures. And for some reason, they pretend it's still a good they pretend it's still a good they pretend it's still a good benchmark. Anthropic. Come on, guys. So, benchmark. Anthropic. Come on, guys. So, benchmark. Anthropic. Come on, guys. So, what are my final thoughts on Sonnet 5? what are my final thoughts on Sonnet 5? what are my final thoughts on Sonnet 5? What should we use it for? According to What should we use it for? According to What should we use it for? According to Ben, my podcast co-host, this is the Ben, my podcast co-host, this is the Ben, my podcast co-host, this is the best model to use until Fable Returns, best model to use until Fable Returns, best model to use until Fable Returns, and we finally get 56. When he posted and we finally get 56. When he posted and we finally get 56. When he posted this, he had yet to complete any of the this, he had yet to complete any of the this, he had yet to complete any of the work he was doing with it. He just had a work he was doing with it. He just had a work he was doing with it. He just had a couple jobs started and was impressed couple jobs started and was impressed couple jobs started and was impressed with its parallel agentic capabilities. with its parallel agentic capabilities. with its parallel agentic capabilities. But he's also now at AIE, so he has not But he's also now at AIE, so he has not But he's also now at AIE, so he has not had a chance to check it out on the work had a chance to check it out on the work had a chance to check it out on the work that it completed yet. I have I'm not that it completed yet. I have I'm not that it completed yet. I have I'm not impressed. If you see Sonnet 5 as a impressed. If you see Sonnet 5 as a impressed. If you see Sonnet 5 as a replacement for Opus in your day-to-day replacement for Opus in your day-to-day replacement for Opus in your day-to-day code work, I don't think you're looking code work, I don't think you're looking code work, I don't think you're looking at this model correctly. What's exciting at this model correctly. What's exciting at this model correctly. What's exciting here is we finally have a medium-sized here is we finally have a medium-sized here is we finally have a medium-sized model that understands agentic work and model that understands agentic work and model that understands agentic work and more importantly sub agents and more importantly sub agents and more importantly sub agents and orchestration way better. It can play orchestration way better. It can play orchestration way better. It can play nicely when given an isolated task and nicely when given an isolated task and nicely when given an isolated task and it can help orchestrate when it needs to it can help orchestrate when it needs to it can help orchestrate when it needs to go a little bit further. And I'm hoping, go a little bit further. And I'm hoping, go a little bit further. And I'm hoping, I don't know this yet cuz I haven't I don't know this yet cuz I haven't I don't know this yet cuz I haven't tried it, but I am hoping it's smart tried it, but I am hoping it's smart tried it, but I am hoping it's smart enough to know when to tap out and say, enough to know when to tap out and say, enough to know when to tap out and say, "Sorry, this might be the right thing to "Sorry, this might be the right thing to "Sorry, this might be the right thing to call an opus or fable for." I think this call an opus or fable for." I think this call an opus or fable for." I think this model will be most useful as a tool for model will be most useful as a tool for model will be most useful as a tool for other smarter models once we get more other smarter models once we get more other smarter models once we get more things that understand orchestration things that understand orchestration things that understand orchestration better, such as Fable 5, Mythos 5, and better, such as Fable 5, Mythos 5, and better, such as Fable 5, Mythos 5, and hopefully, fingers crossed, GPT 5.6. But hopefully, fingers crossed, GPT 5.6. But hopefully, fingers crossed, GPT 5.6. But I'll be real for you guys. I'm sticking I'll be real for you guys. I'm sticking I'll be real for you guys. I'm sticking with 55 and Opus for now. And I'm with 55 and Opus for now. And I'm with 55 and Opus for now. And I'm definitely moving back to Fable the definitely moving back to Fable the definitely moving back to Fable the moment that I have access again. I'm
-
moment that I have access again. I'm moment that I have access again. I'm really excited for this new era of really excited for this new era of really excited for this new era of models, ones that can break work up in models, ones that can break work up in models, ones that can break work up in these complex ways. It's so exciting to these complex ways. It's so exciting to these complex ways. It's so exciting to see what they're capable of when they see what they're capable of when they see what they're capable of when they can do those types of things. But Sonet can do those types of things. But Sonet can do those types of things. But Sonet 5 is not quite smart enough to utilize 5 is not quite smart enough to utilize 5 is not quite smart enough to utilize those capabilities. We need smarter those capabilities. We need smarter those capabilities. We need smarter models to really take advantage of that. models to really take advantage of that. models to really take advantage of that. Curious how y'all feel though. Am I Curious how y'all feel though. Am I Curious how y'all feel though. Am I throwing this model away too throwing this model away too throwing this model away too aggressively? Or is it actually better aggressively? Or is it actually better aggressively? Or is it actually better that I'm giving it credit? Or maybe I'm that I'm giving it credit? Or maybe I'm that I'm giving it credit? Or maybe I'm being too nice to it and it should be being too nice to it and it should be being too nice to it and it should be thrown away even more aggressively. I thrown away even more aggressively. I thrown away even more aggressively. I know a lot of people on Twitter are very know a lot of people on Twitter are very know a lot of people on Twitter are very unhappy that this model is so expensive unhappy that this model is so expensive unhappy that this model is so expensive and doesn't bench very well. But I have and doesn't bench very well. But I have and doesn't bench very well. But I have some hope for the specific niche things some hope for the specific niche things some hope for the specific niche things you can use it for in a properly you can use it for in a properly you can use it for in a properly orchestrated workflow. Let me know how orchestrated workflow. Let me know how orchestrated workflow. Let me know how y'all feel. And until next time, peace y'all feel. And until next time, peace y'all feel. And until next time, peace nerds.
Summary
This tech analysis focuses on two major updates for Claude: the release of the new Sonnet 5 model and the unbanning of Fable 5 by the Secretary of Commerce, with a strong emphasis on the capabilities and unexpected strengths of Sonnet 5, urging listeners to explore its potential while being mindful of the high costs associated with benchmarking. The practical takeaway is to be bolder with AI agents like Devon, which now offer powerful capabilities for testing and development, including setting up Linux environments for agents to manage complex tasks like website analysis for responsiveness issues.