finally a good small model
Read full transcript 33 segments
-
Anthropic is pretty good at making big Anthropic is pretty good at making big models for coding. We've seen that with models for coding. We've seen that with models for coding. We've seen that with models like Fable and Opus and all the models like Fable and Opus and all the models like Fable and Opus and all the way back in the day with models like way back in the day with models like way back in the day with models like Sonnet 3.5 and the idea of tool calls Sonnet 3.5 and the idea of tool calls Sonnet 3.5 and the idea of tool calls kind of making us software engineers use kind of making us software engineers use kind of making us software engineers use AI the way that we do today. But there's AI the way that we do today. But there's AI the way that we do today. But there's always been a bit of a weakness with always been a bit of a weakness with always been a bit of a weakness with their lineup. It's the smaller models. their lineup. It's the smaller models. their lineup. It's the smaller models. And as cool as Haiku 4.5 was when it And as cool as Haiku 4.5 was when it And as cool as Haiku 4.5 was when it came out, it came out in October of last came out, it came out in October of last came out, it came out in October of last year and it was expensive at the time. year and it was expensive at the time. year and it was expensive at the time. Now it's just laughable. And I have put Now it's just laughable. And I have put Now it's just laughable. And I have put a lot of effort into adjusting all of my a lot of effort into adjusting all of my a lot of effort into adjusting all of my tooling, all of my agent and Claude MDs tooling, all of my agent and Claude MDs tooling, all of my agent and Claude MDs and all those things to make sure every and all those things to make sure every and all those things to make sure every tool in my system makes it very, very tool in my system makes it very, very tool in my system makes it very, very clear to all my agents that Haiku 4.5 clear to all my agents that Haiku 4.5 clear to all my agents that Haiku 4.5 should be avoided at all costs because should be avoided at all costs because should be avoided at all costs because it's useless. It's just not very good. it's useless. It's just not very good. it's useless. It's just not very good. That's why I made this post a few weeks That's why I made this post a few weeks That's why I made this post a few weeks ago where I said that Anthropic has no ago where I said that Anthropic has no ago where I said that Anthropic has no small models that are worth using right small models that are worth using right small models that are worth using right now and also dunked on OpenAI saying now and also dunked on OpenAI saying now and also dunked on OpenAI saying they have no large models worth using they have no large models worth using they have no large models worth using cuz at the time they had Astra and 6.0 cuz at the time they had Astra and 6.0 cuz at the time they had Astra and 6.0 Soul. They didn't have 6.1 Soul yet. And Soul. They didn't have 6.1 Soul yet. And Soul. They didn't have 6.1 Soul yet. And of course Google gets the fun jab at the of course Google gets the fun jab at the of course Google gets the fun jab at the end saying none of the Google models are end saying none of the Google models are end saying none of the Google models are worth using. But the thing I want to worth using. But the thing I want to worth using. But the thing I want to focus on here is this idea that focus on here is this idea that focus on here is this idea that Anthropic has no small models worth Anthropic has no small models worth Anthropic has no small models worth using. Historically, Anthropic puts using. Historically, Anthropic puts using. Historically, Anthropic puts almost all their effort into the biggest almost all their effort into the biggest almost all their effort into the biggest model and then poorly distills it into model and then poorly distills it into model and then poorly distills it into the cheaper ones. That's why when they the cheaper ones. That's why when they the cheaper ones. That's why when they had Mythos and Fable, the quality of had Mythos and Fable, the quality of had Mythos and Fable, the quality of Opus plummeted and we ended up with 4.6, Opus plummeted and we ended up with 4.6, Opus plummeted and we ended up with 4.6, 4.7, 4.8 and then eventually Opus 5 4.7, 4.8 and then eventually Opus 5 4.7, 4.8 and then eventually Opus 5 which was terrible. That was all because which was terrible. That was all because which was terrible. That was all because of their focus on Anthropic's large of their focus on Anthropic's large of their focus on Anthropic's large models. That curse has been broken models. That curse has been broken models. That curse has been broken though with both Opus 5.5 and Sonnet though with both Opus 5.5 and Sonnet though with both Opus 5.5 and Sonnet 5.5. Which is why I'm so hyped for 5.5. Which is why I'm so hyped for 5.5. Which is why I'm so hyped for today's introduction of Haiku 5.5.
-
today's introduction of Haiku 5.5. today's introduction of Haiku 5.5. They've been teasing this one for a bit They've been teasing this one for a bit They've been teasing this one for a bit and I I had high hopes but low and I I had high hopes but low and I I had high hopes but low expectations because this is one of the expectations because this is one of the expectations because this is one of the biggest missing pieces. The model the biggest missing pieces. The model the biggest missing pieces. The model the other models use to search for files, to other models use to search for files, to other models use to search for files, to triple check changes, to do all the triple check changes, to do all the triple check changes, to do all the things that a model shouldn't waste things that a model shouldn't waste things that a model shouldn't waste tokens on if the model is expensive. I tokens on if the model is expensive. I tokens on if the model is expensive. I got to the point where I had taught my got to the point where I had taught my got to the point where I had taught my Claude to call Soul in order to save Claude to call Soul in order to save Claude to call Soul in order to save money because I didn't trust it calling money because I didn't trust it calling money because I didn't trust it calling Haiku 4.5. So, how did Haiku 5.5 turn Haiku 4.5. So, how did Haiku 5.5 turn Haiku 4.5. So, how did Haiku 5.5 turn out? I don't want to spoil too much, but out? I don't want to spoil too much, but out? I don't want to spoil too much, but I think it's fair to say I think it's fair to say I think it's fair to say they made something pretty cool. The they made something pretty cool. The they made something pretty cool. The price is right, the performance seems price is right, the performance seems price is right, the performance seems solid, has a handful of tricks that make solid, has a handful of tricks that make solid, has a handful of tricks that make it a little extra special, especially it a little extra special, especially it a little extra special, especially considering how small this model is. But considering how small this model is. But considering how small this model is. But before we can get started, in honor of before we can get started, in honor of before we can get started, in honor of the nature of the small model, we should the nature of the small model, we should the nature of the small model, we should take a really small break for today's take a really small break for today's take a really small break for today's sponsor. One of the painful side effects sponsor. One of the painful side effects sponsor. One of the painful side effects of having such good models for coding is of having such good models for coding is of having such good models for coding is that all of our code bases are getting that all of our code bases are getting that all of our code bases are getting way bigger and more complex. And while way bigger and more complex. And while way bigger and more complex. And while AI review tools are good at helping us AI review tools are good at helping us AI review tools are good at helping us avoid bugs, are they actually helping us avoid bugs, are they actually helping us avoid bugs, are they actually helping us prevent major security issues? And even prevent major security issues? And even prevent major security issues? And even if it does find one, is it actually if it does find one, is it actually if it does find one, is it actually going to let you fix it? Because let's going to let you fix it? Because let's going to let you fix it? Because let's be real, most of these models that are be real, most of these models that are be real, most of these models that are this good just refuse outright once you this good just refuse outright once you this good just refuse outright once you start doing security stuff. That's why start doing security stuff. That's why start doing security stuff. That's why I've been so reliant on today's sponsor.
-
I've been so reliant on today's sponsor. I've been so reliant on today's sponsor. You've already heard of them, it's Code You've already heard of them, it's Code You've already heard of them, it's Code Rabbit, but I'm showing something Rabbit, but I'm showing something Rabbit, but I'm showing something different today. But if you remember a different today. But if you remember a different today. But if you remember a few seconds ago, I said that the models few seconds ago, I said that the models few seconds ago, I said that the models can't really help with the security can't really help with the security can't really help with the security things because they're not allowed to. things because they're not allowed to. things because they're not allowed to. Well, Code Rabbit has made all of the Well, Code Rabbit has made all of the Well, Code Rabbit has made all of the deals they need to and found all the deals they need to and found all the deals they need to and found all the models they have to to get real security models they have to to get real security models they have to to get real security feedback for you as you're working. Not feedback for you as you're working. Not feedback for you as you're working. Not only will they give you real security only will they give you real security only will they give you real security callouts in your pull requests, they'll callouts in your pull requests, they'll callouts in your pull requests, they'll give you the tools needed to fix them. give you the tools needed to fix them. give you the tools needed to fix them. There's just like a little fix button There's just like a little fix button There's just like a little fix button that will run on their side with the that will run on their side with the that will run on their side with the models that are approved to do this. models that are approved to do this. models that are approved to do this. They also have something that's even They also have something that's even They also have something that's even more important, which is the ability to more important, which is the ability to more important, which is the ability to run deep scans across an entire repo. run deep scans across an entire repo. run deep scans across an entire repo. You click this little button in the You click this little button in the You click this little button in the corner, you choose the repo you want to corner, you choose the repo you want to corner, you choose the repo you want to run on, and then half an hour to an hour run on, and then half an hour to an hour run on, and then half an hour to an hour later, you have a ton of real feedback. later, you have a ton of real feedback. later, you have a ton of real feedback. I ended up addressing over half the I ended up addressing over half the I ended up addressing over half the things that were found in both of these things that were found in both of these things that were found in both of these reviews. And if I hop to this findings reviews. And if I hop to this findings reviews. And if I hop to this findings tab, I'm going to need my editor to tab, I'm going to need my editor to tab, I'm going to need my editor to start censoring things because the start censoring things because the start censoring things because the things it finds are actually quite things it finds are actually quite things it finds are actually quite useful to know about. And I'll be real, useful to know about. And I'll be real, useful to know about. And I'll be real, some of these are bad and I need to go some of these are bad and I need to go some of these are bad and I need to go address them. As our software's getting address them. As our software's getting address them. As our software's getting more complex, keeping it secure is only more complex, keeping it secure is only more complex, keeping it secure is only getting more important. Secure your getting more important. Secure your getting more important. Secure your stuff at swyd.link/coderabbit. Before we stuff at swyd.link/coderabbit. Before we stuff at swyd.link/coderabbit. Before we dive into everything about Haiku, I want dive into everything about Haiku, I want dive into everything about Haiku, I want to show one thing quick on Slapolitics.
-
to show one thing quick on Slapolitics. to show one thing quick on Slapolitics. We currently show all of Anthropic's We currently show all of Anthropic's We currently show all of Anthropic's current models. Sadly, we don't have current models. Sadly, we don't have current models. Sadly, we don't have available 5.5 yet, but we do have Opus available 5.5 yet, but we do have Opus available 5.5 yet, but we do have Opus 5.5, Sonnet 5.5, and now Haiku. Let's 5.5, Sonnet 5.5, and now Haiku. Let's 5.5, Sonnet 5.5, and now Haiku. Let's just show the last Haiku model real just show the last Haiku model real just show the last Haiku model real quick. quick. quick. Yeah. Hopefully not much more needs to Yeah. Hopefully not much more needs to Yeah. Hopefully not much more needs to be said. Let's go through the be said. Let's go through the be said. Let's go through the announcement. Introducing Haiku 5.5, the announcement. Introducing Haiku 5.5, the announcement. Introducing Haiku 5.5, the cheapest, fastest, and most capable cheapest, fastest, and most capable cheapest, fastest, and most capable small model we've ever released. Haiku small model we've ever released. Haiku small model we've ever released. Haiku 5.5 is designed for high-volume, 5.5 is designed for high-volume, 5.5 is designed for high-volume, cost-sensitive tasks. It reliably cost-sensitive tasks. It reliably cost-sensitive tasks. It reliably handles quick and repetitive workloads, handles quick and repetitive workloads, handles quick and repetitive workloads, things like summaries, compactions, things like summaries, compactions, things like summaries, compactions, database queries, and classification database queries, and classification database queries, and classification requests. It pairs well with Opus and requests. It pairs well with Opus and requests. It pairs well with Opus and Sonnet as a sub-agent on coding work. Sonnet as a sub-agent on coding work. Sonnet as a sub-agent on coding work. And since it's also our fastest model to And since it's also our fastest model to And since it's also our fastest model to date, it works especially well for date, it works especially well for date, it works especially well for speed-sensitive tasks like live customer speed-sensitive tasks like live customer speed-sensitive tasks like live customer support and browser use. Haiku 5.5 is support and browser use. Haiku 5.5 is support and browser use. Haiku 5.5 is available at a much lower price than available at a much lower price than available at a much lower price than 4.5. On average, it's now about 75% less 4.5. On average, it's now about 75% less 4.5. On average, it's now about 75% less money to run. Along with this launch, money to run. Along with this launch, money to run. Along with this launch, we're making improvements to the value we're making improvements to the value we're making improvements to the value of our model range. We are having the of our model range. We are having the of our model range. We are having the price of Sonnet 5.5's cache reads. Woo! price of Sonnet 5.5's cache reads. Woo! price of Sonnet 5.5's cache reads. Woo! I actually missed that detail. Huge. I actually missed that detail. Huge. I actually missed that detail. Huge. That was the thing I roasted them on in That was the thing I roasted them on in That was the thing I roasted them on in the Sonnet 5.5 video. That is a huge the Sonnet 5.5 video. That is a huge the Sonnet 5.5 video. That is a huge change, actually, that makes the whole change, actually, that makes the whole change, actually, that makes the whole lineup make way more sense. That might lineup make way more sense. That might lineup make way more sense. That might actually be my favorite thing about this actually be my favorite thing about this actually be my favorite thing about this release now. Oh, [ __ ] I actually didn't release now. Oh, [ __ ] I actually didn't release now. Oh, [ __ ] I actually didn't know about that. So, while I sit here in know about that. So, while I sit here in know about that. So, while I sit here in rage that nobody mentioned this to me, rage that nobody mentioned this to me, rage that nobody mentioned this to me, cuz I I'm a huge fan of changes to cache cuz I I'm a huge fan of changes to cache cuz I I'm a huge fan of changes to cache read prices. I have stories I wish I read prices. I have stories I wish I read prices. I have stories I wish I could tell that I can't. If you know, could tell that I can't. If you know, could tell that I can't. If you know, you know.
-
you know. you know. Let's just say I've been deep in the Let's just say I've been deep in the Let's just say I've been deep in the world of cache pricing as of recent. And world of cache pricing as of recent. And world of cache pricing as of recent. And the one complaint I had about Sonnet the one complaint I had about Sonnet the one complaint I had about Sonnet that also made Sonnet 5.5 much less that also made Sonnet 5.5 much less that also made Sonnet 5.5 much less appealing than 6.1 Soul from OpenAI is appealing than 6.1 Soul from OpenAI is appealing than 6.1 Soul from OpenAI is that the cache read price for Sonnet 5.5 that the cache read price for Sonnet 5.5 that the cache read price for Sonnet 5.5 wasn't just too expensive for Sonnet, it wasn't just too expensive for Sonnet, it wasn't just too expensive for Sonnet, it was the same price as Opus. So, the was the same price as Opus. So, the was the same price as Opus. So, the thing that made up about 30 to 50% of thing that made up about 30 to 50% of thing that made up about 30 to 50% of your bill was the same for Sonnet and your bill was the same for Sonnet and your bill was the same for Sonnet and for Opus, which meant the price for Opus, which meant the price for Opus, which meant the price difference between these models was not difference between these models was not difference between these models was not particularly large if you use them for particularly large if you use them for particularly large if you use them for agentic work. Cache reads are essential agentic work. Cache reads are essential agentic work. Cache reads are essential to making these things reasonably cheap. to making these things reasonably cheap. to making these things reasonably cheap. So, with that in mind, let's look at So, with that in mind, let's look at So, with that in mind, let's look at price quick, cuz this is what's exciting price quick, cuz this is what's exciting price quick, cuz this is what's exciting to me. Sonnet 5.5 went from 20 cents per to me. Sonnet 5.5 went from 20 cents per to me. Sonnet 5.5 went from 20 cents per million cache reads to 10 cents, which million cache reads to 10 cents, which million cache reads to 10 cents, which is the same cash read price as a model is the same cash read price as a model is the same cash read price as a model like GPT-61 Soul, finally making those a like GPT-61 Soul, finally making those a like GPT-61 Soul, finally making those a little bit closer in price. I hope little bit closer in price. I hope little bit closer in price. I hope Artificial Analysis re-prices their Artificial Analysis re-prices their Artificial Analysis re-prices their charts with this new price so that I charts with this new price so that I charts with this new price so that I could show the different numbers cuz could show the different numbers cuz could show the different numbers cuz that's a really big improvement, that's a really big improvement, that's a really big improvement, especially for agentic work. Haiku 4.5 especially for agentic work. Haiku 4.5 especially for agentic work. Haiku 4.5 was that same cash read price and then was that same cash read price and then was that same cash read price and then half the price across the board. Haiku half the price across the board. Haiku half the price across the board. Haiku 5.5 is a 10th the cash read price. It's 5.5 is a 10th the cash read price. It's 5.5 is a 10th the cash read price. It's 1 cent per mil tokens read. 12.5 cents 1 cent per mil tokens read. 12.5 cents 1 cent per mil tokens read. 12.5 cents per million cash token writes, 10 cents per million cash token writes, 10 cents per million cash token writes, 10 cents for normal 1 mil read, and 50 cents per for normal 1 mil read, and 50 cents per for normal 1 mil read, and 50 cents per mil out, where previously it was 5 bucks mil out, where previously it was 5 bucks mil out, where previously it was 5 bucks per mil out and Sonnet 5.5 was 10 bucks per mil out and Sonnet 5.5 was 10 bucks per mil out and Sonnet 5.5 was 10 bucks per mil out. This makes it 20x cheaper per mil out. This makes it 20x cheaper per mil out. This makes it 20x cheaper in output tokens and input tokens than in output tokens and input tokens than in output tokens and input tokens than Sonnet, comically cheaper in cash writes Sonnet, comically cheaper in cash writes Sonnet, comically cheaper in cash writes as well, about again, 20x cheaper, and
-
as well, about again, 20x cheaper, and as well, about again, 20x cheaper, and then cash reads are a 10th the price, then cash reads are a 10th the price, then cash reads are a 10th the price, too. There is a catch, though. That is too. There is a catch, though. That is too. There is a catch, though. That is up to a 100k token context. This is a up to a 100k token context. This is a up to a 100k token context. This is a bit of a weird thing because in Tropic bit of a weird thing because in Tropic bit of a weird thing because in Tropic used to charge when you broke 270 or used to charge when you broke 270 or used to charge when you broke 270 or 250k token context and doubled the 250k token context and doubled the 250k token context and doubled the price. They stopped doing that somewhat price. They stopped doing that somewhat price. They stopped doing that somewhat recently. That makes a lot of sense for recently. That makes a lot of sense for recently. That makes a lot of sense for tools like Claude Code because you don't tools like Claude Code because you don't tools like Claude Code because you don't have to like micromanage the context to have to like micromanage the context to have to like micromanage the context to keep your bill low, and cash reads are keep your bill low, and cash reads are keep your bill low, and cash reads are so cheap, it's kind of just fine. But a so cheap, it's kind of just fine. But a so cheap, it's kind of just fine. But a 100k token limit seems very small, 100k token limit seems very small, 100k token limit seems very small, especially for like a long-running especially for like a long-running especially for like a long-running agentic work. I see this as a couple agentic work. I see this as a couple agentic work. I see this as a couple things. And also, I'll just share in things. And also, I'll just share in things. And also, I'll just share in advance here, if you go over 100k advance here, if you go over 100k advance here, if you go over 100k tokens, the price doesn't double, it tokens, the price doesn't double, it tokens, the price doesn't double, it 5x's. So, you go from 50 cents per mil 5x's. So, you go from 50 cents per mil 5x's. So, you go from 50 cents per mil out to 250 per mil out. You go from 10 out to 250 per mil out. You go from 10 out to 250 per mil out. You go from 10 cents per mil in to 50 cents per mil in, cents per mil in to 50 cents per mil in, cents per mil in to 50 cents per mil in, etc. The reason for that, infra-wise, is etc. The reason for that, infra-wise, is etc. The reason for that, infra-wise, is because when you have a certain because when you have a certain because when you have a certain threshold of tokens in the context, that threshold of tokens in the context, that threshold of tokens in the context, that is now much more RAM you need to hold it is now much more RAM you need to hold it is now much more RAM you need to hold it all, which does make it more expensive. all, which does make it more expensive. all, which does make it more expensive. What I think happened here is they made What I think happened here is they made What I think happened here is they made the price for under 100k way cheaper, the price for under 100k way cheaper, the price for under 100k way cheaper, possibly like eating some of their own possibly like eating some of their own possibly like eating some of their own margin there, and they're making up a margin there, and they're making up a margin there, and they're making up a little bit of it with the over 100k. I little bit of it with the over 100k. I little bit of it with the over 100k. I see this a bit differently, though.
-
see this a bit differently, though. see this a bit differently, though. First off, I see this as a method of First off, I see this as a method of First off, I see this as a method of making the main use case, which is quick making the main use case, which is quick making the main use case, which is quick reads and categorization checks like reads and categorization checks like reads and categorization checks like that, as absurdly cheap as possible. that, as absurdly cheap as possible. that, as absurdly cheap as possible. Something like Jev is the real Something like Jev is the real Something like Jev is the real competitor to a small model used in competitor to a small model used in competitor to a small model used in these ways where it's being handed text these ways where it's being handed text these ways where it's being handed text and categorizing, organizing, all these and categorizing, organizing, all these and categorizing, organizing, all these types of things. By making under 100k types of things. By making under 100k types of things. By making under 100k this cheap, they are making it harder this cheap, they are making it harder this cheap, they are making it harder and harder to justify spinning up and harder to justify spinning up and harder to justify spinning up another API when you could just throw it another API when you could just throw it another API when you could just throw it at Haiku. And then the other side is to at Haiku. And then the other side is to at Haiku. And then the other side is to make sure they don't bankrupt themselves make sure they don't bankrupt themselves make sure they don't bankrupt themselves when people use this inside of Claude when people use this inside of Claude when people use this inside of Claude code. This is a way of steering people code. This is a way of steering people code. This is a way of steering people towards the intended use case by making towards the intended use case by making towards the intended use case by making it really, really cheap if they use it it really, really cheap if they use it it really, really cheap if they use it on short, quick response things. But on short, quick response things. But on short, quick response things. But also making it reasonably expensive, also making it reasonably expensive, also making it reasonably expensive, still half the price Haiku was before, still half the price Haiku was before, still half the price Haiku was before, but not for free like it is in the under but not for free like it is in the under but not for free like it is in the under 100k range if you're using it in tools 100k range if you're using it in tools 100k range if you're using it in tools like Claude code. It is worth noting like Claude code. It is worth noting like Claude code. It is worth noting that Anthropic employees like Lydia have that Anthropic employees like Lydia have that Anthropic employees like Lydia have come out publicly and said, "If you want come out publicly and said, "If you want come out publicly and said, "If you want to keep your bill cheap and not worry to keep your bill cheap and not worry to keep your bill cheap and not worry about this, you can manually set Haiku about this, you can manually set Haiku about this, you can manually set Haiku 5.5's auto compact window to 100k so 5.5's auto compact window to 100k so 5.5's auto compact window to 100k so that you're always working within that that you're always working within that that you're always working within that cheaper token pricing. Apparently it's a cheaper token pricing. Apparently it's a cheaper token pricing. Apparently it's a save per model, so this will only apply save per model, so this will only apply save per model, so this will only apply to Haiku including sub-agents. So if you to Haiku including sub-agents. So if you to Haiku including sub-agents. So if you go do this in your Claude code, if go do this in your Claude code, if go do this in your Claude code, if you're working on API billing or you're you're working on API billing or you're you're working on API billing or you're on like Bedrock or something, this on like Bedrock or something, this on like Bedrock or something, this allows you to guarantee the sub-agents allows you to guarantee the sub-agents allows you to guarantee the sub-agents will never cost a shitload of money. So will never cost a shitload of money. So will never cost a shitload of money. So now we need to talk about everything now we need to talk about everything now we need to talk about everything else, all the performance, all the use else, all the performance, all the use else, all the performance, all the use cases, and of course how it performs in cases, and of course how it performs in cases, and of course how it performs in my silly suite of benches. As we can see my silly suite of benches. As we can see my silly suite of benches. As we can see here, it absolutely decimated Haiku 4.5 here, it absolutely decimated Haiku 4.5 here, it absolutely decimated Haiku 4.5 getting more than double the score in getting more than double the score in getting more than double the score in the vast majority of things. And one the vast majority of things. And one the vast majority of things. And one that my friends at Anthropic mentioned
-
that my friends at Anthropic mentioned that my friends at Anthropic mentioned to me is that they were very happy to to me is that they were very happy to to me is that they were very happy to see that Haiku was no longer scoring a see that Haiku was no longer scoring a see that Haiku was no longer scoring a zero on Terminal Bench. They made their zero on Terminal Bench. They made their zero on Terminal Bench. They made their way up from zero to almost 40%. Also way up from zero to almost 40%. Also way up from zero to almost 40%. Also worth noting that as incredible as Luna worth noting that as incredible as Luna worth noting that as incredible as Luna is, and I still think GPT-6 Luna is an is, and I still think GPT-6 Luna is an is, and I still think GPT-6 Luna is an unbelievable model, GPT-6 Luna only got unbelievable model, GPT-6 Luna only got unbelievable model, GPT-6 Luna only got a 16.4%. a 16.4%. a 16.4%. So it more than doubled Luna's score for So it more than doubled Luna's score for So it more than doubled Luna's score for code. So if you're one of those people code. So if you're one of those people code. So if you're one of those people that thought Luna was decent at code, I that thought Luna was decent at code, I that thought Luna was decent at code, I question your judgment. I question a lot question your judgment. I question a lot question your judgment. I question a lot about you, honestly. But at the very about you, honestly. But at the very about you, honestly. But at the very least, you now have a model that is least, you now have a model that is least, you now have a model that is allegedly more than 2x better at coding allegedly more than 2x better at coding allegedly more than 2x better at coding for not too much more money. One of the for not too much more money. One of the for not too much more money. One of the most interesting numbers for me here is most interesting numbers for me here is most interesting numbers for me here is computer use. I've mentioned before that computer use. I've mentioned before that computer use. I've mentioned before that OpenAI has been pretty far ahead in OpenAI has been pretty far ahead in OpenAI has been pretty far ahead in computer use. And also, I don't think computer use. And also, I don't think computer use. And also, I don't think these benches are great because it these benches are great because it these benches are great because it doesn't show the gap as prominent as it doesn't show the gap as prominent as it doesn't show the gap as prominent as it is in my experience. Haiku slaughtered is in my experience. Haiku slaughtered is in my experience. Haiku slaughtered here. It did way better than 6 Luna did here. It did way better than 6 Luna did here. It did way better than 6 Luna did by like almost 40% or so. This makes by like almost 40% or so. This makes by like almost 40% or so. This makes this a very solid small model for doing this a very solid small model for doing this a very solid small model for doing computer use work, potentially. I computer use work, potentially. I computer use work, potentially. I haven't tried it yet. I've heard good haven't tried it yet. I've heard good haven't tried it yet. I've heard good things, though. And when you see the things, though. And when you see the things, though. And when you see the speeds, you'll see why I'm excited.
-
speeds, you'll see why I'm excited. speeds, you'll see why I'm excited. Because Haiku 5.5 is seeing numbers from Because Haiku 5.5 is seeing numbers from Because Haiku 5.5 is seeing numbers from 100 to 200 TPS, depending on the 100 to 200 TPS, depending on the 100 to 200 TPS, depending on the provider. Bias got notified by my chat provider. Bias got notified by my chat provider. Bias got notified by my chat that Prime tweeted he was using Haiku in that Prime tweeted he was using Haiku in that Prime tweeted he was using Haiku in some work he was doing with reasoning some work he was doing with reasoning some work he was doing with reasoning off for automations. He swapped to off for automations. He swapped to off for automations. He swapped to Haiku. It was a drop-in replacement. But Haiku. It was a drop-in replacement. But Haiku. It was a drop-in replacement. But he's actually seeing a worse pass rate he's actually seeing a worse pass rate he's actually seeing a worse pass rate and a much, much higher fail rate, as and a much, much higher fail rate, as and a much, much higher fail rate, as well as it being noticeably slower than well as it being noticeably slower than well as it being noticeably slower than what he was experiencing with Luna. This what he was experiencing with Luna. This what he was experiencing with Luna. This does not surprise me too much because does not surprise me too much because does not surprise me too much because the big thing the big thing the big thing that he did here is he turned off that he did here is he turned off that he did here is he turned off reasoning. He has both 6 Luna and Haiku reasoning. He has both 6 Luna and Haiku reasoning. He has both 6 Luna and Haiku 5.5 on no reasoning. 5.5 on no reasoning. 5.5 on no reasoning. OpenAI has trained their models to be OpenAI has trained their models to be OpenAI has trained their models to be good at that. Anthropic has trained good at that. Anthropic has trained good at that. Anthropic has trained their models to be bad at that. With their models to be bad at that. With their models to be bad at that. With Opus and Sonnet 5.5, if I recall Opus and Sonnet 5.5, if I recall Opus and Sonnet 5.5, if I recall correctly, I know they did with Opus. I correctly, I know they did with Opus. I correctly, I know they did with Opus. I think they did with Sonnet. They turned think they did with Sonnet. They turned think they did with Sonnet. They turned off no reasoning. So, you always have to off no reasoning. So, you always have to off no reasoning. So, you always have to let it reason at least a little bit. let it reason at least a little bit. let it reason at least a little bit. Haiku left no reasoning as an option Haiku left no reasoning as an option Haiku left no reasoning as an option because he will use it for work that because he will use it for work that because he will use it for work that they just want a response in quick. But they just want a response in quick. But they just want a response in quick. But it's not very good at that. it's not very good at that. it's not very good at that. I would honestly say that for like I would honestly say that for like I would honestly say that for like categorization work, if you're turning categorization work, if you're turning categorization work, if you're turning off reasoning anyways, you should just off reasoning anyways, you should just off reasoning anyways, you should just use Jev. I have been playing with the use Jev. I have been playing with the use Jev. I have been playing with the OpenAI decisions API, and I have plenty OpenAI decisions API, and I have plenty OpenAI decisions API, and I have plenty of thoughts to share about that in the of thoughts to share about that in the of thoughts to share about that in the near future. But for now, near future. But for now, near future. But for now, just use Jev unless you need like just use Jev unless you need like just use Jev unless you need like vision. And then Haiku or Luna makes vision. And then Haiku or Luna makes vision. And then Haiku or Luna makes sense. Probably or probably Luna cuz sense. Probably or probably Luna cuz sense. Probably or probably Luna cuz it's cheaper, especially with reasoning it's cheaper, especially with reasoning it's cheaper, especially with reasoning off. But uh off. But uh off. But uh yeah. Take that as you will. For what it yeah. Take that as you will. For what it yeah. Take that as you will. For what it is worth, when I tried out Haiku 55 on is worth, when I tried out Haiku 55 on is worth, when I tried out Haiku 55 on my machines with Claude code and my my machines with Claude code and my my machines with Claude code and my normal Claude code sub, I was seeing normal Claude code sub, I was seeing normal Claude code sub, I was seeing around 180 tokens per second for my just around 180 tokens per second for my just around 180 tokens per second for my just traditional prompting working on things traditional prompting working on things traditional prompting working on things like fish slop, which of course will be like fish slop, which of course will be like fish slop, which of course will be showing fish slop soon. So, while showing fish slop soon. So, while showing fish slop soon. So, while Prime's numbers with reasoning off
-
Prime's numbers with reasoning off Prime's numbers with reasoning off weren't great, Haiku's numbers with weren't great, Haiku's numbers with weren't great, Haiku's numbers with reasoning on do seem very good in this reasoning on do seem very good in this reasoning on do seem very good in this tier, which means it's time to hop into tier, which means it's time to hop into tier, which means it's time to hop into slop politics. Let's look at tokens per slop politics. Let's look at tokens per slop politics. Let's look at tokens per task cuz I am genuinely very curious and task cuz I am genuinely very curious and task cuz I am genuinely very curious and for good reason. It seems like this is for good reason. It seems like this is for good reason. It seems like this is not the usual wasteful [ __ ] show that we not the usual wasteful [ __ ] show that we not the usual wasteful [ __ ] show that we see from these smaller models from see from these smaller models from see from these smaller models from OpenAI. For reference, Sonnet 5.5 on max OpenAI. For reference, Sonnet 5.5 on max OpenAI. For reference, Sonnet 5.5 on max did almost 200,000 tokens per task did almost 200,000 tokens per task did almost 200,000 tokens per task average, putting it at I think the average, putting it at I think the average, putting it at I think the highest ever on Artificial Analysis, highest ever on Artificial Analysis, highest ever on Artificial Analysis, where Fable did 78k tokens, under half where Fable did 78k tokens, under half where Fable did 78k tokens, under half as many. Yeah, it was 2.5 times the as many. Yeah, it was 2.5 times the as many. Yeah, it was 2.5 times the tokens per task with Sonnet 5.5, which tokens per task with Sonnet 5.5, which tokens per task with Sonnet 5.5, which resulted in a cost for Sonnet 55 that resulted in a cost for Sonnet 55 that resulted in a cost for Sonnet 55 that was slightly higher than Fable 51. The was slightly higher than Fable 51. The was slightly higher than Fable 51. The reason for this is an unreasonable reason for this is an unreasonable reason for this is an unreasonable setting. It's right here. It's max. So, setting. It's right here. It's max. So, setting. It's right here. It's max. So, on slop politics, you can click the word on slop politics, you can click the word on slop politics, you can click the word max, which makes it go away, and now we max, which makes it go away, and now we max, which makes it go away, and now we have a chart that makes a lot more sense have a chart that makes a lot more sense have a chart that makes a lot more sense because those max reasoning levels are because those max reasoning levels are because those max reasoning levels are bad. And now we see tokens per task bad. And now we see tokens per task bad. And now we see tokens per task numbers that make way more sense and numbers that make way more sense and numbers that make way more sense and also cost per task numbers that make way also cost per task numbers that make way also cost per task numbers that make way more sense. Where Haiku 55 is now all more sense. Where Haiku 55 is now all more sense. Where Haiku 55 is now all the way on the left here, cheaper than the way on the left here, cheaper than the way on the left here, cheaper than anything else, because I only have X anything else, because I only have X anything else, because I only have X high and max numbers for it. The max high and max numbers for it. The max high and max numbers for it. The max number is still pretty cheap, but more number is still pretty cheap, but more number is still pretty cheap, but more than I would want it to be, especially than I would want it to be, especially than I would want it to be, especially once you add back in models like 61 once you add back in models like 61 once you add back in models like 61 Soul. Because 61 Soul on low is as cheap Soul. Because 61 Soul on low is as cheap Soul. Because 61 Soul on low is as cheap as Haiku on low, largely due to the as Haiku on low, largely due to the as Haiku on low, largely due to the token efficiency difference. But, let's token efficiency difference. But, let's token efficiency difference. But, let's go back to the cost versus intelligence go back to the cost versus intelligence go back to the cost versus intelligence cuz I think this is fun. You also see 45 cuz I think this is fun. You also see 45 cuz I think this is fun. You also see 45 Haiku in here, which I am going to turn Haiku in here, which I am going to turn Haiku in here, which I am going to turn off because we don't care about that off because we don't care about that off because we don't care about that model anymore. It can finally be left model anymore. It can finally be left model anymore. It can finally be left where it belongs, the grave. And we'll
-
where it belongs, the grave. And we'll where it belongs, the grave. And we'll notice some weird things here. All this notice some weird things here. All this notice some weird things here. All this data is from Artificial Analysis and it data is from Artificial Analysis and it data is from Artificial Analysis and it is far from representative of everything is far from representative of everything is far from representative of everything these models could or would or should be these models could or would or should be these models could or would or should be used for, but you'll see that aggregate used for, but you'll see that aggregate used for, but you'll see that aggregate across all their benches, Sonnet 5.5 across all their benches, Sonnet 5.5 across all their benches, Sonnet 5.5 performs worse than Opus at the same if performs worse than Opus at the same if performs worse than Opus at the same if not a slightly higher cost, which meant not a slightly higher cost, which meant not a slightly higher cost, which meant that Opus and Sonnet aren't as that Opus and Sonnet aren't as that Opus and Sonnet aren't as complimentary as one might hope. But, if complimentary as one might hope. But, if complimentary as one might hope. But, if we look all the way to left here with we look all the way to left here with we look all the way to left here with Haiku 5.5, things are looking a lot more Haiku 5.5, things are looking a lot more Haiku 5.5, things are looking a lot more promising. Because even though these promising. Because even though these promising. Because even though these scores are close, scores are close, scores are close, Opus was almost five times as expensive Opus was almost five times as expensive Opus was almost five times as expensive as the run was with Haiku. as the run was with Haiku. as the run was with Haiku. Pretty solid. I have a fun feature here Pretty solid. I have a fun feature here Pretty solid. I have a fun feature here as well, the Pareto line, which is the as well, the Pareto line, which is the as well, the Pareto line, which is the best performance at a given price point. best performance at a given price point. best performance at a given price point. And this line moves around based on like And this line moves around based on like And this line moves around based on like what is the best you can do at that what is the best you can do at that what is the best you can do at that level of intelligence. And you can see level of intelligence. And you can see level of intelligence. And you can see some very interesting things here. Of some very interesting things here. Of some very interesting things here. Of course, as you expect, Haiku 5.5 does course, as you expect, Haiku 5.5 does course, as you expect, Haiku 5.5 does very well here with X-High and Max being very well here with X-High and Max being very well here with X-High and Max being the best value in that range within the best value in that range within the best value in that range within Anthropic's line of models. Sonnet 5.5 Anthropic's line of models. Sonnet 5.5 Anthropic's line of models. Sonnet 5.5 sneaks in randomly with high because it sneaks in randomly with high because it sneaks in randomly with high because it is smarter than Opus low and cheaper is smarter than Opus low and cheaper is smarter than Opus low and cheaper than Opus medium. But, after Sonnet than Opus medium. But, after Sonnet than Opus medium. But, after Sonnet high, all the best value, according to high, all the best value, according to high, all the best value, according to this bench, is just Opus 5.5 medium, this bench, is just Opus 5.5 medium, this bench, is just Opus 5.5 medium, high, X-High, Max.
-
high, X-High, Max. high, X-High, Max. But, here is where I must admit I was But, here is where I must admit I was But, here is where I must admit I was slightly misleading you because I only slightly misleading you because I only slightly misleading you because I only am showing Anthropic models right now. am showing Anthropic models right now. am showing Anthropic models right now. In fact, I'm only showing Anthropic In fact, I'm only showing Anthropic In fact, I'm only showing Anthropic models that are the latest for each of models that are the latest for each of models that are the latest for each of these lines. And also, Fable 5.1 doesn't these lines. And also, Fable 5.1 doesn't these lines. And also, Fable 5.1 doesn't even light up anymore cuz there's no even light up anymore cuz there's no even light up anymore cuz there's no reason to use that when you look at the reason to use that when you look at the reason to use that when you look at the Pareto line here. Ready for where things Pareto line here. Ready for where things Pareto line here. Ready for where things start to get more interesting though? start to get more interesting though? start to get more interesting though? First off, we can switch to linear First off, we can switch to linear First off, we can switch to linear pricing where you see the gap isn't pricing where you see the gap isn't pricing where you see the gap isn't quite as big as it seems because the log quite as big as it seems because the log quite as big as it seems because the log pricing makes things seem further away pricing makes things seem further away pricing makes things seem further away on the cheap side and closer on the on the cheap side and closer on the on the cheap side and closer on the expensive side than they really are. And expensive side than they really are. And expensive side than they really are. And you can see here that Sonnet X-High is a you can see here that Sonnet X-High is a you can see here that Sonnet X-High is a fine value at $2.75 per task. But, then fine value at $2.75 per task. But, then fine value at $2.75 per task. But, then Sonnet 5.5 Max makes literally no sense Sonnet 5.5 Max makes literally no sense Sonnet 5.5 Max makes literally no sense at $7.60 at $7.60 at $7.60 per task. Do not if you're using Max per task. Do not if you're using Max per task. Do not if you're using Max reasoning levels, I have questions. But, reasoning levels, I have questions. But, reasoning levels, I have questions. But, I also will say very confidently, do not I also will say very confidently, do not I also will say very confidently, do not use Sonnet 5.5 on Max, it makes no sense use Sonnet 5.5 on Max, it makes no sense use Sonnet 5.5 on Max, it makes no sense at all. But, again, still just looking at all. But, again, still just looking at all. But, again, still just looking at Anthropic models. What happens if we at Anthropic models. What happens if we at Anthropic models. What happens if we put on like Gemini 4 Argon? Oh, not put on like Gemini 4 Argon? Oh, not put on like Gemini 4 Argon? Oh, not much. How about 38 flash? Still not much. How about 38 flash? Still not much. How about 38 flash? Still not much. In fact, it almost looks like we much. In fact, it almost looks like we much. In fact, it almost looks like we took the Haiku line and moved it to the took the Haiku line and moved it to the took the Haiku line and moved it to the right and down. So, we don't really need right and down. So, we don't really need right and down. So, we don't really need those anymore. We'll just turn those those anymore. We'll just turn those those anymore. We'll just turn those back off. How about back off. How about back off. How about uh I don't know, Grok 47?
-
uh I don't know, Grok 47? uh I don't know, Grok 47? Still not very interesting. Still below Still not very interesting. Still below Still not very interesting. Still below all of this. all of this. all of this. How about 6 Astra? How about 6 Astra? How about 6 Astra? Oh, they got a point on. Astra on the Oh, they got a point on. Astra on the Oh, they got a point on. Astra on the low is around the same intelligence as low is around the same intelligence as low is around the same intelligence as Sonnet 55 on high. So, that adds to the Sonnet 55 on high. So, that adds to the Sonnet 55 on high. So, that adds to the Pareto line. But, here's where things Pareto line. But, here's where things Pareto line. But, here's where things get more interesting. Let's put on GPT-6 get more interesting. Let's put on GPT-6 get more interesting. Let's put on GPT-6 Luna. Luna. Luna. Oh. Huh. That just extended the line a Oh. Huh. That just extended the line a Oh. Huh. That just extended the line a whole bunch to the left because Luna is whole bunch to the left because Luna is whole bunch to the left because Luna is cheaper than Haiku, always. cheaper than Haiku, always. cheaper than Haiku, always. It's also dumber, always, at least with It's also dumber, always, at least with It's also dumber, always, at least with the numbers they have. But, these are the numbers they have. But, these are the numbers they have. But, these are complementary. This is a pretty stable complementary. This is a pretty stable complementary. This is a pretty stable line. If you turn on the Pareto line, it line. If you turn on the Pareto line, it line. If you turn on the Pareto line, it has all of Luna, all of Haiku, then it has all of Luna, all of Haiku, then it has all of Luna, all of Haiku, then it briefly dips to Astra, then Sonnet, then briefly dips to Astra, then Sonnet, then briefly dips to Astra, then Sonnet, then Opus for the top. And if you were to Opus for the top. And if you were to Opus for the top. And if you were to turn off Astra turn off Astra turn off Astra and Sonnet, you have a pretty clear, and Sonnet, you have a pretty clear, and Sonnet, you have a pretty clear, simple line here for the best value per simple line here for the best value per simple line here for the best value per dollar. dollar. dollar. But, here is where things get messy. But, here is where things get messy. But, here is where things get messy. I'm going to turn off the Pareto line to I'm going to turn off the Pareto line to I'm going to turn off the Pareto line to make this clearer. I'm also going to make this clearer. I'm also going to make this clearer. I'm also going to turn off Fable and Sonnet cuz they don't turn off Fable and Sonnet cuz they don't turn off Fable and Sonnet cuz they don't really make sense here. really make sense here. really make sense here. Now, what happens when we turn on 61 Now, what happens when we turn on 61 Now, what happens when we turn on 61 Sol? Oh.
-
That's unfortunate. That's unfortunate. Turn on the Pareto line. Turn on the Pareto line. Turn on the Pareto line. And I guess technically Haiku is And I guess technically Haiku is And I guess technically Haiku is slightly smarter for meaningfully more slightly smarter for meaningfully more slightly smarter for meaningfully more money there, but then money there, but then money there, but then just barely more. Going from just barely more. Going from just barely more. Going from 21.3 cents to 21.4 cents, you get a huge 21.3 cents to 21.4 cents, you get a huge 21.3 cents to 21.4 cents, you get a huge jump in intelligence. And if we were to jump in intelligence. And if we were to jump in intelligence. And if we were to just turn Haiku off here, oh, that's a just turn Haiku off here, oh, that's a just turn Haiku off here, oh, that's a much smoother line. much smoother line. much smoother line. So, if you're actually trying to So, if you're actually trying to So, if you're actually trying to optimize based on the benchmark scores optimize based on the benchmark scores optimize based on the benchmark scores for the best intelligence per dollar, for the best intelligence per dollar, for the best intelligence per dollar, your stock would look like 6 Luna, 61 your stock would look like 6 Luna, 61 your stock would look like 6 Luna, 61 Sol, and Opus 55. There is one more Sol, and Opus 55. There is one more Sol, and Opus 55. There is one more hidden gem, diamond in the rough, that hidden gem, diamond in the rough, that hidden gem, diamond in the rough, that I'm surprised people aren't talking I'm surprised people aren't talking I'm surprised people aren't talking about as much. about as much. about as much. Mimo V26 Pro is an underrated open Mimo V26 Pro is an underrated open Mimo V26 Pro is an underrated open weight model. Pretty damn good for the weight model. Pretty damn good for the weight model. Pretty damn good for the price. But other than that, price. But other than that, price. But other than that, nothing really comes above this line. nothing really comes above this line. nothing really comes above this line. So, if you're trying to get the absolute So, if you're trying to get the absolute So, if you're trying to get the absolute most value and you're paying API prices, most value and you're paying API prices, most value and you're paying API prices, 6 Luna 61 Soul Opus gets you pretty much 6 Luna 61 Soul Opus gets you pretty much 6 Luna 61 Soul Opus gets you pretty much all of that. That said, Haiku 55 is neck all of that. That said, Haiku 55 is neck all of that. That said, Haiku 55 is neck and neck now. Are Soul and Luna still and neck now. Are Soul and Luna still and neck now. Are Soul and Luna still better per dollar? Slightly, usually.
-
better per dollar? Slightly, usually. better per dollar? Slightly, usually. However, if you are trying to do However, if you are trying to do However, if you are trying to do everything through one provider, like everything through one provider, like everything through one provider, like you want to just use your Anthropic you want to just use your Anthropic you want to just use your Anthropic inference cuz you have a crazy spend inference cuz you have a crazy spend inference cuz you have a crazy spend target with them and you just use their target with them and you just use their target with them and you just use their APIs, or you're subscribed to Claude APIs, or you're subscribed to Claude APIs, or you're subscribed to Claude Code and don't want to also subscribe to Code and don't want to also subscribe to Code and don't want to also subscribe to Codex, or you want models that Codex, or you want models that Codex, or you want models that understand each other better cuz you're understand each other better cuz you're understand each other better cuz you're using Opus for everything and when Opus using Opus for everything and when Opus using Opus for everything and when Opus prompts Soul it's worse than when Opus prompts Soul it's worse than when Opus prompts Soul it's worse than when Opus prompts Haiku. I don't know how true prompts Haiku. I don't know how true prompts Haiku. I don't know how true that is, but it's real. There's a lot of that is, but it's real. There's a lot of that is, but it's real. There's a lot of reasons that you would want a model in reasons that you would want a model in reasons that you would want a model in that range in this family. And for once, that range in this family. And for once, that range in this family. And for once, it's no longer irresponsible to use it's no longer irresponsible to use it's no longer irresponsible to use Haiku like it was before. Haiku 4.5 was Haiku like it was before. Haiku 4.5 was Haiku like it was before. Haiku 4.5 was so bad and so unnecessarily expensive so bad and so unnecessarily expensive so bad and so unnecessarily expensive that there was no justifiable reason to that there was no justifiable reason to that there was no justifiable reason to use it. When I had my tweet that use it. When I had my tweet that use it. When I had my tweet that Anthropic has no small models worth Anthropic has no small models worth Anthropic has no small models worth using, OpenAI has no large models worth using, OpenAI has no large models worth using, OpenAI has no large models worth using, and Google has none that worth using, and Google has none that worth using, and Google has none that worth using, somebody replied, "All three using, somebody replied, "All three using, somebody replied, "All three statements are false." To which I statements are false." To which I statements are false." To which I replied, "LMAO, this guy uses Haiku replied, "LMAO, this guy uses Haiku replied, "LMAO, this guy uses Haiku 4.5." And then he got torn to teeny, 4.5." And then he got torn to teeny, 4.5." And then he got torn to teeny, tiny pieces in my replies because there tiny pieces in my replies because there tiny pieces in my replies because there is literally no reason to use Haiku 45 is literally no reason to use Haiku 45 is literally no reason to use Haiku 45 at this point in September of this year, at this point in September of this year, at this point in September of this year, but also right now especially. Haiku 45 but also right now especially. Haiku 45 but also right now especially. Haiku 45 was barely usable at the time and was barely usable at the time and was barely usable at the time and quickly became garbage and just wasn't quickly became garbage and just wasn't quickly became garbage and just wasn't touched because Anthropic was too touched because Anthropic was too touched because Anthropic was too focused on the big models. Now they're focused on the big models. Now they're focused on the big models. Now they're not. And while I do wish this model not. And while I do wish this model not. And while I do wish this model pushed the Pareto line a bit more, I am pushed the Pareto line a bit more, I am pushed the Pareto line a bit more, I am thankful Anthropic is embracing models thankful Anthropic is embracing models thankful Anthropic is embracing models in this range because together, these in this range because together, these in this range because together, these all actually look quite reasonable. Also all actually look quite reasonable. Also all actually look quite reasonable. Also remember, Sonnet 55 got way cheaper with remember, Sonnet 55 got way cheaper with remember, Sonnet 55 got way cheaper with the cash read price change, so that the cash read price change, so that the cash read price change, so that might actually adjust things even more.
-
might actually adjust things even more. might actually adjust things even more. Good news, Artificial Analysis updated Good news, Artificial Analysis updated Good news, Artificial Analysis updated their data, their data, their data, which means I've updated mine. And now which means I've updated mine. And now which means I've updated mine. And now Sonnet is shifted quite a bit to the Sonnet is shifted quite a bit to the Sonnet is shifted quite a bit to the left, which makes it much more left, which makes it much more left, which makes it much more compelling on the Pareto line. Actually, compelling on the Pareto line. Actually, compelling on the Pareto line. Actually, not that much more. It does make this not that much more. It does make this not that much more. It does make this bottom Sonnet 5.5 high feel less awkward bottom Sonnet 5.5 high feel less awkward bottom Sonnet 5.5 high feel less awkward as a dip. But yeah, that is improved. as a dip. But yeah, that is improved. as a dip. But yeah, that is improved. Sonnet feels less bad now by quite a bit Sonnet feels less bad now by quite a bit Sonnet feels less bad now by quite a bit with that change. Okay, enough of the with that change. Okay, enough of the with that change. Okay, enough of the bench maxing. Let's talk about the model bench maxing. Let's talk about the model bench maxing. Let's talk about the model and how it actually works. Okay, one and how it actually works. Okay, one and how it actually works. Okay, one last fun thing from Anthropic's release last fun thing from Anthropic's release last fun thing from Anthropic's release notes here that is actually quite cool. notes here that is actually quite cool. notes here that is actually quite cool. They want to encourage people using They want to encourage people using They want to encourage people using Claude code to build more things with Claude code to build more things with Claude code to build more things with the Claude APIs. A lot of people, myself the Claude APIs. A lot of people, myself the Claude APIs. A lot of people, myself included, were scared of this included, were scared of this included, were scared of this announcement thinking this might be the announcement thinking this might be the announcement thinking this might be the end of using your sub and tools like T3 end of using your sub and tools like T3 end of using your sub and tools like T3 Code cuz yes, in T3 Code you can use Code cuz yes, in T3 Code you can use Code cuz yes, in T3 Code you can use your sub with Anthropic, OpenAI, or your sub with Anthropic, OpenAI, or your sub with Anthropic, OpenAI, or whatever else. If it's Claude Code or whatever else. If it's Claude Code or whatever else. If it's Claude Code or Codex installed on your computer, it Codex installed on your computer, it Codex installed on your computer, it will just work with T3 Code. And a lot will just work with T3 Code. And a lot will just work with T3 Code. And a lot of us are scared cuz they had said they of us are scared cuz they had said they of us are scared cuz they had said they would kill that, and they haven't yet. would kill that, and they haven't yet. would kill that, and they haven't yet. But they did a cool thing here, actually But they did a cool thing here, actually But they did a cool thing here, actually really cool thing, where if you're really cool thing, where if you're really cool thing, where if you're subscribed as a 20x user, so you're on subscribed as a 20x user, so you're on subscribed as a 20x user, so you're on the $200 plan, you get $200 in API the $200 plan, you get $200 in API the $200 plan, you get $200 in API credit every month. 5x users get $100 in credit every month. 5x users get $100 in credit every month. 5x users get $100 in credit. And this is for the API. So credit. And this is for the API. So credit. And this is for the API. So you're effectively able to get an API you're effectively able to get an API you're effectively able to get an API key like you would if you were paying, key like you would if you were paying, key like you would if you were paying, but they pre-load it with 100 or $200 but they pre-load it with 100 or $200 but they pre-load it with 100 or $200 every month to throw at whatever. So if every month to throw at whatever. So if every month to throw at whatever. So if you want to vibe code an app in Claude you want to vibe code an app in Claude you want to vibe code an app in Claude code that has Haiku or Sonnet built into code that has Haiku or Sonnet built into code that has Haiku or Sonnet built into it and then send it to your friends, you it and then send it to your friends, you it and then send it to your friends, you now have $200 of credit for that, which now have $200 of credit for that, which now have $200 of credit for that, which I actually think is really nice of them.
-
I actually think is really nice of them. I actually think is really nice of them. I I I Do I say these things? Yeah, [ __ ] it. I Do I say these things? Yeah, [ __ ] it. I Do I say these things? Yeah, [ __ ] it. I chatted with them a lot about this back chatted with them a lot about this back chatted with them a lot about this back in the day cuz they wanted to make sure in the day cuz they wanted to make sure in the day cuz they wanted to make sure they didn't screw up things with how they didn't screw up things with how they didn't screw up things with how like Claude Code SDK and Agent SDK like Claude Code SDK and Agent SDK like Claude Code SDK and Agent SDK integrations were going to end up. They integrations were going to end up. They integrations were going to end up. They were going to do something a lot were going to do something a lot were going to do something a lot stupider with this. stupider with this. stupider with this. And I'm thankful they talked to people And I'm thankful they talked to people And I'm thankful they talked to people like me and that they listened because like me and that they listened because like me and that they listened because they turned what could have been an L they turned what could have been an L they turned what could have been an L into a really genuinely cool thing. That into a really genuinely cool thing. That into a really genuinely cool thing. That this is If you guys have noticed that this is If you guys have noticed that this is If you guys have noticed that I'm being nicer to Anthropic, it's cuz I'm being nicer to Anthropic, it's cuz I'm being nicer to Anthropic, it's cuz they're actually listening now and this they're actually listening now and this they're actually listening now and this is a great example of one of those is a great example of one of those is a great example of one of those things. Ooh, actually one more fun small things. Ooh, actually one more fun small things. Ooh, actually one more fun small detail here. They're updating the Claude detail here. They're updating the Claude detail here. They're updating the Claude Python and TypeScript SDKs to add Python and TypeScript SDKs to add Python and TypeScript SDKs to add support for computer use and browser support for computer use and browser support for computer use and browser use. That is potentially very useful for use. That is potentially very useful for use. That is potentially very useful for us. T3 code might level up from this. I us. T3 code might level up from this. I us. T3 code might level up from this. I just saw Elon reply to them with just saw Elon reply to them with just saw Elon reply to them with congrats on another great model. I would congrats on another great model. I would congrats on another great model. I would be surprised if Haiku 5.5 wasn't nicer be surprised if Haiku 5.5 wasn't nicer be surprised if Haiku 5.5 wasn't nicer to use than Grok 4.7 if I'm being real. to use than Grok 4.7 if I'm being real. to use than Grok 4.7 if I'm being real. Oh, I missed the artificial analysis Oh, I missed the artificial analysis Oh, I missed the artificial analysis jump chart. jump chart. jump chart. This is hilarious. This is hilarious. This is hilarious. Beautiful. From the furthest right end Beautiful. From the furthest right end Beautiful. From the furthest right end here to actually competitive. They are a here to actually competitive. They are a here to actually competitive. They are a bit more brutal showing it is not on the bit more brutal showing it is not on the bit more brutal showing it is not on the Pareto line, but uh yeah, the data is Pareto line, but uh yeah, the data is Pareto line, but uh yeah, the data is updated. It's more interesting than they updated. It's more interesting than they updated. It's more interesting than they give it credit. Regardless, I'm tired of give it credit. Regardless, I'm tired of give it credit. Regardless, I'm tired of talking about benchmarks. I want to talking about benchmarks. I want to talking about benchmarks. I want to showcase the awesome things this model showcase the awesome things this model showcase the awesome things this model can do. And not just by itself, but when can do. And not just by itself, but when can do. And not just by itself, but when it's used alongside models like Opus 5.5 it's used alongside models like Opus 5.5 it's used alongside models like Opus 5.5 and Sonnet 5.5. The best use cases for and Sonnet 5.5. The best use cases for and Sonnet 5.5. The best use cases for Haiku 5.5 aren't clicking it in the UI Haiku 5.5 aren't clicking it in the UI Haiku 5.5 aren't clicking it in the UI or selecting it in Claude code and using or selecting it in Claude code and using or selecting it in Claude code and using it. It's letting your agents use it and it. It's letting your agents use it and it. It's letting your agents use it and letting your applications use it.
-
letting your applications use it. letting your applications use it. Building it in so when data comes in, Building it in so when data comes in, Building it in so when data comes in, Haiku will audit it and categorize it. Haiku will audit it and categorize it. Haiku will audit it and categorize it. Or having Sonnet or Opus call Haiku for Or having Sonnet or Opus call Haiku for Or having Sonnet or Opus call Haiku for testing different things out, exploring testing different things out, exploring testing different things out, exploring the code base to find stuff. Anthropic the code base to find stuff. Anthropic the code base to find stuff. Anthropic made a really cool demo here where they made a really cool demo here where they made a really cool demo here where they had Opus 5.5 by itself try to build a had Opus 5.5 by itself try to build a had Opus 5.5 by itself try to build a safe egg drop in this emulation of like safe egg drop in this emulation of like safe egg drop in this emulation of like egg drop mechanics that they built. And egg drop mechanics that they built. And egg drop mechanics that they built. And on the other side, they had Opus 5.5 on the other side, they had Opus 5.5 on the other side, they had Opus 5.5 plus 10 Haiku 5.5 sub agents do it. And plus 10 Haiku 5.5 sub agents do it. And plus 10 Haiku 5.5 sub agents do it. And I think the results here are actually I think the results here are actually I think the results here are actually quite interesting. quite interesting. quite interesting. The thing that you might notice pretty The thing that you might notice pretty The thing that you might notice pretty quickly here is that Opus is only trying quickly here is that Opus is only trying quickly here is that Opus is only trying one at a time cuz it's running single one at a time cuz it's running single one at a time cuz it's running single threaded. And if you were to have it run threaded. And if you were to have it run threaded. And if you were to have it run multiple threads, it would end up multiple threads, it would end up multiple threads, it would end up wasting way more money. Where on the wasting way more money. Where on the wasting way more money. Where on the right, we have Opus plus 10 Haiku sub right, we have Opus plus 10 Haiku sub right, we have Opus plus 10 Haiku sub agents and it's able to try more things agents and it's able to try more things agents and it's able to try more things at the same time and save a ton of money at the same time and save a ton of money at the same time and save a ton of money and get to a result faster as well. and get to a result faster as well. and get to a result faster as well. There you go. There you go. There you go. It took 3 and 1/2 minutes for Opus 5.5 It took 3 and 1/2 minutes for Opus 5.5 It took 3 and 1/2 minutes for Opus 5.5 by itself to succeed at this task. It by itself to succeed at this task. It by itself to succeed at this task. It only designed 25 attempts and it cost only designed 25 attempts and it cost only designed 25 attempts and it cost $0.47 to do. But with Opus 5.5 $0.47 to do. But with Opus 5.5 $0.47 to do. But with Opus 5.5 commanding 10 Haiku instances, it was commanding 10 Haiku instances, it was commanding 10 Haiku instances, it was able to do it in under a minute with 86 able to do it in under a minute with 86 able to do it in under a minute with 86 attempts. Way more attempts because it's attempts. Way more attempts because it's attempts. Way more attempts because it's able to get more renditions out when able to get more renditions out when able to get more renditions out when it's that much cheaper and faster and it's that much cheaper and faster and it's that much cheaper and faster and only cost $0.14. So, for the types of only cost $0.14. So, for the types of only cost $0.14. So, for the types of tasks where the majority of answers are tasks where the majority of answers are tasks where the majority of answers are wrong, like if you're looking through a wrong, like if you're looking through a wrong, like if you're looking through a code base to find one file out of a code base to find one file out of a code base to find one file out of a bunch and you can't just do like a rip bunch and you can't just do like a rip bunch and you can't just do like a rip grep call to find it. Haiku will be way grep call to find it. Haiku will be way grep call to find it. Haiku will be way cheaper and faster cuz it can do a lot cheaper and faster cuz it can do a lot cheaper and faster cuz it can do a lot of things at once without causing crazy of things at once without causing crazy of things at once without causing crazy bills. So, for the subset of tasks where
-
bills. So, for the subset of tasks where bills. So, for the subset of tasks where more attempts is more valuable than more attempts is more valuable than more attempts is more valuable than smart attempts, where you'd rather try smart attempts, where you'd rather try smart attempts, where you'd rather try 10 times and verify each result than 10 times and verify each result than 10 times and verify each result than spend a little more time trying once and spend a little more time trying once and spend a little more time trying once and hoping it's right. If the gap between a hoping it's right. If the gap between a hoping it's right. If the gap between a random try from a dumb model and a random try from a dumb model and a random try from a dumb model and a thoughtful try from a smart model isn't thoughtful try from a smart model isn't thoughtful try from a smart model isn't that big, Haiku 5.5 is going to that big, Haiku 5.5 is going to that big, Haiku 5.5 is going to slaughter. This is a small subset of slaughter. This is a small subset of slaughter. This is a small subset of work and I am hoping Opus is smart work and I am hoping Opus is smart work and I am hoping Opus is smart enough to know what that subset is. I enough to know what that subset is. I enough to know what that subset is. I have not pushed it hard enough to know have not pushed it hard enough to know have not pushed it hard enough to know yet, but I am going to go edit my Claude yet, but I am going to go edit my Claude yet, but I am going to go edit my Claude MD to no longer force everything to be MD to no longer force everything to be MD to no longer force everything to be Opus or sole sub agents because this is Opus or sole sub agents because this is Opus or sole sub agents because this is cheap enough that it would make real cheap enough that it would make real cheap enough that it would make real sense. But you guys aren't here just to sense. But you guys aren't here just to sense. But you guys aren't here just to listen to me rant about sub agent listen to me rant about sub agent listen to me rant about sub agent architecture. You're here to see what architecture. You're here to see what architecture. You're here to see what the model's capable of. So, let's see the model's capable of. So, let's see the model's capable of. So, let's see some of the fun demos people have been some of the fun demos people have been some of the fun demos people have been making with Haiku 5.5. Multiple people making with Haiku 5.5. Multiple people making with Haiku 5.5. Multiple people from my community have been throwing from my community have been throwing from my community have been throwing together video demos cuz everyone knows together video demos cuz everyone knows together video demos cuz everyone knows Opus 5.5 the model got way better at Opus 5.5 the model got way better at Opus 5.5 the model got way better at programmatically creating video. And programmatically creating video. And programmatically creating video. And this example came from Arshan in the this example came from Arshan in the this example came from Arshan in the community.
-
>> [music] >> [music] [bell] >> This is >> This is insanely impressive for the money. insanely impressive for the money. insanely impressive for the money. God damn. Yeah, I am very impressed. And this cost Yeah, I am very impressed. And this cost 60 cents an API cost to make. 60 cents an API cost to make. 60 cents an API cost to make. Was built in under 20 minutes. That's Was built in under 20 minutes. That's Was built in under 20 minutes. That's insane. A lot of why Haiku could make a insane. A lot of why Haiku could make a insane. A lot of why Haiku could make a video that good is that the Anthropic video that good is that the Anthropic video that good is that the Anthropic models have pretty good like taste built models have pretty good like taste built models have pretty good like taste built in is the best I can put it especially in is the best I can put it especially in is the best I can put it especially around design stuff. And I think they around design stuff. And I think they around design stuff. And I think they just have a lot of like really good just have a lot of like really good just have a lot of like really good design data and have obsessively trimmed design data and have obsessively trimmed design data and have obsessively trimmed that data set down to give the model that data set down to give the model that data set down to give the model good reference points. And you can see good reference points. And you can see good reference points. And you can see that even more so with the front-end that even more so with the front-end that even more so with the front-end design capabilities using which AI by design capabilities using which AI by design capabilities using which AI by our friend Dara. our friend Dara. our friend Dara. This is so comically better than Grok. This is so comically better than Grok. This is so comically better than Grok. Oh, that's cool actually. The little Oh, that's cool actually. The little Oh, that's cool actually. The little animation for how that text came and I animation for how that text came and I animation for how that text came and I haven't seen anything like that in one haven't seen anything like that in one haven't seen anything like that in one of these demos before actually. It's different, but it's it's cool. It's It's different, but it's it's cool. It's like a new thing. This is the blueprinty like a new thing. This is the blueprinty like a new thing. This is the blueprinty one. I don't love this diagram. It's one. I don't love this diagram. It's one. I don't love this diagram. It's cool it like has hover behaviors and cool it like has hover behaviors and cool it like has hover behaviors and changes things underneath.
-
changes things underneath. changes things underneath. But uh But uh But uh not my favorite. This is kind of cute. Not my favorite. I This is kind of cute. Not my favorite. I like the fonts it chose. It's not like the fonts it chose. It's not like the fonts it chose. It's not picking cringe fonts as much as models picking cringe fonts as much as models picking cringe fonts as much as models used to. used to. used to. And this one's fine. If we turn off the design scale, how If we turn off the design scale, how much worse is it? much worse is it? much worse is it? Yeah, now now we're back into classic LM Yeah, now now we're back into classic LM Yeah, now now we're back into classic LM design. It seems like the new Claude design. It seems like the new Claude design. It seems like the new Claude design skill helps their models a design skill helps their models a design skill helps their models a shitload right now. Yeah, the gap to shitload right now. Yeah, the gap to shitload right now. Yeah, the gap to that to this is hilarious. that to this is hilarious. that to this is hilarious. I had uninstalled the design skill. I I had uninstalled the design skill. I I had uninstalled the design skill. I definitely need to bring it back cuz it definitely need to bring it back cuz it definitely need to bring it back cuz it is it is much better now than it was is it is much better now than it was is it is much better now than it was before. These are great. And just like before. These are great. And just like before. These are great. And just like for comparison, we can look at Grok 4/7. for comparison, we can look at Grok 4/7. for comparison, we can look at Grok 4/7. >> [laughter] >> [laughter] >> [laughter] >> With the design skill off, it's even >> With the design skill off, it's even >> With the design skill off, it's even worse. My god. Like This one still kills me. I I'm still This one still kills me. I I'm still very very amused by this particular one.
-
very very amused by this particular one. very very amused by this particular one. Oh, man. Oh, man. Oh, man. I'm sorry. I just I have to laugh a I'm sorry. I just I have to laugh a I'm sorry. I just I have to laugh a little, okay? So, when in chat said that little, okay? So, when in chat said that little, okay? So, when in chat said that they legitimately think they might have they legitimately think they might have they legitimately think they might have made better sites in raw HTML in made better sites in raw HTML in made better sites in raw HTML in elementary school than what they were elementary school than what they were elementary school than what they were seeing from Grok there, I don't seeing from Grok there, I don't seeing from Grok there, I don't disagree. Regardless, like for the the disagree. Regardless, like for the the disagree. Regardless, like for the the people who are picking Haiku as their people who are picking Haiku as their people who are picking Haiku as their model for cost-sensitive reasons or like model for cost-sensitive reasons or like model for cost-sensitive reasons or like you're using it in the $8 plan with open you're using it in the $8 plan with open you're using it in the $8 plan with open code and such, this is usable. It's not code and such, this is usable. It's not code and such, this is usable. It's not great. Like I'm not going to sit here great. Like I'm not going to sit here great. Like I'm not going to sit here and pretend that this is all you'll ever and pretend that this is all you'll ever and pretend that this is all you'll ever need doing front-end design and building need doing front-end design and building need doing front-end design and building applications, but like applications, but like applications, but like a model this cheap being decent is a model this cheap being decent is a model this cheap being decent is impressive. And it means like the $20 impressive. And it means like the $20 impressive. And it means like the $20 Claude sub might be usable for code. I'm Claude sub might be usable for code. I'm Claude sub might be usable for code. I'm not going to spin one up to test it. I'm not going to spin one up to test it. I'm not going to spin one up to test it. I'm not getting more Claude accounts even not getting more Claude accounts even not getting more Claude accounts even for experiments. I think I have too many for experiments. I think I have too many for experiments. I think I have too many already if I'm being real with y'all, already if I'm being real with y'all, already if I'm being real with y'all, but uh regardless of all that, good but uh regardless of all that, good but uh regardless of all that, good designs. Out of curiosity, I'm throwing designs. Out of curiosity, I'm throwing designs. Out of curiosity, I'm throwing this model at one of my favorite tasks, this model at one of my favorite tasks, this model at one of my favorite tasks, which is auditing my open pull requests. which is auditing my open pull requests. which is auditing my open pull requests. T3 code has over 1,500 PR's right now. T3 code has over 1,500 PR's right now. T3 code has over 1,500 PR's right now. Actually, it's just under 1,500 PR's. Actually, it's just under 1,500 PR's. Actually, it's just under 1,500 PR's. It's probably 1,500 by the time I finish It's probably 1,500 by the time I finish It's probably 1,500 by the time I finish the sentence cuz we have so many coming the sentence cuz we have so many coming the sentence cuz we have so many coming in at all times. So, I have Haiku 55 in at all times. So, I have Haiku 55 in at all times. So, I have Haiku 55 auditing all of them. Specifically in auditing all of them. Specifically in auditing all of them. Specifically in this case, it is addressing all the ones this case, it is addressing all the ones this case, it is addressing all the ones that have performance issues like PR's that have performance issues like PR's that have performance issues like PR's that are fixing performance-related that are fixing performance-related that are fixing performance-related stuff. And I have another thread here stuff. And I have another thread here stuff. And I have another thread here where it's just going to will go through where it's just going to will go through where it's just going to will go through and categorize all of them. Sentiently and categorize all of them. Sentiently and categorize all of them. Sentiently told to use six holes for sub-agents and told to use six holes for sub-agents and told to use six holes for sub-agents and corrected it later. You get the idea corrected it later. You get the idea corrected it later. You get the idea though. This is going to go through all though. This is going to go through all though. This is going to go through all my PR's and make me a nice HTML page my PR's and make me a nice HTML page my PR's and make me a nice HTML page showing them when it's done. While we showing them when it's done. While we showing them when it's done. While we wait for that though, I have my favorite wait for that though, I have my favorite wait for that though, I have my favorite demo.
-
demo. demo. We got a new build of Fish Slap. We got a new build of Fish Slap. We got a new build of Fish Slap. Oof. Oof. Oof. This one has some problems. This one has some problems. This one has some problems. First off, First off, First off, no mouse move. no mouse move. no mouse move. Only moving with space, shift, and WASD. Only moving with space, shift, and WASD. Only moving with space, shift, and WASD. You got to turn. You got to turn. You got to turn. No sound either. I don't know if that's No sound either. I don't know if that's No sound either. I don't know if that's cuz I have it muted. Nope, it's just no cuz I have it muted. Nope, it's just no cuz I have it muted. Nope, it's just no sound. Yeah, it's a little rough. Yeah, it's a little rough. This is roughly the same tier as the This is roughly the same tier as the This is roughly the same tier as the Grok version. Grok version. Grok version. Which was offensively bad. I will say Which was offensively bad. I will say Which was offensively bad. I will say for this one, other than the lack of for this one, other than the lack of for this one, other than the lack of mouse move, mouse move, mouse move, there is no like a just like a there is no like a just like a there is no like a just like a straight-up incorrect offensively bad straight-up incorrect offensively bad straight-up incorrect offensively bad thing. So, for comparison with um GPD6 thing. So, for comparison with um GPD6 thing. So, for comparison with um GPD6 Astra, the demo it made was stunning, Astra, the demo it made was stunning, Astra, the demo it made was stunning, but it got the like core movement wrong. but it got the like core movement wrong. but it got the like core movement wrong. Like the mouse movement was way too slow Like the mouse movement was way too slow Like the mouse movement was way too slow and it felt awful. And the gameplay's and it felt awful. And the gameplay's and it felt awful. And the gameplay's like core loop was a little screwy and like core loop was a little screwy and like core loop was a little screwy and there were like pretty apparent bugs in there were like pretty apparent bugs in there were like pretty apparent bugs in it.
-
it. it. This version doesn't seem to have like This version doesn't seem to have like This version doesn't seem to have like like other than the mouse move thing, like other than the mouse move thing, like other than the mouse move thing, where it just doesn't work at all. where it just doesn't work at all. where it just doesn't work at all. Nothing else is like immediately like Nothing else is like immediately like Nothing else is like immediately like this is just [ __ ] wrong. this is just [ __ ] wrong. this is just [ __ ] wrong. Which is good. It's not It's not making Which is good. It's not It's not making Which is good. It's not It's not making bad decisions. It's just not making bad decisions. It's just not making bad decisions. It's just not making impressive ones. Like, "Oh, yeah, the impressive ones. Like, "Oh, yeah, the impressive ones. Like, "Oh, yeah, the corners are fucked." That's a good corners are fucked." That's a good corners are fucked." That's a good point, actually. Yeah, the way that this point, actually. Yeah, the way that this point, actually. Yeah, the way that this corner works is just bad. Okay, yeah, corner works is just bad. Okay, yeah, corner works is just bad. Okay, yeah, I'm taking back what I was saying there. I'm taking back what I was saying there. I'm taking back what I was saying there. This does have some [ __ ] stuff. It's This does have some [ __ ] stuff. It's This does have some [ __ ] stuff. It's subtle, but it has it. subtle, but it has it. subtle, but it has it. Also, the timing of the second alien Also, the timing of the second alien Also, the timing of the second alien attack made no sense. It shouldn't have attack made no sense. It shouldn't have attack made no sense. It shouldn't have been there that early. It is not been there that early. It is not been there that early. It is not honoring the original source the way I honoring the original source the way I honoring the original source the way I was hoping it would. was hoping it would. was hoping it would. It's not as like immediately offensively It's not as like immediately offensively It's not as like immediately offensively pathetic as a lot of other models do pathetic as a lot of other models do pathetic as a lot of other models do things. Like it's It's lows aren't as things. Like it's It's lows aren't as things. Like it's It's lows aren't as low as Astra's, but it's highs are a low as Astra's, but it's highs are a low as Astra's, but it's highs are a tenth as high as Astra's, if that's I tenth as high as Astra's, if that's I tenth as high as Astra's, if that's I think that's the simplest way I can put think that's the simplest way I can put think that's the simplest way I can put it. That bill of fish slop was about a it. That bill of fish slop was about a it. That bill of fish slop was about a dollar of usage total. It was a shitload dollar of usage total. It was a shitload dollar of usage total. It was a shitload of cash read tokens, 9.73 of cash read tokens, 9.73 of cash read tokens, 9.73 million of them. A bunch of 1-hour cash million of them. A bunch of 1-hour cash million of them. A bunch of 1-hour cash writes, because it always writes 1 hour, writes, because it always writes 1 hour, writes, because it always writes 1 hour, which is annoying. New inputs, which is annoying. New inputs, which is annoying. New inputs, apparently only 102 tokens cuz it was apparently only 102 tokens cuz it was apparently only 102 tokens cuz it was caching constantly. And the output caching constantly. And the output caching constantly. And the output tokens were about 30 cents. 47 of 51 tokens were about 30 cents. 47 of 51 tokens were about 30 cents. 47 of 51 requests exceeded 100k input tokens, so requests exceeded 100k input tokens, so requests exceeded 100k input tokens, so they used Haiku's higher price tier. If they used Haiku's higher price tier. If they used Haiku's higher price tier. If I had capped the context, the game I had capped the context, the game I had capped the context, the game probably would have come out much worse, probably would have come out much worse, probably would have come out much worse, but it would have been way cheaper. Look but it would have been way cheaper. Look but it would have been way cheaper. Look at that. The first PR review is in. They at that. The first PR review is in. They at that. The first PR review is in. They found 40 performance-related open PRs.
-
found 40 performance-related open PRs. found 40 performance-related open PRs. Most are server or mobile fixes. Only a Most are server or mobile fixes. Only a Most are server or mobile fixes. Only a few are clean and high impact. Top of few are clean and high impact. Top of few are clean and high impact. Top of the list below is merge ready cuz GitHub the list below is merge ready cuz GitHub the list below is merge ready cuz GitHub reports clean. So it thinks that they're reports clean. So it thinks that they're reports clean. So it thinks that they're ready just cuz GitHub said so. It didn't ready just cuz GitHub said so. It didn't ready just cuz GitHub said so. It didn't seem to audit them thoroughly itself. seem to audit them thoroughly itself. seem to audit them thoroughly itself. Replaying a command it no longer freezes Replaying a command it no longer freezes Replaying a command it no longer freezes the server. Retry went from 18 seconds the server. Retry went from 18 seconds the server. Retry went from 18 seconds plus to 6 milliseconds. I'll do what I plus to 6 milliseconds. I'll do what I plus to 6 milliseconds. I'll do what I usually do here, which is I open up usually do here, which is I open up usually do here, which is I open up Opus. I paste the PR and I say something Opus. I paste the PR and I say something Opus. I paste the PR and I say something along the lines of this seems like it is along the lines of this seems like it is along the lines of this seems like it is a change worth landing. Can you figure a change worth landing. Can you figure a change worth landing. Can you figure out if that's the case, fix anything out if that's the case, fix anything out if that's the case, fix anything that you think needs to be fixed, and that you think needs to be fixed, and that you think needs to be fixed, and get this merged if you're confident it's get this merged if you're confident it's get this merged if you're confident it's ready to go. Full send. And now either ready to go. Full send. And now either ready to go. Full send. And now either that PR or something like it will be that PR or something like it will be that PR or something like it will be merged momentarily. merged momentarily. merged momentarily. The full send full stop combo has The full send full stop combo has The full send full stop combo has changed how I built. I'm so happy with changed how I built. I'm so happy with changed how I built. I'm so happy with it. While I wait for this PR summary it. While I wait for this PR summary it. While I wait for this PR summary view to generate, I want to answer one view to generate, I want to answer one view to generate, I want to answer one of the questions I see happening a lot of the questions I see happening a lot of the questions I see happening a lot in my chat. Will I use this model for in my chat. Will I use this model for in my chat. Will I use this model for code? No. I I want to be clear about code? No. I I want to be clear about code? No. I I want to be clear about this cuz I think you guys are thinking this cuz I think you guys are thinking this cuz I think you guys are thinking about these things wrong in general. about these things wrong in general. about these things wrong in general. Coding is not what makes LLMs expensive. Coding is not what makes LLMs expensive. Coding is not what makes LLMs expensive. It's not the process of going from a It's not the process of going from a It's not the process of going from a prompt to a JavaScript file you can prompt to a JavaScript file you can prompt to a JavaScript file you can execute that is the majority of your execute that is the majority of your execute that is the majority of your costs. Your costs come from everything costs. Your costs come from everything costs. Your costs come from everything else largely. Like once the context is else largely. Like once the context is else largely. Like once the context is in the model and it's cached, and in the model and it's cached, and in the model and it's cached, and generating the right code file and generating the right code file and generating the right code file and putting it in the right place is putting it in the right place is putting it in the right place is relatively cheap. Getting the context relatively cheap. Getting the context relatively cheap. Getting the context needed to do that is expensive.
-
needed to do that is expensive. needed to do that is expensive. Verifying that it worked is expensive. Verifying that it worked is expensive. Verifying that it worked is expensive. Writing a shitload of code to test all Writing a shitload of code to test all Writing a shitload of code to test all the edges around it, to hook into GitHub the edges around it, to hook into GitHub the edges around it, to hook into GitHub so that you can monitor the PR when it's so that you can monitor the PR when it's so that you can monitor the PR when it's up, and all these other things, those up, and all these other things, those up, and all these other things, those are expensive. Regenerating it eight are expensive. Regenerating it eight are expensive. Regenerating it eight times because you got it wrong the first times because you got it wrong the first times because you got it wrong the first seven, that's expensive. Those things seven, that's expensive. Those things seven, that's expensive. Those things are expensive. But the act of going from are expensive. But the act of going from are expensive. But the act of going from all of the context and details needed to all of the context and details needed to all of the context and details needed to write code to the code file in your code write code to the code file in your code write code to the code file in your code base, even on the most expensive Fable, base, even on the most expensive Fable, base, even on the most expensive Fable, that is a few pennies. It is not that that is a few pennies. It is not that that is a few pennies. It is not that expensive to generate 500 output tokens expensive to generate 500 output tokens expensive to generate 500 output tokens once the input tokens are cached. What once the input tokens are cached. What once the input tokens are cached. What I'm trying to say here is I don't get I'm trying to say here is I don't get I'm trying to say here is I don't get why people are looking for a dumb model why people are looking for a dumb model why people are looking for a dumb model to do the code part. The code is cheap. to do the code part. The code is cheap. to do the code part. The code is cheap. The expensive part is all of the prep The expensive part is all of the prep The expensive part is all of the prep before you write the code file and all before you write the code file and all before you write the code file and all the verification after. And Haiku is a the verification after. And Haiku is a the verification after. And Haiku is a tool your models can use to do those tool your models can use to do those tool your models can use to do those parts more cheaply is valuable. And if parts more cheaply is valuable. And if parts more cheaply is valuable. And if that middle part, the coding, has edges that middle part, the coding, has edges that middle part, the coding, has edges around it, like you want to try five around it, like you want to try five around it, like you want to try five versions of the thing and you have a versions of the thing and you have a versions of the thing and you have a verification system where you know if it verification system where you know if it verification system where you know if it worked or not trivially, where you can worked or not trivially, where you can worked or not trivially, where you can run code to know, Haiku generating five run code to know, Haiku generating five run code to know, Haiku generating five or eight versions to test could be worth or eight versions to test could be worth or eight versions to test could be worth it. But for most problems, the smarter it. But for most problems, the smarter it. But for most problems, the smarter model writing the code is not that model writing the code is not that model writing the code is not that expensive and is preferred in general.
-
expensive and is preferred in general. expensive and is preferred in general. When the model is unsure about a theory When the model is unsure about a theory When the model is unsure about a theory and it wants to vet four things to and it wants to vet four things to and it wants to vet four things to confirm it or wants to go look at eight confirm it or wants to go look at eight confirm it or wants to go look at eight apps in the world and like random apps in the world and like random apps in the world and like random open-source projects to compare against open-source projects to compare against open-source projects to compare against or read through a bunch of documentation or read through a bunch of documentation or read through a bunch of documentation to figure out some facts about some idea to figure out some facts about some idea to figure out some facts about some idea you have. That is when Haiku is you have. That is when Haiku is you have. That is when Haiku is incredible. But the execution of the incredible. But the execution of the incredible. But the execution of the plan, sure, if you write a really plan, sure, if you write a really plan, sure, if you write a really thorough detailed plan with a really thorough detailed plan with a really thorough detailed plan with a really smart model, a dumb model might be able smart model, a dumb model might be able smart model, a dumb model might be able to implement parts of it. Cool. Awesome. to implement parts of it. Cool. Awesome. to implement parts of it. Cool. Awesome. You already spent the money on the plan, You already spent the money on the plan, You already spent the money on the plan, though. though. though. Oof. This is an L for me. Oof. This is an L for me. Oof. This is an L for me. I hit a rate limit on GitHub, probably I hit a rate limit on GitHub, probably I hit a rate limit on GitHub, probably my fault cuz I was doing a lot of these my fault cuz I was doing a lot of these my fault cuz I was doing a lot of these types of things across a lot of data, types of things across a lot of data, types of things across a lot of data, but the script saved the error body as but the script saved the error body as but the script saved the error body as though it was API data and then failed though it was API data and then failed though it was API data and then failed because it would hit errors when it ran because it would hit errors when it ran because it would hit errors when it ran jq on that. The parallel per PR fetches jq on that. The parallel per PR fetches jq on that. The parallel per PR fetches that this run that this run that this run did used most of my 5,000 requests per did used most of my 5,000 requests per did used most of my 5,000 requests per hour, so now I'm locked out of T3 code hour, so now I'm locked out of T3 code hour, so now I'm locked out of T3 code GitHub integrations for a bit. The GitHub integrations for a bit. The GitHub integrations for a bit. The PR-star.json PR-star.json PR-star.json files parsed fine in aggregate, so their files parsed fine in aggregate, so their files parsed fine in aggregate, so their data is valid. Now it's waiting for the data is valid. Now it's waiting for the data is valid. Now it's waiting for the quota to reset, then it will refetch the quota to reset, then it will refetch the quota to reset, then it will refetch the 181 PRs and fix the script so it rejects 181 PRs and fix the script so it rejects 181 PRs and fix the script so it rejects responses that aren't JSON arrays. Then responses that aren't JSON arrays. Then responses that aren't JSON arrays. Then it will recount CI and review. Cool.
-
it will recount CI and review. Cool. it will recount CI and review. Cool. Except for the fact that it trigger Except for the fact that it trigger Except for the fact that it trigger anything to spin itself back up. It said anything to spin itself back up. It said anything to spin itself back up. It said that it's going to wait until the rate that it's going to wait until the rate that it's going to wait until the rate limits up, but it didn't trigger a wait. limits up, but it didn't trigger a wait. limits up, but it didn't trigger a wait. We would show that in the UI if it did. We would show that in the UI if it did. We would show that in the UI if it did. So, this thread is just done until I So, this thread is just done until I So, this thread is just done until I tell it to keep going. tell it to keep going. tell it to keep going. So, So, So, yeah, this this is why I don't like dumb yeah, this this is why I don't like dumb yeah, this this is why I don't like dumb models. models. models. Small models will corner themselves like Small models will corner themselves like Small models will corner themselves like this, make something harder for not just this, make something harder for not just this, make something harder for not just themselves, but for me. Like, this just themselves, but for me. Like, this just themselves, but for me. Like, this just ruined my T3 code experience for at ruined my T3 code experience for at ruined my T3 code experience for at least the next 50 minutes until my quota least the next 50 minutes until my quota least the next 50 minutes until my quota gets reset so I can go back to like gets reset so I can go back to like gets reset so I can go back to like doing PRs in side of T3 code. But, since doing PRs in side of T3 code. But, since doing PRs in side of T3 code. But, since my rate limits are [ __ ] anyways, let's my rate limits are [ __ ] anyways, let's my rate limits are [ __ ] anyways, let's show off how I would actually do this. I show off how I would actually do this. I show off how I would actually do this. I took the same prompt from before, but I took the same prompt from before, but I took the same prompt from before, but I changed two things. changed two things. changed two things. First, I told it to use lots of First, I told it to use lots of First, I told it to use lots of sub-agents and workflows to review these sub-agents and workflows to review these sub-agents and workflows to review these PRs and to use Haiku 55 for all PRs and to use Haiku 55 for all PRs and to use Haiku 55 for all sub-agents throughout the work. To say sub-agents throughout the work. To say sub-agents throughout the work. To say it in one more sentence, minimize work it in one more sentence, minimize work it in one more sentence, minimize work that you do yourself outside of actually that you do yourself outside of actually that you do yourself outside of actually getting the data and orchestrating the getting the data and orchestrating the getting the data and orchestrating the sub-agents. And now, the other change. sub-agents. And now, the other change. sub-agents. And now, the other change. I'm going to run this on Opus, because I'm going to run this on Opus, because I'm going to run this on Opus, because unlike Haiku, Opus is a smart model, and unlike Haiku, Opus is a smart model, and unlike Haiku, Opus is a smart model, and it should not only be able to not hit my it should not only be able to not hit my it should not only be able to not hit my rate limits, it should find ways to work rate limits, it should find ways to work rate limits, it should find ways to work around them, too. Here is where it's around them, too. Here is where it's around them, too. Here is where it's going to hit those rate limits, and it's going to hit those rate limits, and it's going to hit those rate limits, and it's probably going to find a way to work probably going to find a way to work probably going to find a way to work around it. Here it's Here's Opus being around it. Here it's Here's Opus being around it. Here it's Here's Opus being smart.
-
smart. smart. 1,495 open PRs is a huge number to 1,495 open PRs is a huge number to 1,495 open PRs is a huge number to tackle with sub-agents. Maybe batches of tackle with sub-agents. Maybe batches of tackle with sub-agents. Maybe batches of around 15 PRs per Haiku agent will mean around 15 PRs per Haiku agent will mean around 15 PRs per Haiku agent will mean roughly 100 agents running. Beyond the roughly 100 agents running. Beyond the roughly 100 agents running. Beyond the typical small-scale workflow guidelines, typical small-scale workflow guidelines, typical small-scale workflow guidelines, but the user explicitly wants, quote, but the user explicitly wants, quote, but the user explicitly wants, quote, lots of sub-agents. Look at that. It's lots of sub-agents. Look at that. It's lots of sub-agents. Look at that. It's actually being thoughtful. This is going actually being thoughtful. This is going actually being thoughtful. This is going to take longer than I plan on sitting to take longer than I plan on sitting to take longer than I plan on sitting here to show you guys, but uh here to show you guys, but uh here to show you guys, but uh you get the idea. you get the idea. you get the idea. I think that Haiku is incredibly useful I think that Haiku is incredibly useful I think that Haiku is incredibly useful as a tool that you insert into your as a tool that you insert into your as a tool that you insert into your code, and that your agents can use as code, and that your agents can use as code, and that your agents can use as well, but I would not select this model well, but I would not select this model well, but I would not select this model as a tool I use in my coding agents, as a tool I use in my coding agents, as a tool I use in my coding agents, because I'm on the $200 tier, and I because I'm on the $200 tier, and I because I'm on the $200 tier, and I would rather just use Opus. And as big would rather just use Opus. And as big would rather just use Opus. And as big as the price gap is, you have to as the price gap is, you have to as the price gap is, you have to remember that when a model can't do remember that when a model can't do remember that when a model can't do something, and it ends up doing lots of something, and it ends up doing lots of something, and it ends up doing lots of things in the process, it might end up things in the process, it might end up things in the process, it might end up using more money. Cuz if Opus can solve using more money. Cuz if Opus can solve using more money. Cuz if Opus can solve a problem in 100k tokens, and Haiku a problem in 100k tokens, and Haiku a problem in 100k tokens, and Haiku burns a million before it figures out it burns a million before it figures out it burns a million before it figures out it can't, Haiku ends up more expensive. So, can't, Haiku ends up more expensive. So, can't, Haiku ends up more expensive. So, for a lot of work, not only is Haiku not for a lot of work, not only is Haiku not for a lot of work, not only is Haiku not able to do it, Haiku will cost more able to do it, Haiku will cost more able to do it, Haiku will cost more money as it fails to do it. And as money as it fails to do it. And as money as it fails to do it. And as things have continued to improve across things have continued to improve across things have continued to improve across the industry, in particular cost the industry, in particular cost the industry, in particular cost performance with both Anthropic and open performance with both Anthropic and open performance with both Anthropic and open AI, we've ended up in a fun place where AI, we've ended up in a fun place where AI, we've ended up in a fun place where the price gap between Opus 5 low and the price gap between Opus 5 low and the price gap between Opus 5 low and Haiku 5 5 max is not as big as you would Haiku 5 5 max is not as big as you would Haiku 5 5 max is not as big as you would think. It's 20 cents to 55 cents. So, think. It's 20 cents to 55 cents. So, think. It's 20 cents to 55 cents. So, it's around 2x more money for Opus if it's around 2x more money for Opus if it's around 2x more money for Opus if you're willing to deal with low. Also, you're willing to deal with low. Also, you're willing to deal with low. Also, we got the rest of the data for Haiku 5 we got the rest of the data for Haiku 5 we got the rest of the data for Haiku 5 5 because they finally finished running 5 because they finally finished running 5 because they finally finished running it, which makes Sol and Luna still quite
-
it, which makes Sol and Luna still quite it, which makes Sol and Luna still quite compelling with the Pareto line. But, uh compelling with the Pareto line. But, uh compelling with the Pareto line. But, uh it does fill out for the Anthropic it does fill out for the Anthropic it does fill out for the Anthropic lineup quite well here, where lineup quite well here, where lineup quite well here, where have all of Haiku for the cheap, and have all of Haiku for the cheap, and have all of Haiku for the cheap, and most of Opus for the less cheap. most of Opus for the less cheap. most of Opus for the less cheap. This is a compelling lineup. If the only This is a compelling lineup. If the only This is a compelling lineup. If the only models you used were Opus and Haiku, models you used were Opus and Haiku, models you used were Opus and Haiku, you're probably not missing out on a you're probably not missing out on a you're probably not missing out on a whole lot. Now, Sonnet, I showed you whole lot. Now, Sonnet, I showed you whole lot. Now, Sonnet, I showed you guys Sonnet just doesn't offer much with guys Sonnet just doesn't offer much with guys Sonnet just doesn't offer much with the Pareto line. Actually, it offers the Pareto line. Actually, it offers the Pareto line. Actually, it offers nothing now. Maybe just barely on high nothing now. Maybe just barely on high nothing now. Maybe just barely on high there, interesting. But, yeah, I I think there, interesting. But, yeah, I I think there, interesting. But, yeah, I I think you can avoid Sonnet entirely and just you can avoid Sonnet entirely and just you can avoid Sonnet entirely and just use Opus and Haiku and have a very, very use Opus and Haiku and have a very, very use Opus and Haiku and have a very, very good time with your Anthropic models. good time with your Anthropic models. good time with your Anthropic models. But, that leaves us with a question. But, that leaves us with a question. But, that leaves us with a question. When should you use Sol? When should you use Sol? When should you use Sol? Well, first off, as you see here, Sol's Well, first off, as you see here, Sol's Well, first off, as you see here, Sol's actually quite a good bridge from Haiku actually quite a good bridge from Haiku actually quite a good bridge from Haiku over to Opus. It fits in the middle over to Opus. It fits in the middle over to Opus. It fits in the middle there. But, that's not how I would think there. But, that's not how I would think there. But, that's not how I would think of Sol. I wouldn't think of Sol as for of Sol. I wouldn't think of Sol as for of Sol. I wouldn't think of Sol as for things that Haiku is too dumb for, and things that Haiku is too dumb for, and things that Haiku is too dumb for, and Opus is too expensive for. I'd think of Opus is too expensive for. I'd think of Opus is too expensive for. I'd think of Sol as a different family with similar Sol as a different family with similar Sol as a different family with similar intelligence. Imagine that you have a intelligence. Imagine that you have a intelligence. Imagine that you have a team of people that work on the same team of people that work on the same team of people that work on the same things with you every day, and you have things with you every day, and you have things with you every day, and you have a really hard problem you have to solve.
-
a really hard problem you have to solve. a really hard problem you have to solve. And you ask everyone around you, and you And you ask everyone around you, and you And you ask everyone around you, and you all, for the most part, agree on what all, for the most part, agree on what all, for the most part, agree on what the path is. That makes sense because the path is. That makes sense because the path is. That makes sense because you guys work together every day, so you guys work together every day, so you guys work together every day, so you've started to think about things in you've started to think about things in you've started to think about things in a similar way. Soul thinks about things a similar way. Soul thinks about things a similar way. Soul thinks about things in a different way. It's like you're in a different way. It's like you're in a different way. It's like you're asking your friend who works at a asking your friend who works at a asking your friend who works at a different company for their thoughts, different company for their thoughts, different company for their thoughts, and it's a unique perspective. And I and it's a unique perspective. And I and it's a unique perspective. And I found open AI models bring a unique found open AI models bring a unique found open AI models bring a unique perspective in particular with their perspective in particular with their perspective in particular with their thoroughness auditing code. I love 61 thoroughness auditing code. I love 61 thoroughness auditing code. I love 61 Soul as an auditor. I have set up skills Soul as an auditor. I have set up skills Soul as an auditor. I have set up skills for all my agents where they use Opus as for all my agents where they use Opus as for all my agents where they use Opus as the thing that does the coding, and then the thing that does the coding, and then the thing that does the coding, and then they consult with Soul for opinions they consult with Soul for opinions they consult with Soul for opinions before they put up the PR, and I'm before they put up the PR, and I'm before they put up the PR, and I'm landing code so much faster simply landing code so much faster simply landing code so much faster simply because Soul catches all the dumb things because Soul catches all the dumb things because Soul catches all the dumb things Opus might miss. And by the time the PR Opus might miss. And by the time the PR Opus might miss. And by the time the PR is up, it's already been scrubbed by is up, it's already been scrubbed by is up, it's already been scrubbed by Opus and by Soul. So, I personally am Opus and by Soul. So, I personally am Opus and by Soul. So, I personally am sticking to my Opus and Soul duo, but sticking to my Opus and Soul duo, but sticking to my Opus and Soul duo, but when Opus needs to go collect a bunch of when Opus needs to go collect a bunch of when Opus needs to go collect a bunch of data or read through tons of [ __ ] or data or read through tons of [ __ ] or data or read through tons of [ __ ] or just categorize things, Haiku is going just categorize things, Haiku is going just categorize things, Haiku is going to be an incredible addition to my to be an incredible addition to my to be an incredible addition to my portfolio of models. I won't be the one portfolio of models. I won't be the one portfolio of models. I won't be the one to pick it, but I trust Opus to do a to pick it, but I trust Opus to do a to pick it, but I trust Opus to do a good enough job picking it at the times good enough job picking it at the times good enough job picking it at the times where it makes sense. And I think the where it makes sense. And I think the where it makes sense. And I think the result of all of this is a very result of all of this is a very result of all of this is a very compelling set of models. This isn't compelling set of models. This isn't compelling set of models. This isn't going to take too long. I am curious going to take too long. I am curious going to take too long. I am curious what the results are, but uh what the results are, but uh what the results are, but uh I think I've said all I have to on this I think I've said all I have to on this I think I've said all I have to on this model release. It's a useful tool. It's model release. It's a useful tool. It's model release. It's a useful tool. It's a nice change to see Anthropic making a nice change to see Anthropic making a nice change to see Anthropic making small models that are actually usable small models that are actually usable small models that are actually usable and useful for real world stuff, and and useful for real world stuff, and and useful for real world stuff, and it's still honestly a bit trippy seeing it's still honestly a bit trippy seeing it's still honestly a bit trippy seeing an Anthropic model so far to the left on
-
an Anthropic model so far to the left on an Anthropic model so far to the left on the artificial analysis cost versus the artificial analysis cost versus the artificial analysis cost versus intelligence. Right there, neck and neck intelligence. Right there, neck and neck intelligence. Right there, neck and neck with Luna. Actually, I mentioned earlier with Luna. Actually, I mentioned earlier with Luna. Actually, I mentioned earlier the like Pareto line stuff. We only had the like Pareto line stuff. We only had the like Pareto line stuff. We only had X high and max for Haiku at the time. X high and max for Haiku at the time. X high and max for Haiku at the time. Now that they have the rest, it's so Now that they have the rest, it's so Now that they have the rest, it's so close to Luna that this no longer feels close to Luna that this no longer feels close to Luna that this no longer feels like a big missing piece of the like a big missing piece of the like a big missing piece of the Anthropic portfolio. And with that, I Anthropic portfolio. And with that, I Anthropic portfolio. And with that, I will go back to my opening statement. will go back to my opening statement. will go back to my opening statement. Anthropic has no small models that are Anthropic has no small models that are Anthropic has no small models that are worth using right now. Anthropic now has worth using right now. Anthropic now has worth using right now. Anthropic now has one small model that is sometimes worth one small model that is sometimes worth one small model that is sometimes worth using, and that is a very nice change. using, and that is a very nice change. using, and that is a very nice change. And with that, all I have left to say is And with that, all I have left to say is And with that, all I have left to say is that I can't wait for Fable 5.5. I that I can't wait for Fable 5.5. I that I can't wait for Fable 5.5. I understand that they're probably not understand that they're probably not understand that they're probably not going to put it out for a while cuz it going to put it out for a while cuz it going to put it out for a while cuz it kind of is outside of the spirit of kind of is outside of the spirit of kind of is outside of the spirit of pacing the frontier, but I'm so happy pacing the frontier, but I'm so happy pacing the frontier, but I'm so happy with this 5.5 lineup right now that I'm with this 5.5 lineup right now that I'm with this 5.5 lineup right now that I'm really hopeful Fable meets the bar really hopeful Fable meets the bar really hopeful Fable meets the bar they've set with model releases like they've set with model releases like they've set with model releases like Opus, Sonnet, and Haiku as part of this Opus, Sonnet, and Haiku as part of this Opus, Sonnet, and Haiku as part of this line. Anthropic is dominating across the line. Anthropic is dominating across the line. Anthropic is dominating across the field now, and if I was OpenAI, I would field now, and if I was OpenAI, I would field now, and if I was OpenAI, I would be very, very scared, and I would be be very, very scared, and I would be be very, very scared, and I would be praying that the next training run for praying that the next training run for praying that the next training run for Astra comes out unbelievable. Because Astra comes out unbelievable. Because Astra comes out unbelievable. Because right now I'm honestly struggling a bit right now I'm honestly struggling a bit right now I'm honestly struggling a bit to use my Codex subs because my to use my Codex subs because my to use my Codex subs because my Anthropic and Claude subs are just so Anthropic and Claude subs are just so Anthropic and Claude subs are just so much more useful. Let me know if you much more useful. Let me know if you much more useful. Let me know if you guys agree, if I'm over exaggerating guys agree, if I'm over exaggerating guys agree, if I'm over exaggerating this gap, or if it really does feel that this gap, or if it really does feel that this gap, or if it really does feel that way to y'all. Until next time.
-
way to y'all. Until next time. way to y'all. Until next time. Peace, nerds.
No summary available yet.
View original episode ↗