GPT-5.6: The Review
Read full transcript 27 segments
-
GBT 56 is here and it's time for a real GBT 56 is here and it's time for a real review. As I've mentioned in other review. As I've mentioned in other review. As I've mentioned in other videos, I've been using it a ton, but I videos, I've been using it a ton, but I videos, I've been using it a ton, but I don't want to just share my opinions don't want to just share my opinions don't want to just share my opinions anymore. I've talked about that in anymore. I've talked about that in anymore. I've talked about that in plenty of stuff and I will in the near plenty of stuff and I will in the near plenty of stuff and I will in the near future. I'm here for hard numbers and to future. I'm here for hard numbers and to future. I'm here for hard numbers and to do my best aggregating what other people do my best aggregating what other people do my best aggregating what other people think, as well as offering a little bit think, as well as offering a little bit think, as well as offering a little bit of advice on how to use the model, how of advice on how to use the model, how of advice on how to use the model, how to pick between the sole, Luna, and to pick between the sole, Luna, and to pick between the sole, Luna, and Terra versions, the different reasoning Terra versions, the different reasoning Terra versions, the different reasoning levels, ultra, pro, and all of the chaos levels, ultra, pro, and all of the chaos levels, ultra, pro, and all of the chaos OpenAI has given us. Because when you OpenAI has given us. Because when you OpenAI has given us. Because when you multiply out all the different options multiply out all the different options multiply out all the different options in the pro version and all of that, in the pro version and all of that, in the pro version and all of that, there's like 30 plus choices you have there's like 30 plus choices you have there's like 30 plus choices you have here and it's not easy to get right. And here and it's not easy to get right. And here and it's not easy to get right. And I want to do my best to help you while I want to do my best to help you while I want to do my best to help you while also explaining what the strengths and also explaining what the strengths and also explaining what the strengths and weaknesses of this model are. And the weaknesses of this model are. And the weaknesses of this model are. And the strengths are showing themselves pretty strengths are showing themselves pretty strengths are showing themselves pretty clearly when you look at charts like clearly when you look at charts like clearly when you look at charts like deepswe where 56 soul on max got the deepswe where 56 soul on max got the deepswe where 56 soul on max got the highest score ever at 73% while costing highest score ever at 73% while costing highest score ever at 73% while costing under half as much as fable did on its under half as much as fable did on its under half as much as fable did on its max equivalent $22 per task versus $839. max equivalent $22 per task versus $839. max equivalent $22 per task versus $839. And again, Soul got a higher score. And again, Soul got a higher score. And again, Soul got a higher score. Benchmarks are far from the only thing Benchmarks are far from the only thing Benchmarks are far from the only thing that matters here, though. The sentiment that matters here, though. The sentiment that matters here, though. The sentiment on this model is wild. From people on this model is wild. From people on this model is wild. From people saying it's the best thing they've ever saying it's the best thing they've ever saying it's the best thing they've ever used to it's changing how they write used to it's changing how they write used to it's changing how they write software to others saying Fable is so software to others saying Fable is so software to others saying Fable is so much better that they basically stopped much better that they basically stopped much better that they basically stopped using 5.6. There's a wide range of using 5.6. There's a wide range of using 5.6. There's a wide range of opinions and things to go over here and opinions and things to go over here and opinions and things to go over here and I'll do my best to cover all of it right I'll do my best to cover all of it right I'll do my best to cover all of it right after a quick break for today's sponsor.
-
after a quick break for today's sponsor. after a quick break for today's sponsor. I've built a lot of different solutions I've built a lot of different solutions I've built a lot of different solutions to a lot of different problems. I have to a lot of different problems. I have to a lot of different problems. I have so many repos that they're hard to keep so many repos that they're hard to keep so many repos that they're hard to keep track of, but all the ones that matter track of, but all the ones that matter track of, but all the ones that matter have one thing in common. And no, it's have one thing in common. And no, it's have one thing in common. And no, it's not that they use TypeScript. Some of not that they use TypeScript. Some of not that they use TypeScript. Some of them even use other things. It's this them even use other things. It's this them even use other things. It's this little hedgehog that just snuck into my little hedgehog that just snuck into my little hedgehog that just snuck into my inbox. Post HTOG is truly something inbox. Post HTOG is truly something inbox. Post HTOG is truly something else. They're an all-in-one suite of else. They're an all-in-one suite of else. They're an all-in-one suite of product tools that make it way easier to product tools that make it way easier to product tools that make it way easier to understand your users and give them good understand your users and give them good understand your users and give them good experiences. From their analytics to experiences. From their analytics to experiences. From their analytics to feature flags, their data warehouses, feature flags, their data warehouses, feature flags, their data warehouses, and so much other stuff, they were and so much other stuff, they were and so much other stuff, they were always the obvious solution for knowing always the obvious solution for knowing always the obvious solution for knowing your users better when you were building your users better when you were building your users better when you were building real products. And they always had a real products. And they always had a real products. And they always had a nice attitude, too. But recently, their nice attitude, too. But recently, their nice attitude, too. But recently, their attitudes have been changing because of attitudes have been changing because of attitudes have been changing because of AI. And unlike most of the companies in AI. And unlike most of the companies in AI. And unlike most of the companies in the space that are just using AI to give the space that are just using AI to give the space that are just using AI to give you automatic insights that aren't you automatic insights that aren't you automatic insights that aren't particularly useful, Postthog is going particularly useful, Postthog is going particularly useful, Postthog is going full hog with this one, they're leaning full hog with this one, they're leaning full hog with this one, they're leaning in. These guys replace the default in. These guys replace the default in. These guys replace the default dashboard with a chat box that I dashboard with a chat box that I dashboard with a chat box that I honestly thought I would hate. If you honestly thought I would hate. If you honestly thought I would hate. If you don't know this about me, I spent a lot don't know this about me, I spent a lot don't know this about me, I spent a lot of time in analytics dashboards of time in analytics dashboards of time in analytics dashboards throughout my career. I would often use throughout my career. I would often use throughout my career. I would often use them to settle arguments I was in at big them to settle arguments I was in at big them to settle arguments I was in at big companies. and learning things like companies. and learning things like companies. and learning things like click house, SQL and mode and all those click house, SQL and mode and all those click house, SQL and mode and all those tools was super useful for me when I was tools was super useful for me when I was tools was super useful for me when I was working at a real company. As such, at working at a real company. As such, at working at a real company. As such, at my companies, everyone just kind of my companies, everyone just kind of my companies, everyone just kind of expects me to be the guy to go do the expects me to be the guy to go do the expects me to be the guy to go do the analytics. That changed because of how analytics. That changed because of how analytics. That changed because of how much better this tool is. You can just much better this tool is. You can just much better this tool is. You can just ask it to create a chart and get useful ask it to create a chart and get useful ask it to create a chart and get useful insights and it does. Like T3 code is insights and it does. Like T3 code is insights and it does. Like T3 code is used by Linux and Mac and Windows used by Linux and Mac and Windows used by Linux and Mac and Windows people. Let's just ask about this.
-
people. Let's just ask about this. people. Let's just ask about this. What's the split like between Windows, What's the split like between Windows, What's the split like between Windows, Mac, and Linux for T3 code users? Is Mac, and Linux for T3 code users? Is Mac, and Linux for T3 code users? Is there any unique insights you have about there any unique insights you have about there any unique insights you have about how those different groups use T3 code how those different groups use T3 code how those different groups use T3 code differently? Maybe one of them uses the differently? Maybe one of them uses the differently? Maybe one of them uses the app more heavily than the other and now app more heavily than the other and now app more heavily than the other and now it can respond like an agent, but also it can respond like an agent, but also it can respond like an agent, but also write SQL to get information, generate write SQL to get information, generate write SQL to get information, generate charts, and more. Okay, we now have charts, and more. Okay, we now have charts, and more. Okay, we now have Insights, and I'm already learning Insights, and I'm already learning Insights, and I'm already learning things about my users. Apparently, Linux things about my users. Apparently, Linux things about my users. Apparently, Linux had a slump two weeks ago, but it is on had a slump two weeks ago, but it is on had a slump two weeks ago, but it is on the line of overtaking Mac on Apple the line of overtaking Mac on Apple the line of overtaking Mac on Apple Silicon. And also Intel Mac is basically Silicon. And also Intel Mac is basically Silicon. And also Intel Mac is basically none of our usage. I should probably none of our usage. I should probably none of our usage. I should probably just drop it cuz like let's be real, just drop it cuz like let's be real, just drop it cuz like let's be real, Intel Mac is no ple please don't use my Intel Mac is no ple please don't use my Intel Mac is no ple please don't use my app with an Intel Mac. Just go upgrade. app with an Intel Mac. Just go upgrade. app with an Intel Mac. Just go upgrade. It's time. But we can also see how many It's time. But we can also see how many It's time. But we can also see how many threads the average user makes across threads the average user makes across threads the average user makes across platforms. And you can see that Linux platforms. And you can see that Linux platforms. And you can see that Linux users are making more and more threads. users are making more and more threads. users are making more and more threads. Mac OS users are slightly so, but the Mac OS users are slightly so, but the Mac OS users are slightly so, but the Linux users are becoming more and more Linux users are becoming more and more Linux users are becoming more and more valuable. As you might have guessed from valuable. As you might have guessed from valuable. As you might have guessed from my recent Linux arc, you can see the my recent Linux arc, you can see the my recent Linux arc, you can see the average project count distribution per average project count distribution per average project count distribution per OS as well. This is really cool stuff OS as well. This is really cool stuff OS as well. This is really cool stuff for me to learn from. And in a world for me to learn from. And in a world for me to learn from. And in a world where our products are going to become where our products are going to become where our products are going to become more and more self-improving, they need more and more self-improving, they need more and more self-improving, they need to know what to improve. So, you need to know what to improve. So, you need to know what to improve. So, you need the data on what your users are doing.
-
the data on what your users are doing. the data on what your users are doing. And that's what PostG is building for. A And that's what PostG is building for. A And that's what PostG is building for. A future where you know enough about your future where you know enough about your future where you know enough about your users to autonomously improve things for users to autonomously improve things for users to autonomously improve things for them. So, if you're ready to better them. So, if you're ready to better them. So, if you're ready to better understand your users, check them out understand your users, check them out understand your users, check them out now at sidv.link/posto. now at sidv.link/posto. now at sidv.link/posto. Let's start with this unnecessarily Let's start with this unnecessarily Let's start with this unnecessarily beautiful looking blog post. We got all beautiful looking blog post. We got all beautiful looking blog post. We got all three of the celestial bodies they name three of the celestial bodies they name three of the celestial bodies they name things after, the sun, earth, and moon. things after, the sun, earth, and moon. things after, the sun, earth, and moon. Curious if they end up sticking with Curious if they end up sticking with Curious if they end up sticking with Soul, Terra, and Luna going forward. Who Soul, Terra, and Luna going forward. Who Soul, Terra, and Luna going forward. Who knows? So, let's dive into what OpenAI knows? So, let's dive into what OpenAI knows? So, let's dive into what OpenAI has to say. I've actually not read this has to say. I've actually not read this has to say. I've actually not read this yet. We're launching the 5.6 family of yet. We're launching the 5.6 family of yet. We're launching the 5.6 family of models for general availability models for general availability models for general availability following their limited preview. Our new following their limited preview. Our new following their limited preview. Our new flagship, Soul, alongside Terra, which flagship, Soul, alongside Terra, which flagship, Soul, alongside Terra, which is a balanced model for everyday work, is a balanced model for everyday work, is a balanced model for everyday work, and Luna, which is their most and Luna, which is their most and Luna, which is their most costefficient model. The concept of costefficient model. The concept of costefficient model. The concept of dropping three models like this at once dropping three models like this at once dropping three models like this at once is a bit much, and I think it's going to is a bit much, and I think it's going to is a bit much, and I think it's going to confuse a lot of people, and I'll do my confuse a lot of people, and I'll do my confuse a lot of people, and I'll do my best to explain how to think of and use best to explain how to think of and use best to explain how to think of and use all of them later. Most of this article all of them later. Most of this article all of them later. Most of this article is probably going to be focused on Soul, is probably going to be focused on Soul, is probably going to be focused on Soul, though, which I would expect. It's the though, which I would expect. It's the though, which I would expect. It's the big one. It's the one that matters. Soul big one. It's the one that matters. Soul big one. It's the one that matters. Soul sets a new standard for both sets a new standard for both sets a new standard for both intelligence and efficiency, achieving intelligence and efficiency, achieving intelligence and efficiency, achieving state-of-the-art results across coding, state-of-the-art results across coding, state-of-the-art results across coding, knowledge work, cyber security, and knowledge work, cyber security, and knowledge work, cyber security, and science while outperforming previous and science while outperforming previous and science while outperforming previous and competing Frontier models with fewer competing Frontier models with fewer competing Frontier models with fewer tokens and at lower estimated costs. If tokens and at lower estimated costs. If tokens and at lower estimated costs. If you're curious how they make their you're curious how they make their you're curious how they make their models so token efficient, I have one of models so token efficient, I have one of models so token efficient, I have one of my personal favorite videos out about my personal favorite videos out about my personal favorite videos out about that that I put out a few weeks ago.
-
that that I put out a few weeks ago. that that I put out a few weeks ago. Check it out if you haven't. The result Check it out if you haven't. The result Check it out if you haven't. The result of all this work is stronger performance of all this work is stronger performance of all this work is stronger performance per dollar, more successful work at the per dollar, more successful work at the per dollar, more successful work at the same spend, or comparable results at a same spend, or comparable results at a same spend, or comparable results at a lower total cost. We also introduced a lower total cost. We also introduced a lower total cost. We also introduced a new way to accelerate the most demanding new way to accelerate the most demanding new way to accelerate the most demanding work, which is Ultra, their highest work, which is Ultra, their highest work, which is Ultra, their highest capability setting, which coordinates capability setting, which coordinates capability setting, which coordinates multiple agents in parallel to finish multiple agents in parallel to finish multiple agents in parallel to finish complex tasks faster. Going to do a complex tasks faster. Going to do a complex tasks faster. Going to do a dedicated Ultra video as well, cuz dedicated Ultra video as well, cuz dedicated Ultra video as well, cuz people really struggle to understand people really struggle to understand people really struggle to understand what it is and what it brings. Stronger what it is and what it brings. Stronger what it is and what it brings. Stronger computer use and design judgment make 56 computer use and design judgment make 56 computer use and design judgment make 56 soul our most polished collaborator yet, soul our most polished collaborator yet, soul our most polished collaborator yet, helping it inspect, refine, and deliver helping it inspect, refine, and deliver helping it inspect, refine, and deliver readyto-use results. We've trained 56 to readyto-use results. We've trained 56 to readyto-use results. We've trained 56 to get more useful work from every token. get more useful work from every token. get more useful work from every token. On the agents last exam, which is a On the agents last exam, which is a On the agents last exam, which is a newer bench that they've been working newer bench that they've been working newer bench that they've been working on, which is an evaluation of longunning on, which is an evaluation of longunning on, which is an evaluation of longunning professional workflows across 55 fields, professional workflows across 55 fields, professional workflows across 55 fields, Saul set a new high of 53.6, eclipsing Saul set a new high of 53.6, eclipsing Saul set a new high of 53.6, eclipsing Fable 5 with adaptive reasoning by 13 Fable 5 with adaptive reasoning by 13 Fable 5 with adaptive reasoning by 13 points. Even at medium reasoning, it points. Even at medium reasoning, it points. Even at medium reasoning, it beats Fable 5 by 11 points at roughly beats Fable 5 by 11 points at roughly beats Fable 5 by 11 points at roughly one quarter of the estimated cost. That one quarter of the estimated cost. That one quarter of the estimated cost. That efficiency extends to smaller models, efficiency extends to smaller models, efficiency extends to smaller models, which are essential to making which are essential to making which are essential to making intelligence more abundant and intelligence more abundant and intelligence more abundant and affordable. 56 Terra and Luna outperform affordable. 56 Terra and Luna outperform affordable. 56 Terra and Luna outperform Fable 5 at around 116th the cost. On the Fable 5 at around 116th the cost. On the Fable 5 at around 116th the cost. On the artificial analysis intelligence index, artificial analysis intelligence index, artificial analysis intelligence index, a broad measure of intelligence spanning a broad measure of intelligence spanning a broad measure of intelligence spanning agentic work, coding, scientific agentic work, coding, scientific agentic work, coding, scientific reasoning, and general capabilities, 56 reasoning, and general capabilities, 56 reasoning, and general capabilities, 56 soul with max reasoning comes within one soul with max reasoning comes within one soul with max reasoning comes within one point of fable 5 while completing tasks point of fable 5 while completing tasks point of fable 5 while completing tasks 61% less time at roughly half of the 61% less time at roughly half of the 61% less time at roughly half of the estimated cost. I'm guessing for agent estimated cost. I'm guessing for agent estimated cost. I'm guessing for agent last exam that there's a lot of computer last exam that there's a lot of computer last exam that there's a lot of computer use stuff here. Yeah, most major fields use stuff here. Yeah, most major fields use stuff here. Yeah, most major fields professional work performed on a professional work performed on a professional work performed on a computer because OpenAI's models are way
-
computer because OpenAI's models are way computer because OpenAI's models are way better at computer use in general, but better at computer use in general, but better at computer use in general, but especially now with 56. God, this agent especially now with 56. God, this agent especially now with 56. God, this agent last exam score is terrifying. Soul on X last exam score is terrifying. Soul on X last exam score is terrifying. Soul on X high outperformed soul on max. That high outperformed soul on max. That high outperformed soul on max. That checks out for reasons we'll discuss. checks out for reasons we'll discuss. checks out for reasons we'll discuss. The X high version got that 53.6% The X high version got that 53.6% The X high version got that 53.6% at around 760 bucks. Meanwhile, the best at around 760 bucks. Meanwhile, the best at around 760 bucks. Meanwhile, the best run Fable got was at uh $3,985 run Fable got was at uh $3,985 run Fable got was at uh $3,985 and it got a 45. Oh, no. That was Opus. and it got a 45. Oh, no. That was Opus. and it got a 45. Oh, no. That was Opus. Fable's in the middle here with adaptive Fable's in the middle here with adaptive Fable's in the middle here with adaptive reasoning at $2,300. So, Opus reasoning at $2,300. So, Opus reasoning at $2,300. So, Opus outperformed, but also outcost, too. It outperformed, but also outcost, too. It outperformed, but also outcost, too. It has crazy safeguards, yada yada, you has crazy safeguards, yada yada, you has crazy safeguards, yada yada, you know, all that. Let's talk about the know, all that. Let's talk about the know, all that. Let's talk about the efficiency and performance on demand. On efficiency and performance on demand. On efficiency and performance on demand. On the artificial analysis coding agent the artificial analysis coding agent the artificial analysis coding agent index, 56 soul with max reasoning sets a index, 56 soul with max reasoning sets a index, 56 soul with max reasoning sets a new state-of-the-art at 20 points. That new state-of-the-art at 20 points. That new state-of-the-art at 20 points. That is a pretty big leap actually. Yeah, the is a pretty big leap actually. Yeah, the is a pretty big leap actually. Yeah, the next highest was Fable at a 77 and then next highest was Fable at a 77 and then next highest was Fable at a 77 and then Grock build at a 76. Grock build was Grock build at a 76. Grock build was Grock build at a 76. Grock build was tied with 55 if I recall, but 56 is now tied with 55 if I recall, but 56 is now tied with 55 if I recall, but 56 is now meaningfully leading specifically on max meaningfully leading specifically on max meaningfully leading specifically on max though. Terra also tied with Fable though. Terra also tied with Fable though. Terra also tied with Fable there, which is kind of crazy. I want to there, which is kind of crazy. I want to there, which is kind of crazy. I want to see costs. 56 Soul was so expensive it see costs. 56 Soul was so expensive it see costs. 56 Soul was so expensive it even beat out really expensive options even beat out really expensive options even beat out really expensive options like GLM52. If that confuses you because like GLM52. If that confuses you because like GLM52. If that confuses you because you heard GLM52 is cheap, you need to you heard GLM52 is cheap, you need to you heard GLM52 is cheap, you need to watch more of my videos. Man, I covered watch more of my videos. Man, I covered watch more of my videos. Man, I covered that a lot at this point. GLM52 is so that a lot at this point. GLM52 is so that a lot at this point. GLM52 is so token inefficient that even at its token inefficient that even at its token inefficient that even at its cheaper cost per token, it ends up being cheaper cost per token, it ends up being cheaper cost per token, it ends up being more expensive and way slower. Sadly, it more expensive and way slower. Sadly, it more expensive and way slower. Sadly, it looks like they don't have scores in for
-
looks like they don't have scores in for looks like they don't have scores in for medium, high, and other options for Soul medium, high, and other options for Soul medium, high, and other options for Soul just yet. But Terra, which was the the just yet. But Terra, which was the the just yet. But Terra, which was the the third highest score, if I recall, was third highest score, if I recall, was third highest score, if I recall, was $2.76, making it cheaper than almost any $2.76, making it cheaper than almost any $2.76, making it cheaper than almost any other Frontier thing they've tested. other Frontier thing they've tested. other Frontier thing they've tested. It's neck andneck with 55 while getting It's neck andneck with 55 while getting It's neck andneck with 55 while getting a slightly higher score. It seems like a slightly higher score. It seems like a slightly higher score. It seems like Terra is an underrated GOAT for a lot of Terra is an underrated GOAT for a lot of Terra is an underrated GOAT for a lot of this. As I was reading though, they got this. As I was reading though, they got this. As I was reading though, they got state-of-the-art on Terminal Bench 21 as state-of-the-art on Terminal Bench 21 as state-of-the-art on Terminal Bench 21 as well as Deepsw SWE. And they show a lot well as Deepsw SWE. And they show a lot well as Deepsw SWE. And they show a lot of these numbers here. And again, 56 of these numbers here. And again, 56 of these numbers here. And again, 56 Soul, the benchmark leader, the Soul, the benchmark leader, the Soul, the benchmark leader, the state-of-the-art, got the highest score state-of-the-art, got the highest score state-of-the-art, got the highest score we've ever seen on this on high for we've ever seen on this on high for we've ever seen on this on high for 1,400 bucks, whereas Fable cost $3,700 1,400 bucks, whereas Fable cost $3,700 1,400 bucks, whereas Fable cost $3,700 and got a slightly lower score. Oh, no. and got a slightly lower score. Oh, no. and got a slightly lower score. Oh, no. They're basically tied. 77.2 versus They're basically tied. 77.2 versus They're basically tied. 77.2 versus 77.1. Yeah, not great. But then X High 77.1. Yeah, not great. But then X High 77.1. Yeah, not great. But then X High scored 78. And of course, the Max, which scored 78. And of course, the Max, which scored 78. And of course, the Max, which got an 80. I think the cost difference got an 80. I think the cost difference got an 80. I think the cost difference between these isn't worth it and we'll between these isn't worth it and we'll between these isn't worth it and we'll talk about that later on. But uh 56 high talk about that later on. But uh 56 high talk about that later on. But uh 56 high very very good option and it seems like very very good option and it seems like very very good option and it seems like Terra on XH high and max might also be Terra on XH high and max might also be Terra on XH high and max might also be as well. 56 can write and run as well. 56 can write and run as well. 56 can write and run lightweight programs that coordinate lightweight programs that coordinate lightweight programs that coordinate tools. Big deal there. It can process tools. Big deal there. It can process tools. Big deal there. It can process intermediate results, monitor progress, intermediate results, monitor progress, intermediate results, monitor progress, and choose the next action as work and choose the next action as work and choose the next action as work unfolds. This lets tool heavy tasks unfolds. This lets tool heavy tasks unfolds. This lets tool heavy tasks advance with fewer tokens, fewer model advance with fewer tokens, fewer model advance with fewer tokens, fewer model round trips, and less guidance. Instead round trips, and less guidance. Instead round trips, and less guidance. Instead of requiring devs to script every step of requiring devs to script every step of requiring devs to script every step or pass every tool response back through or pass every tool response back through or pass every tool response back through the model, the programmatic tool calling
-
the model, the programmatic tool calling the model, the programmatic tool calling option that's built into the responses option that's built into the responses option that's built into the responses API can now filter large amounts of API can now filter large amounts of API can now filter large amounts of intermediate data, retaining only what intermediate data, retaining only what intermediate data, retaining only what matters, and adapt its workflow along matters, and adapt its workflow along matters, and adapt its workflow along the way. Check out my other videos on the way. Check out my other videos on the way. Check out my other videos on code mode. If you don't know what code mode. If you don't know what code mode. If you don't know what programmatic tool calling is, the super programmatic tool calling is, the super programmatic tool calling is, the super quick TLDDR is that when you have the quick TLDDR is that when you have the quick TLDDR is that when you have the agent call each tool individually, you agent call each tool individually, you agent call each tool individually, you end up filling context with nonsense. If end up filling context with nonsense. If end up filling context with nonsense. If you do it programmatically, you can do you do it programmatically, you can do you do it programmatically, you can do things like query the database, filter, things like query the database, filter, things like query the database, filter, bind to the three rows that matter, and bind to the three rows that matter, and bind to the three rows that matter, and then only pass those back to the model then only pass those back to the model then only pass those back to the model instead. They talk a bit about the max instead. They talk a bit about the max instead. They talk a bit about the max option here, which basically turns off option here, which basically turns off option here, which basically turns off the super efficient post training. I'm the super efficient post training. I'm the super efficient post training. I'm sure that it still does the Grug mode in sure that it still does the Grug mode in sure that it still does the Grug mode in its reasoning traces, but it lets the its reasoning traces, but it lets the its reasoning traces, but it lets the model go way, way longer. So much longer model go way, way longer. So much longer model go way, way longer. So much longer that it ate through my 5 hour window in that it ate through my 5 hour window in that it ate through my 5 hour window in like 50 minutes, not even. Probably like 50 minutes, not even. Probably like 50 minutes, not even. Probably closer to 30. Be very careful with both closer to 30. Be very careful with both closer to 30. Be very careful with both max and ultra. We'll talk about them in max and ultra. We'll talk about them in max and ultra. We'll talk about them in the future. The next section we have is the future. The next section we have is the future. The next section we have is design. And I know this is one of the design. And I know this is one of the design. And I know this is one of the questions y'all have been asking about questions y'all have been asking about questions y'all have been asking about the most. Is it better at design? From the most. Is it better at design? From the most. Is it better at design? From my experience, it is definitely better my experience, it is definitely better my experience, it is definitely better than GPT55, although that is not saying than GPT55, although that is not saying than GPT55, although that is not saying much. Can it make good novel designs?
-
much. Can it make good novel designs? much. Can it make good novel designs? That's that's a little tougher. If you That's that's a little tougher. If you That's that's a little tougher. If you steer it carefully, you can get good steer it carefully, you can get good steer it carefully, you can get good designs out of it, but you can also get designs out of it, but you can also get designs out of it, but you can also get some absolute slop. And believe me, I've some absolute slop. And believe me, I've some absolute slop. And believe me, I've seen some slop. Even just this UI that seen some slop. Even just this UI that seen some slop. Even just this UI that it made for summarizing all the work I it made for summarizing all the work I it made for summarizing all the work I did with the model, which uh yeah, did a did with the model, which uh yeah, did a did with the model, which uh yeah, did a lot of work with this model and the lot of work with this model and the lot of work with this model and the window I had to use it for. I had to do window I had to use it for. I had to do window I had to use it for. I had to do some refinement passes and this is as some refinement passes and this is as some refinement passes and this is as unugly as I could make it quickly. Our unugly as I could make it quickly. Our unugly as I could make it quickly. Our boy Dra did some designs using 56. Thank boy Dra did some designs using 56. Thank boy Dra did some designs using 56. Thank you again, Dar, for the help here. And you again, Dar, for the help here. And you again, Dar, for the help here. And you can see the things it made. They're you can see the things it made. They're you can see the things it made. They're definitely less bad than I would have definitely less bad than I would have definitely less bad than I would have expected from an OpenAI model. This is expected from an OpenAI model. This is expected from an OpenAI model. This is also all with the design skill on. So I also all with the design skill on. So I also all with the design skill on. So I will turn that off and see what we got will turn that off and see what we got will turn that off and see what we got instead. Oh god, that first one here. I instead. Oh god, that first one here. I instead. Oh god, that first one here. I hate what it did with the font there. hate what it did with the font there. hate what it did with the font there. That's awful. This is fine. This is This That's awful. This is fine. This is This That's awful. This is fine. This is This has some ideas that are cool in it, but has some ideas that are cool in it, but has some ideas that are cool in it, but I don't love it. And this is garbage. I don't love it. And this is garbage. I don't love it. And this is garbage. Yeah. Again, without steering, this Yeah. Again, without steering, this Yeah. Again, without steering, this model is not going to make things that model is not going to make things that model is not going to make things that look good. To compare with 55 though, look good. To compare with 55 though, look good. To compare with 55 though, yeah, all of these came out nearly yeah, all of these came out nearly yeah, all of these came out nearly identical before. So, it's it's it's not identical before. So, it's it's it's not identical before. So, it's it's it's not [ __ ] the same way, but it's still not [ __ ] the same way, but it's still not [ __ ] the same way, but it's still not good. It still needs a lot of good. It still needs a lot of good. It still needs a lot of handholding. Thankfully, much like other handholding. Thankfully, much like other handholding. Thankfully, much like other OpenAI models, it listens when you tell OpenAI models, it listens when you tell OpenAI models, it listens when you tell it [ __ ] so you can steer it towards it [ __ ] so you can steer it towards it [ __ ] so you can steer it towards better designs if you have an idea in better designs if you have an idea in better designs if you have an idea in your head, but it's not going to bring your head, but it's not going to bring your head, but it's not going to bring you a good design. You have to point it you a good design. You have to point it you a good design. You have to point it towards one. I made this museum website.
-
towards one. I made this museum website. towards one. I made this museum website. It really likes doing this marquee style It really likes doing this marquee style It really likes doing this marquee style thing. I've seen that on a ton of things thing. I've seen that on a ton of things thing. I've seen that on a ton of things it designed. It's pretty good at 3D. I it designed. It's pretty good at 3D. I it designed. It's pretty good at 3D. I talk about this a bit in my all the talk about this a bit in my all the talk about this a bit in my all the things I built video. I'm surprised at things I built video. I'm surprised at things I built video. I'm surprised at how well it can like handle 3D spaces, how well it can like handle 3D spaces, how well it can like handle 3D spaces, but it's far from like really good at but it's far from like really good at but it's far from like really good at design. design. design. It does not surprise me at all that the It does not surprise me at all that the It does not surprise me at all that the end to end knowledge work is way better, end to end knowledge work is way better, end to end knowledge work is way better, too, because again, computer use is a too, because again, computer use is a too, because again, computer use is a massive improvement, which tends to be a massive improvement, which tends to be a massive improvement, which tends to be a huge part of these types of work. I got huge part of these types of work. I got huge part of these types of work. I got two more Macs because I wanted to let two more Macs because I wanted to let two more Macs because I wanted to let Codeex control them entirely, and I do Codeex control them entirely, and I do Codeex control them entirely, and I do not regret it at all. It has been not regret it at all. It has been not regret it at all. It has been awesome. I'm almost at the point where awesome. I'm almost at the point where awesome. I'm almost at the point where I'm willing to give it access to my I'm willing to give it access to my I'm willing to give it access to my Gmail, and I never thought I would get Gmail, and I never thought I would get Gmail, and I never thought I would get there. Here's the browse comp benchmark, there. Here's the browse comp benchmark, there. Here's the browse comp benchmark, and they finally fixed this. So, it and they finally fixed this. So, it and they finally fixed this. So, it actually has the different options. We actually has the different options. We actually has the different options. We have Soul Ultra here, industryleading at have Soul Ultra here, industryleading at have Soul Ultra here, industryleading at $12.17 per task, but a 92% pass. Google $12.17 per task, but a 92% pass. Google $12.17 per task, but a 92% pass. Google and Anthropic didn't give costs for and Anthropic didn't give costs for and Anthropic didn't give costs for theirs, so we don't actually know what theirs, so we don't actually know what theirs, so we don't actually know what it costs them to run, but you can see it costs them to run, but you can see it costs them to run, but you can see the massive improvement here. both how the massive improvement here. both how the massive improvement here. both how much cheaper the runs are with Soul to much cheaper the runs are with Soul to much cheaper the runs are with Soul to complete the work compared to things complete the work compared to things complete the work compared to things like 55, but also the improvement in like 55, but also the improvement in like 55, but also the improvement in score. If you look at this with latency, score. If you look at this with latency, score. If you look at this with latency, it's even cooler. You can see something it's even cooler. You can see something it's even cooler. You can see something like 56 medium can complete the tasks in like 56 medium can complete the tasks in like 56 medium can complete the tasks in 2 minutes instead of 10, which feels so 2 minutes instead of 10, which feels so 2 minutes instead of 10, which feels so much better. But also, when you use much better. But also, when you use much better. But also, when you use ultra, you're going to be burning a ultra, you're going to be burning a ultra, you're going to be burning a shitload of tokens. So, be careful about shitload of tokens. So, be careful about shitload of tokens. So, be careful about that. Yeah. Wait, how is 56 soul so much that. Yeah. Wait, how is 56 soul so much that. Yeah. Wait, how is 56 soul so much cheaper when it did so many more output cheaper when it did so many more output cheaper when it did so many more output tokens? Is it just super input tokens? Is it just super input tokens? Is it just super input tokenheavy? That's a little weird. I tokenheavy? That's a little weird. I tokenheavy? That's a little weird. I want more info on those numbers later.
-
want more info on those numbers later. want more info on those numbers later. It's also good at spreadsheets. I can't It's also good at spreadsheets. I can't It's also good at spreadsheets. I can't believe they put this in as like a believe they put this in as like a believe they put this in as like a sincere example with this opening slide. sincere example with this opening slide. sincere example with this opening slide. Pushing the frontier on cyber and Pushing the frontier on cyber and Pushing the frontier on cyber and science. It did well in exploit bench science. It did well in exploit bench science. It did well in exploit bench and exploit gym yada yada. It's smart. and exploit gym yada yada. It's smart. and exploit gym yada yada. It's smart. It's good at these things. It's not as It's good at these things. It's not as It's good at these things. It's not as good as Mythos 5 was allegedly. Again, good as Mythos 5 was allegedly. Again, good as Mythos 5 was allegedly. Again, we don't have access. We can't test we don't have access. We can't test we don't have access. We can't test that. But it is still way better than it that. But it is still way better than it that. But it is still way better than it was before. seems like cyber security was before. seems like cyber security was before. seems like cyber security stuff mythos is still the best but we stuff mythos is still the best but we stuff mythos is still the best but we can't use it. So yeah, Ganbench Pro and can't use it. So yeah, Ganbench Pro and can't use it. So yeah, Ganbench Pro and Lifesside Bench at Slaughter and Medcam Lifesside Bench at Slaughter and Medcam Lifesside Bench at Slaughter and Medcam Bench it also seems to have done very Bench it also seems to have done very Bench it also seems to have done very well on. And now they talk about how well on. And now they talk about how well on. And now they talk about how they used it internally which I think is they used it internally which I think is they used it internally which I think is one of the most interesting pieces of one of the most interesting pieces of one of the most interesting pieces of these blog posts. In particular, this these blog posts. In particular, this these blog posts. In particular, this section here over the past 6 months the section here over the past 6 months the section here over the past 6 months the share of research compute devoted to share of research compute devoted to share of research compute devoted to internal coding inference grew 100fold internal coding inference grew 100fold internal coding inference grew 100fold while internal agentic token usage while internal agentic token usage while internal agentic token usage increased approximately 22fold. These increased approximately 22fold. These increased approximately 22fold. These adoption metrics do not measure research adoption metrics do not measure research adoption metrics do not measure research progress on their own, but they show how progress on their own, but they show how progress on their own, but they show how rapidly AI assistance is increasing for rapidly AI assistance is increasing for rapidly AI assistance is increasing for research and across other teams like research and across other teams like research and across other teams like sales, marketing, user ops, finance, and sales, marketing, user ops, finance, and sales, marketing, user ops, finance, and more. So, it looks like the research more. So, it looks like the research more. So, it looks like the research team is using the models way more than team is using the models way more than team is using the models way more than they ever have. 5.6's daily output they ever have. 5.6's daily output they ever have. 5.6's daily output tokens per active researcher was more tokens per active researcher was more tokens per active researcher was more than twice the highest levels observed than twice the highest levels observed than twice the highest levels observed with 55. And that lines up with my with 55. And that lines up with my with 55. And that lines up with my experience, too. I've never burned quite experience, too. I've never burned quite experience, too. I've never burned quite as many tokens as I did with 56. Not as many tokens as I did with 56. Not as many tokens as I did with 56. Not because it's super inefficient, just cuz because it's super inefficient, just cuz because it's super inefficient, just cuz I want to throw it at everything. There I want to throw it at everything. There I want to throw it at everything. There are some problems with the safety stuff.
-
are some problems with the safety stuff. are some problems with the safety stuff. I've even experienced this. I had it I've even experienced this. I had it I've even experienced this. I had it block a request I did trying to clean up block a request I did trying to clean up block a request I did trying to clean up my lake bed code base. They called out my lake bed code base. They called out my lake bed code base. They called out that compared with previous models, 56 that compared with previous models, 56 that compared with previous models, 56 souls cyber safeguards block roughly 10 souls cyber safeguards block roughly 10 souls cyber safeguards block roughly 10 times more potential harmful activity. times more potential harmful activity. times more potential harmful activity. That is bad. Because these measures can That is bad. Because these measures can That is bad. Because these measures can create friction for benign use. We create friction for benign use. We create friction for benign use. We provide an option in chatgpt and codecs provide an option in chatgpt and codecs provide an option in chatgpt and codecs to easily retry prompts on lower to easily retry prompts on lower to easily retry prompts on lower capability models. This feature was also capability models. This feature was also capability models. This feature was also really broken when they first shipped it really broken when they first shipped it really broken when they first shipped it and I gave them a lot of feedback, but and I gave them a lot of feedback, but and I gave them a lot of feedback, but it seems to be better since they plan to it seems to be better since they plan to it seems to be better since they plan to keep shifting this around over time and keep shifting this around over time and keep shifting this around over time and making it less likely to block. But for making it less likely to block. But for making it less likely to block. But for now, it's pretty aggressive. They're now, it's pretty aggressive. They're now, it's pretty aggressive. They're rolling out all of the models to rolling out all of the models to rolling out all of the models to basically all the paid accounts, freeand basically all the paid accounts, freeand basically all the paid accounts, freeand go users, so people on that like I think go users, so people on that like I think go users, so people on that like I think it's five or $8 tier, they only get it's five or $8 tier, they only get it's five or $8 tier, they only get Terara. They don't get Soul, but they Terara. They don't get Soul, but they Terara. They don't get Soul, but they get Terara and obviously Luna. And then get Terara and obviously Luna. And then get Terara and obviously Luna. And then everything else available to pretty much everything else available to pretty much everything else available to pretty much everyone else, especially in codecs. And everyone else, especially in codecs. And everyone else, especially in codecs. And the API rates are what we discussed the API rates are what we discussed the API rates are what we discussed before. $5 per mill in, 30 per mill out before. $5 per mill in, 30 per mill out before. $5 per mill in, 30 per mill out for Soul. 250 per mill in and 15 per for Soul. 250 per mill in and 15 per for Soul. 250 per mill in and 15 per mill out for Terra and dollar per mill mill out for Terra and dollar per mill mill out for Terra and dollar per mill in, $6 per mill out for Luna. They also in, $6 per mill out for Luna. They also in, $6 per mill out for Luna. They also have better and more predictable prompt have better and more predictable prompt have better and more predictable prompt cashing, but that comes at the cost of cashing, but that comes at the cost of cashing, but that comes at the cost of them billing for cash rates, which they them billing for cash rates, which they them billing for cash rates, which they didn't do before, which is going to be a didn't do before, which is going to be a didn't do before, which is going to be a net increase in cost, sadly. So, by the net increase in cost, sadly. So, by the net increase in cost, sadly. So, by the judge of this, the model slaughters, and judge of this, the model slaughters, and judge of this, the model slaughters, and it's obviously so much better, right?
-
it's obviously so much better, right? it's obviously so much better, right? Well, let's go talk about what others Well, let's go talk about what others Well, let's go talk about what others had to say. One of the most interesting had to say. One of the most interesting had to say. One of the most interesting details is how long people have had it details is how long people have had it details is how long people have had it for. Many of the early testers have had for. Many of the early testers have had for. Many of the early testers have had it since May 27th. But also, when the it since May 27th. But also, when the it since May 27th. But also, when the original announcement happened, a lot of original announcement happened, a lot of original announcement happened, a lot of the early testers lost access. So, we the early testers lost access. So, we the early testers lost access. So, we had to go through the strange experience had to go through the strange experience had to go through the strange experience of getting used to this model and then of getting used to this model and then of getting used to this model and then losing access and having to get over the losing access and having to get over the losing access and having to get over the fact that we didn't have it anymore. And fact that we didn't have it anymore. And fact that we didn't have it anymore. And falling back has never felt worse. Dax falling back has never felt worse. Dax falling back has never felt worse. Dax said the following. I've never hyped a said the following. I've never hyped a said the following. I've never hyped a model release. We're generally model release. We're generally model release. We're generally conservative with how we use these conservative with how we use these conservative with how we use these things, but 56 has had a massive impact things, but 56 has had a massive impact things, but 56 has had a massive impact on our team. We're using five times the on our team. We're using five times the on our team. We're using five times the tokens that we used to. It's not even tokens that we used to. It's not even tokens that we used to. It's not even smarter than Fable or anything, but it's smarter than Fable or anything, but it's smarter than Fable or anything, but it's just so reliable and fun to use. He also just so reliable and fun to use. He also just so reliable and fun to use. He also calls out how miserable people were when calls out how miserable people were when calls out how miserable people were when they lost access. Jay, his boss, said they lost access. Jay, his boss, said they lost access. Jay, his boss, said even more on this. I don't want to talk even more on this. I don't want to talk even more on this. I don't want to talk too much about 5.6 versus Fable because too much about 5.6 versus Fable because too much about 5.6 versus Fable because that's going to be a deeper dive in the that's going to be a deeper dive in the that's going to be a deeper dive in the near future. So, for now, we're just near future. So, for now, we're just near future. So, for now, we're just going to cover 56 itself, but I want to going to cover 56 itself, but I want to going to cover 56 itself, but I want to talk about this. They tested early talk about this. They tested early talk about this. They tested early versions of 56 for a couple weeks. Had a versions of 56 for a couple weeks. Had a versions of 56 for a couple weeks. Had a great time. It felt like a step change great time. It felt like a step change great time. It felt like a step change improvement enabling new workflows. They improvement enabling new workflows. They improvement enabling new workflows. They tried Fable and don't think it's not as tried Fable and don't think it's not as tried Fable and don't think it's not as good. Personally would have taken the good. Personally would have taken the good. Personally would have taken the experience with a grain of salt. Tends experience with a grain of salt. Tends experience with a grain of salt. Tends to be a bias on trying something new.
-
to be a bias on trying something new. to be a bias on trying something new. Fable and 56 are taken away because of Fable and 56 are taken away because of Fable and 56 are taken away because of regulatory issues. Team is literally regulatory issues. Team is literally regulatory issues. Team is literally depressed that 56 is gone. We're looking depressed that 56 is gone. We're looking depressed that 56 is gone. We're looking for anything that could even partly for anything that could even partly for anything that could even partly replace it. I actually had the same replace it. I actually had the same replace it. I actually had the same thing where I tried to make skills to thing where I tried to make skills to thing where I tried to make skills to distill the things I liked about 56 into distill the things I liked about 56 into distill the things I liked about 56 into 55. Not needed at all anymore. Fable 55. Not needed at all anymore. Fable 55. Not needed at all anymore. Fable comes back and here's where it gets comes back and here's where it gets comes back and here's where it gets interesting. You would think Fable would interesting. You would think Fable would interesting. You would think Fable would be enough, but no. the team is still be enough, but no. the team is still be enough, but no. the team is still depressed that 56 isn't available and depressed that 56 isn't available and depressed that 56 isn't available and then 56 is back and it's immediately then 56 is back and it's immediately then 56 is back and it's immediately clear to them that it's just better than clear to them that it's just better than clear to them that it's just better than Fable. I have a lot of controversial Fable. I have a lot of controversial Fable. I have a lot of controversial thoughts there. I wanted to talk about thoughts there. I wanted to talk about thoughts there. I wanted to talk about how much they it sucked losing the how much they it sucked losing the how much they it sucked losing the models though because it made it really models though because it made it really models though because it made it really clear how much better they were. I like clear how much better they were. I like clear how much better they were. I like what Max had to say here. He thinks the what Max had to say here. He thinks the what Max had to say here. He thinks the most impressive part about this model is most impressive part about this model is most impressive part about this model is that it never gives up. If you throw it that it never gives up. If you throw it that it never gives up. If you throw it in Max's reasoning, it will just keep in Max's reasoning, it will just keep in Max's reasoning, it will just keep working until it's done. And as such, it working until it's done. And as such, it working until it's done. And as such, it is his favorite model by far. And this I is his favorite model by far. And this I is his favorite model by far. And this I absolutely agree with. It's the thing absolutely agree with. It's the thing absolutely agree with. It's the thing that makes 56 feel special. It'll just that makes 56 feel special. It'll just that makes 56 feel special. It'll just keep going. It will try so goddamn hard keep going. It will try so goddamn hard keep going. It will try so goddamn hard to solve the problem more than any model to solve the problem more than any model to solve the problem more than any model I've ever used. And that's so different I've ever used. And that's so different I've ever used. And that's so different from how 55 behaved. It's almost like from how 55 behaved. It's almost like from how 55 behaved. It's almost like jarring in a way. 55 was really quick to jarring in a way. 55 was really quick to jarring in a way. 55 was really quick to stop and be like, "Okay, I finished one stop and be like, "Okay, I finished one stop and be like, "Okay, I finished one and a half of your sevenstep plan. Can I and a half of your sevenstep plan. Can I and a half of your sevenstep plan. Can I keep going over and over?" Or it would keep going over and over?" Or it would keep going over and over?" Or it would have something bad in context and get have something bad in context and get have something bad in context and get lost and die. I knew 56 would fix a lot lost and die. I knew 56 would fix a lot lost and die. I knew 56 would fix a lot of the problems I have with 5'5. I did of the problems I have with 5'5. I did of the problems I have with 5'5. I did not think a post-training run and an RL not think a post-training run and an RL not think a post-training run and an RL pass would be able to make the model go pass would be able to make the model go pass would be able to make the model go from, yeah, that's smart, but a little from, yeah, that's smart, but a little from, yeah, that's smart, but a little annoying to, "Holy [ __ ] this is
-
annoying to, "Holy [ __ ] this is annoying to, "Holy [ __ ] this is unbelievably capable." Because remember, unbelievably capable." Because remember, unbelievably capable." Because remember, that's all 56 is. It's not a new base that's all 56 is. It's not a new base that's all 56 is. It's not a new base model. It's not a new pre-training. It's model. It's not a new pre-training. It's model. It's not a new pre-training. It's not more parameters than it was before. not more parameters than it was before. not more parameters than it was before. It is a refinement on 5.5, which is why It is a refinement on 5.5, which is why It is a refinement on 5.5, which is why its capability is so insanely impressive its capability is so insanely impressive its capability is so insanely impressive and gets me even more excited for the and gets me even more excited for the and gets me even more excited for the next pre-training run they do. And I'm next pre-training run they do. And I'm next pre-training run they do. And I'm almost positive they're working on GPT6 almost positive they're working on GPT6 almost positive they're working on GPT6 right now. I did hear rumors that GPT6 right now. I did hear rumors that GPT6 right now. I did hear rumors that GPT6 would come out end of month. I'm going would come out end of month. I'm going would come out end of month. I'm going to call [ __ ] on that now. There is to call [ __ ] on that now. There is to call [ __ ] on that now. There is no way they're going to be able to get no way they're going to be able to get no way they're going to be able to get that done in time. Just realistically that done in time. Just realistically that done in time. Just realistically speaking. More reviews. Mitchell, the speaking. More reviews. Mitchell, the speaking. More reviews. Mitchell, the creator of Terraform and Ghosty, creator of Terraform and Ghosty, creator of Terraform and Ghosty, absolute legend, has had early access as absolute legend, has had early access as absolute legend, has had early access as well. Souls as default. It's faster. well. Souls as default. It's faster. well. Souls as default. It's faster. plans and judges just as good as Fable plans and judges just as good as Fable plans and judges just as good as Fable and he thinks it produces better overall and he thinks it produces better overall and he thinks it produces better overall work. He calls out a few things Fable is work. He calls out a few things Fable is work. He calls out a few things Fable is good at which I'll save for the future good at which I'll save for the future good at which I'll save for the future because again dedicated Fable versus because again dedicated Fable versus because again dedicated Fable versus video coming soon. Tim from the Next.js video coming soon. Tim from the Next.js video coming soon. Tim from the Next.js team said that he's been testing 56 Soul team said that he's been testing 56 Soul team said that he's been testing 56 Soul for over two months. It's incredibly for over two months. It's incredibly for over two months. It's incredibly good in his day-to-day work on Next. It good in his day-to-day work on Next. It good in his day-to-day work on Next. It understands architecture trade-offs. It understands architecture trade-offs. It understands architecture trade-offs. It can investigate complicated Next issues. can investigate complicated Next issues. can investigate complicated Next issues. It considers other areas of the codebase It considers other areas of the codebase It considers other areas of the codebase when fixing bugs. It needs very little when fixing bugs. It needs very little when fixing bugs. It needs very little guidance and short prompts are enough.
-
guidance and short prompts are enough. guidance and short prompts are enough. There's some big refactors of the next There's some big refactors of the next There's some big refactors of the next server that it implemented end to end server that it implemented end to end server that it implemented end to end with him pointing at high level possible with him pointing at high level possible with him pointing at high level possible improvements. The peers are ready to improvements. The peers are ready to improvements. The peers are ready to merge after next 163 has been released. merge after next 163 has been released. merge after next 163 has been released. Pretty nuts. And then Cory had the Pretty nuts. And then Cory had the Pretty nuts. And then Cory had the following to say, "I'm really confused following to say, "I'm really confused following to say, "I'm really confused by OpenAI strategy for 5.6. I can't find by OpenAI strategy for 5.6. I can't find by OpenAI strategy for 5.6. I can't find the date the price jumps by 50%, the the date the price jumps by 50%, the the date the price jumps by 50%, the date that it leaves the inclusion in date that it leaves the inclusion in date that it leaves the inclusion in subscriptions where it's probably subscriptions where it's probably subscriptions where it's probably missing. And what their employees say is missing. And what their employees say is missing. And what their employees say is somehow aligning with what the docs say, somehow aligning with what the docs say, somehow aligning with what the docs say, too. Clearly something's up. For those too. Clearly something's up. For those too. Clearly something's up. For those who suck at sarcasm, this is a hilarious who suck at sarcasm, this is a hilarious who suck at sarcasm, this is a hilarious burn on Anthropic for doing all of those burn on Anthropic for doing all of those burn on Anthropic for doing all of those things incorrectly constantly. An things incorrectly constantly. An things incorrectly constantly. An opening eye employee replied, "Sorry to opening eye employee replied, "Sorry to opening eye employee replied, "Sorry to disappoint, but you can use 100% of your disappoint, but you can use 100% of your disappoint, but you can use 100% of your quota on the plan that you're paying for quota on the plan that you're paying for quota on the plan that you're paying for on 56 forever. Additionally, the price on 56 forever. Additionally, the price on 56 forever. Additionally, the price is going to stay the same." To which is going to stay the same." To which is going to stay the same." To which Cory replied that they need to hire a VP Cory replied that they need to hire a VP Cory replied that they need to hire a VP of rug pulling. One more review that of rug pulling. One more review that of rug pulling. One more review that I'll summarize pretty quick from the I'll summarize pretty quick from the I'll summarize pretty quick from the guys over at Every. They lost access and guys over at Every. They lost access and guys over at Every. They lost access and felt like they were going insane, like felt like they were going insane, like felt like they were going insane, like they were trying to shoot a basketball they were trying to shoot a basketball they were trying to shoot a basketball that's twice as heavy. Soul is their that's twice as heavy. Soul is their that's twice as heavy. Soul is their favorite model to work with. It's really favorite model to work with. It's really favorite model to work with. It's really fast, which changes how you use it. It fast, which changes how you use it. It fast, which changes how you use it. It finds the context it needs. This is a finds the context it needs. This is a finds the context it needs. This is a huge difference with 55, which was so huge difference with 55, which was so huge difference with 55, which was so much worse at that in particular. One much worse at that in particular. One much worse at that in particular. One thread can carry a project in thread can carry a project in thread can carry a project in production. Yes, again, it doesn't lose production. Yes, again, it doesn't lose production. Yes, again, it doesn't lose track of what it's doing in a thread. I track of what it's doing in a thread. I track of what it's doing in a thread. I probably have way fewer threads than I probably have way fewer threads than I probably have way fewer threads than I used to with 56, but I'm doing way more used to with 56, but I'm doing way more used to with 56, but I'm doing way more work cuz the compaction [ __ ] works work cuz the compaction [ __ ] works work cuz the compaction [ __ ] works again. It's great. It plans well, but it again. It's great. It plans well, but it again. It's great. It plans well, but it may build too much. We'll talk about may build too much. We'll talk about may build too much. We'll talk about this momentarily. And it works best when this momentarily. And it works best when this momentarily. And it works best when you plan to steer. Soul gets better when you plan to steer. Soul gets better when you plan to steer. Soul gets better when the surrounding system supplies sources.
-
the surrounding system supplies sources. the surrounding system supplies sources. examples, style guides, and clear examples, style guides, and clear examples, style guides, and clear outcomes. Review its choices and outcomes. Review its choices and outcomes. Review its choices and redirect it as the work changes. I'm redirect it as the work changes. I'm redirect it as the work changes. I'm going to talk a bit about this though. going to talk a bit about this though. going to talk a bit about this though. Obviously, focusing on soul for now. Obviously, focusing on soul for now. Obviously, focusing on soul for now. I'll talk about the others in a bit. I'll talk about the others in a bit. I'll talk about the others in a bit. Don't worry. Hopefully, we'll establish Don't worry. Hopefully, we'll establish Don't worry. Hopefully, we'll establish at this point. GBD56 Soul is a at this point. GBD56 Soul is a at this point. GBD56 Soul is a phenomenal model, and if you use AI for phenomenal model, and if you use AI for phenomenal model, and if you use AI for writing code, it absolutely has a place writing code, it absolutely has a place writing code, it absolutely has a place in your toolbox. That doesn't mean we've in your toolbox. That doesn't mean we've in your toolbox. That doesn't mean we've answered the question, though. The answered the question, though. The answered the question, though. The question, of course, being should this question, of course, being should this question, of course, being should this be the model that consumes all of your be the model that consumes all of your be the model that consumes all of your usage? Should this be the model that usage? Should this be the model that usage? Should this be the model that runs all of your other things? Should runs all of your other things? Should runs all of your other things? Should this model command Fable or should Fable this model command Fable or should Fable this model command Fable or should Fable command soul? That is again a dedicated command soul? That is again a dedicated command soul? That is again a dedicated video coming up. But I want to set the video coming up. But I want to set the video coming up. But I want to set the groundwork for it by talking about its groundwork for it by talking about its groundwork for it by talking about its strengths and weaknesses. First strength strengths and weaknesses. First strength strengths and weaknesses. First strength is that it's determined as hell. If you is that it's determined as hell. If you is that it's determined as hell. If you give this model a task and it possibly give this model a task and it possibly give this model a task and it possibly can complete it, it'll find a way to. can complete it, it'll find a way to. can complete it, it'll find a way to. Whether or not you want it to, it will Whether or not you want it to, it will Whether or not you want it to, it will find a way to. It is better at front find a way to. It is better at front find a way to. It is better at front end. It's still not better than other end. It's still not better than other end. It's still not better than other labs, but it's better than 55. It's labs, but it's better than 55. It's labs, but it's better than 55. It's industryleading at computer use, which industryleading at computer use, which industryleading at computer use, which honestly is one of my favorite things honestly is one of my favorite things honestly is one of my favorite things that I never thought I would love. I that I never thought I would love. I that I never thought I would love. I just let it control my computer and it just let it control my computer and it just let it control my computer and it it blows me away what it can do. I it blows me away what it can do. I it blows me away what it can do. I forgot how much worse 55 was at it until forgot how much worse 55 was at it until forgot how much worse 55 was at it until it was taken away and I just kind of it was taken away and I just kind of it was taken away and I just kind of stopped using my computer because I was stopped using my computer because I was stopped using my computer because I was so much less happy with my experience.
-
so much less happy with my experience. so much less happy with my experience. It was kind of jarring to have it taken It was kind of jarring to have it taken It was kind of jarring to have it taken away and then get it back cuz that away and then get it back cuz that away and then get it back cuz that window without it was rough. It's also window without it was rough. It's also window without it was rough. It's also really efficient, even compared to other really efficient, even compared to other really efficient, even compared to other models in a similar capability tier. It models in a similar capability tier. It models in a similar capability tier. It uses way fewer tokens and comes out way uses way fewer tokens and comes out way uses way fewer tokens and comes out way cheaper at a better base price as well, cheaper at a better base price as well, cheaper at a better base price as well, which also means it's really fast, even which also means it's really fast, even which also means it's really fast, even without the Cerebra stuff. If you're not without the Cerebra stuff. If you're not without the Cerebra stuff. If you're not familiar, OpenAI promised us that there familiar, OpenAI promised us that there familiar, OpenAI promised us that there would be a version of Soul on the would be a version of Soul on the would be a version of Soul on the Cerebras hosting that would be up to 750 Cerebras hosting that would be up to 750 Cerebras hosting that would be up to 750 tokens per second instead of the usual tokens per second instead of the usual tokens per second instead of the usual 40 to 60. And I'm very excited to try 40 to 60. And I'm very excited to try 40 to 60. And I'm very excited to try that cuz it already feels so fast. The that cuz it already feels so fast. The that cuz it already feels so fast. The fast mode isn't that though. The fast fast mode isn't that though. The fast fast mode isn't that though. The fast mode is just more provisioning on the mode is just more provisioning on the mode is just more provisioning on the normal Nvidia based inference. So if normal Nvidia based inference. So if normal Nvidia based inference. So if you're using fast mode and you're like you're using fast mode and you're like you're using fast mode and you're like that's not as fast as I expected, it's that's not as fast as I expected, it's that's not as fast as I expected, it's cuz that's not the Cerebrus hosting cuz that's not the Cerebrus hosting cuz that's not the Cerebrus hosting that's apparently coming soon. Also on that's apparently coming soon. Also on that's apparently coming soon. Also on the topic of coding, it's way better at the topic of coding, it's way better at the topic of coding, it's way better at mobile dev stuff. It's also really good mobile dev stuff. It's also really good mobile dev stuff. It's also really good at like navigating around a problem at like navigating around a problem at like navigating around a problem space. So things like environment setup, space. So things like environment setup, space. So things like environment setup, using other machines, controlling things using other machines, controlling things using other machines, controlling things over SSH, provisioning, orchestrating, over SSH, provisioning, orchestrating, over SSH, provisioning, orchestrating, that type of stuff it's great at. And on that type of stuff it's great at. And on that type of stuff it's great at. And on that note, his capability of that note, his capability of that note, his capability of orchestration things with sub agents in orchestration things with sub agents in orchestration things with sub agents in particular is nextgen. The only models particular is nextgen. The only models particular is nextgen. The only models that have this taste right now, this that have this taste right now, this that have this taste right now, this capability of understanding how to break capability of understanding how to break capability of understanding how to break up their work for sub agents are 56 up their work for sub agents are 56 up their work for sub agents are 56 soul, fable 5/OS, soul, fable 5/OS, soul, fable 5/OS, and sonnet 5. Potentially Terra can do and sonnet 5. Potentially Terra can do and sonnet 5. Potentially Terra can do this, but I haven't had a chance to try this, but I haven't had a chance to try this, but I haven't had a chance to try yet. But this new generation of yet. But this new generation of yet. But this new generation of orchestration tasks are a huge strength orchestration tasks are a huge strength orchestration tasks are a huge strength that only a couple models have. And now
-
that only a couple models have. And now that only a couple models have. And now OpenAI ones do too. And that's enough OpenAI ones do too. And that's enough OpenAI ones do too. And that's enough for this to be your default model for a for this to be your default model for a for this to be your default model for a lot of reasons. Other things really lot of reasons. Other things really lot of reasons. Other things really quick that I could think of, uh, way quick that I could think of, uh, way quick that I could think of, uh, way better at compaction. This is an better at compaction. This is an better at compaction. This is an important one because in codeex you important one because in codeex you important one because in codeex you still don't get to use the full million still don't get to use the full million still don't get to use the full million token context window. It still limits to token context window. It still limits to token context window. It still limits to around 200k. So for it to do these long around 200k. So for it to do these long around 200k. So for it to do these long running things, it needs to be able to running things, it needs to be able to running things, it needs to be able to manage the context. Now that it could manage the context. Now that it could manage the context. Now that it could compact and not lose track of what it's compact and not lose track of what it's compact and not lose track of what it's doing, it is way better at those types doing, it is way better at those types doing, it is way better at those types of long runs where before I would of long runs where before I would of long runs where before I would basically never let 55 run for a while basically never let 55 run for a while basically never let 55 run for a while because it would just get lost and do because it would just get lost and do because it would just get lost and do something stupid. On that note, it's something stupid. On that note, it's something stupid. On that note, it's much better about context pollution than much better about context pollution than much better about context pollution than it used to be. This is a huge problem I it used to be. This is a huge problem I it used to be. This is a huge problem I had with 55 where if it read the wrong had with 55 where if it read the wrong had with 55 where if it read the wrong file and put something in its history, file and put something in its history, file and put something in its history, it would just lose track of its goals it would just lose track of its goals it would just lose track of its goals and be miserable to work with. That's and be miserable to work with. That's and be miserable to work with. That's resolved now. And it's also better at resolved now. And it's also better at resolved now. And it's also better at understanding intent now. when you tell understanding intent now. when you tell understanding intent now. when you tell it to do something, it doesn't make a it to do something, it doesn't make a it to do something, it doesn't make a bad assumption anywhere near as bad assumption anywhere near as bad assumption anywhere near as aggressively as it used to. Apparently, aggressively as it used to. Apparently, aggressively as it used to. Apparently, the context it has by default is 350K, the context it has by default is 350K, the context it has by default is 350K, which is better. Cool. It did not used which is better. Cool. It did not used which is better. Cool. It did not used to be that high. So, now that we have to be that high. So, now that we have to be that high. So, now that we have the strengths, I think it's time to the strengths, I think it's time to the strengths, I think it's time to write the weaknesses.
-
write the weaknesses. write the weaknesses. First and foremost, by default, it First and foremost, by default, it First and foremost, by default, it writes way too much code. You got to writes way too much code. You got to writes way too much code. You got to adjust your system prompts and your adjust your system prompts and your adjust your system prompts and your skills and a bunch of those things to skills and a bunch of those things to skills and a bunch of those things to convince it to not do that. To tone down convince it to not do that. To tone down convince it to not do that. To tone down its overeagerness to just write its overeagerness to just write its overeagerness to just write everything. This model will turn a everything. This model will turn a everything. This model will turn a fiveline change into a 300line file fiveline change into a 300line file fiveline change into a 300line file rewrite in 2,000 lines of tests. It rewrite in 2,000 lines of tests. It rewrite in 2,000 lines of tests. It loves writing unnecessary tests. Like loves writing unnecessary tests. Like loves writing unnecessary tests. Like far too many of them. I can't tell you far too many of them. I can't tell you far too many of them. I can't tell you how many times I've had to have other how many times I've had to have other how many times I've had to have other models come in and clean up because 56 models come in and clean up because 56 models come in and clean up because 56 did too much. it was too convinced those did too much. it was too convinced those did too much. it was too convinced those things were necessary. I also hinted at things were necessary. I also hinted at things were necessary. I also hinted at this before, but it's a little too this before, but it's a little too this before, but it's a little too determined to the point where it will determined to the point where it will determined to the point where it will break things if they're in the way. This break things if they're in the way. This break things if they're in the way. This model loves to work around the problems model loves to work around the problems model loves to work around the problems that it runs into sometimes in ways that that it runs into sometimes in ways that that it runs into sometimes in ways that are a little too clever, like when it are a little too clever, like when it are a little too clever, like when it can't launch some things that doesn't can't launch some things that doesn't can't launch some things that doesn't have pseudo permissions. So, it finds have pseudo permissions. So, it finds have pseudo permissions. So, it finds something else that does and then something else that does and then something else that does and then convinces it to launch it. Weird sketchy convinces it to launch it. Weird sketchy convinces it to launch it. Weird sketchy [ __ ] I've seen this model do. It means [ __ ] I've seen this model do. It means [ __ ] I've seen this model do. It means that it doesn't get blocked, but it also that it doesn't get blocked, but it also that it doesn't get blocked, but it also means I kind of want to run it in a VM means I kind of want to run it in a VM means I kind of want to run it in a VM more often than not. As I mentioned more often than not. As I mentioned more often than not. As I mentioned before, it's far from frontier at before, it's far from frontier at before, it's far from frontier at design. It's better, but it doesn't design. It's better, but it doesn't design. It's better, but it doesn't bring exciting designs to me very often.
-
bring exciting designs to me very often. bring exciting designs to me very often. And on that note, it's also pretty bad And on that note, it's also pretty bad And on that note, it's also pretty bad at understanding its own limitations. It at understanding its own limitations. It at understanding its own limitations. It is eager to use its tools to confirm is eager to use its tools to confirm is eager to use its tools to confirm information it's not sure about, but if information it's not sure about, but if information it's not sure about, but if it thinks something is the case that it thinks something is the case that it thinks something is the case that isn't, it will fight you tooth and nail isn't, it will fight you tooth and nail isn't, it will fight you tooth and nail on that. It does not like being wrong on that. It does not like being wrong on that. It does not like being wrong and it does not know when it is wrong. and it does not know when it is wrong. and it does not know when it is wrong. It was only come up a few times for me, It was only come up a few times for me, It was only come up a few times for me, but it was really annoying when it did but it was really annoying when it did but it was really annoying when it did and it often sent it down crazy rabbit and it often sent it down crazy rabbit and it often sent it down crazy rabbit holes where it would try to to come up holes where it would try to to come up holes where it would try to to come up with some novel solution to a problem with some novel solution to a problem with some novel solution to a problem that didn't actually exist. Back to the that didn't actually exist. Back to the that didn't actually exist. Back to the determined part though, cuz there's determined part though, cuz there's determined part though, cuz there's other important pieces here. It can burn other important pieces here. It can burn other important pieces here. It can burn tokens aggressively if it doesn't have a tokens aggressively if it doesn't have a tokens aggressively if it doesn't have a clear stopping point. If it thinks it clear stopping point. If it thinks it clear stopping point. If it thinks it can eventually get somewhere, it will go can eventually get somewhere, it will go can eventually get somewhere, it will go and go and go, even without having like and go and go, even without having like and go and go, even without having like a goal that it's running against, a goal that it's running against, a goal that it's running against, there's a lot of prompts where 55 would there's a lot of prompts where 55 would there's a lot of prompts where 55 would have given up and stopped, and 56 will have given up and stopped, and 56 will have given up and stopped, and 56 will just keep trying. That could result in just keep trying. That could result in just keep trying. That could result in massive token burn, especially if you massive token burn, especially if you massive token burn, especially if you combine that with fast motor ultra. This combine that with fast motor ultra. This combine that with fast motor ultra. This model is the first time I have gotten model is the first time I have gotten model is the first time I have gotten into my limits on codeex ever. It's also into my limits on codeex ever. It's also into my limits on codeex ever. It's also confusing how many options we have with confusing how many options we have with confusing how many options we have with it. Not just Soul, but when you combine it. Not just Soul, but when you combine it. Not just Soul, but when you combine the different reasoning efforts on Soul the different reasoning efforts on Soul the different reasoning efforts on Soul and Soul Pro alongside Terra and Luna, and Soul Pro alongside Terra and Luna, and Soul Pro alongside Terra and Luna, there's a lot to decide between, which there's a lot to decide between, which there's a lot to decide between, which can be rough. There's one more thing I can be rough. There's one more thing I can be rough. There's one more thing I want to say about the weaknesses that's want to say about the weaknesses that's want to say about the weaknesses that's hard to put into words. The simplest I hard to put into words. The simplest I hard to put into words. The simplest I can put it is that the model's capable, can put it is that the model's capable, can put it is that the model's capable, but not necessarily thoughtful. It does but not necessarily thoughtful. It does but not necessarily thoughtful. It does whatever it has to, but it doesn't whatever it has to, but it doesn't whatever it has to, but it doesn't always think about what it's doing. and always think about what it's doing. and always think about what it's doing. and the results are that it just writes more the results are that it just writes more the results are that it just writes more and more and more code or tries again and more and more code or tries again and more and more code or tries again and again and again to work around
-
and again and again to work around and again and again to work around something instead of taking the step something instead of taking the step something instead of taking the step back and rethinking it. So on one hand I back and rethinking it. So on one hand I back and rethinking it. So on one hand I trust this model more than ever to go trust this model more than ever to go trust this model more than ever to go off and do [ __ ] but on the other hand, off and do [ __ ] but on the other hand, off and do [ __ ] but on the other hand, I also check in a lot to make sure it's I also check in a lot to make sure it's I also check in a lot to make sure it's not getting lost in the sauce. I think not getting lost in the sauce. I think not getting lost in the sauce. I think that's most of what I have to say here. that's most of what I have to say here. that's most of what I have to say here. Is it the best model in the world? You Is it the best model in the world? You Is it the best model in the world? You have to wait for my video comparing Soul have to wait for my video comparing Soul have to wait for my video comparing Soul and Fable for that. But I do want to and Fable for that. But I do want to and Fable for that. But I do want to talk a little more about the options you talk a little more about the options you talk a little more about the options you have. Ultra is going to have to wait for have. Ultra is going to have to wait for have. Ultra is going to have to wait for later because there's a lot more to say later because there's a lot more to say later because there's a lot more to say there. But when we're talking about there. But when we're talking about there. But when we're talking about Soul, Terra, and Luna, as well as all of Soul, Terra, and Luna, as well as all of Soul, Terra, and Luna, as well as all of their different reasoning options, I their different reasoning options, I their different reasoning options, I understand why you might be getting understand why you might be getting understand why you might be getting confused. For code tasks, I'm going to confused. For code tasks, I'm going to confused. For code tasks, I'm going to try my best to resolve this with Deepsw try my best to resolve this with Deepsw try my best to resolve this with Deepsw S. This benchmark is great, and there's S. This benchmark is great, and there's S. This benchmark is great, and there's some really useful info to get here, some really useful info to get here, some really useful info to get here, even if it's a little clogged up, even if it's a little clogged up, even if it's a little clogged up, despite the fact that I just put a lot despite the fact that I just put a lot despite the fact that I just put a lot of time into hiding all the things that of time into hiding all the things that of time into hiding all the things that aren't really relevant anymore. I know a aren't really relevant anymore. I know a aren't really relevant anymore. I know a lot of people really liked 55 medium and lot of people really liked 55 medium and lot of people really liked 55 medium and are confused about what they should go are confused about what they should go are confused about what they should go to now. Well, I have good news because to now. Well, I have good news because to now. Well, I have good news because almost every version of Terra ends up almost every version of Terra ends up almost every version of Terra ends up cheaper than 55 medium was. And even 56 cheaper than 55 medium was. And even 56 cheaper than 55 medium was. And even 56 on medium is cheaper than 55 was. So 56 on medium is cheaper than 55 was. So 56 on medium is cheaper than 55 was. So 56 medium is probably a good enough medium is probably a good enough medium is probably a good enough replacement for what you're used to with replacement for what you're used to with replacement for what you're used to with 55 medium. But how do you pick between 55 medium. But how do you pick between 55 medium. But how do you pick between Terra on XHigh and Soul on medium? I Terra on XHigh and Soul on medium? I Terra on XHigh and Soul on medium? I don't [ __ ] know. I only had access to don't [ __ ] know. I only had access to don't [ __ ] know. I only had access to soul when I was testing. That said, it soul when I was testing. That said, it soul when I was testing. That said, it seems like the spacing across options seems like the spacing across options seems like the spacing across options with Terra is really good. Specifically, with Terra is really good. Specifically, with Terra is really good. Specifically, the jump from X high to max is actually the jump from X high to max is actually the jump from X high to max is actually a meaningful score bump. Whereas with
-
a meaningful score bump. Whereas with a meaningful score bump. Whereas with Soul, the bump from Xai to max wasn't Soul, the bump from Xai to max wasn't Soul, the bump from Xai to max wasn't that big. I'll also throw in Luna here, that big. I'll also throw in Luna here, that big. I'll also throw in Luna here, and you can see how much messier things and you can see how much messier things and you can see how much messier things get when it's there. A lot of its scores get when it's there. A lot of its scores get when it's there. A lot of its scores are low, but it also just kind of gets are low, but it also just kind of gets are low, but it also just kind of gets in the way a lot. Somehow Luna on max in the way a lot. Somehow Luna on max in the way a lot. Somehow Luna on max scored really really well, but that scored really really well, but that scored really really well, but that doesn't mean I think you should use it. doesn't mean I think you should use it. doesn't mean I think you should use it. So here's how I would think about these So here's how I would think about these So here's how I would think about these options. I will start with Luna. Luna is options. I will start with Luna. Luna is options. I will start with Luna. Luna is not for you as a dev. It's really cheap, not for you as a dev. It's really cheap, not for you as a dev. It's really cheap, really fast, really capable, but the really fast, really capable, but the really fast, really capable, but the point of Luna is to be one of the things point of Luna is to be one of the things point of Luna is to be one of the things that is orchestrated by a smarter agent that is orchestrated by a smarter agent that is orchestrated by a smarter agent or for it to do things like bulk data or for it to do things like bulk data or for it to do things like bulk data processing, title generation, stuff like processing, title generation, stuff like processing, title generation, stuff like that. We'll probably start using Luna that. We'll probably start using Luna that. We'll probably start using Luna for doing like branch naming and title for doing like branch naming and title for doing like branch naming and title gen inside of T3 code by default. So you gen inside of T3 code by default. So you gen inside of T3 code by default. So you won't have to go click any buttons for won't have to go click any buttons for won't have to go click any buttons for it. It'll just work. And I'm very it. It'll just work. And I'm very it. It'll just work. And I'm very excited that it is at the point it is at excited that it is at the point it is at excited that it is at the point it is at and that it is so damn good. But it is and that it is so damn good. But it is and that it is so damn good. But it is not there for you to click on it in a not there for you to click on it in a not there for you to click on it in a drop down. The point of Luna is to be drop down. The point of Luna is to be drop down. The point of Luna is to be surprisingly capable at that really surprisingly capable at that really surprisingly capable at that really really cheap price tier. You should let really cheap price tier. You should let really cheap price tier. You should let the models be the ones to call Luna, not the models be the ones to call Luna, not the models be the ones to call Luna, not you. And maybe if you're doing like your you. And maybe if you're doing like your you. And maybe if you're doing like your own programmatic stuff like you're an own programmatic stuff like you're an own programmatic stuff like you're an analyzing data or chat reads or stuff analyzing data or chat reads or stuff analyzing data or chat reads or stuff like that, you could use Luna in code, like that, you could use Luna in code, like that, you could use Luna in code, but I wouldn't select it in a drop down but I wouldn't select it in a drop down but I wouldn't select it in a drop down other than experimentation. So now we other than experimentation. So now we other than experimentation. So now we need to talk about Terra and Soul need to talk about Terra and Soul need to talk about Terra and Soul because these are the pieces that are because these are the pieces that are because these are the pieces that are much more interesting. Soul is the much more interesting. Soul is the much more interesting. Soul is the smartest model ever depending on who you smartest model ever depending on who you smartest model ever depending on who you ask. It is surprisingly efficient for ask. It is surprisingly efficient for ask. It is surprisingly efficient for the price and you should definitely the price and you should definitely the price and you should definitely throw it at work when you're not sure if
-
throw it at work when you're not sure if throw it at work when you're not sure if a model is capable of solving it. If you a model is capable of solving it. If you a model is capable of solving it. If you are planning on setting off a task and are planning on setting off a task and are planning on setting off a task and expected to take more than 10 minutes, expected to take more than 10 minutes, expected to take more than 10 minutes, Soul's almost certainly the right Soul's almost certainly the right Soul's almost certainly the right choice. But what about Terra? Again, choice. But what about Terra? Again, choice. But what about Terra? Again, I've only tested Terara for about 20 or I've only tested Terara for about 20 or I've only tested Terara for about 20 or so prompts over the last 2 days cuz I so prompts over the last 2 days cuz I so prompts over the last 2 days cuz I did not have access during the previous did not have access during the previous did not have access during the previous early access window. So, what do I early access window. So, what do I early access window. So, what do I recommend Terara for? First and recommend Terara for? First and recommend Terara for? First and foremost, coding on a budget. It is foremost, coding on a budget. It is foremost, coding on a budget. It is incredibly capable for its price. So if incredibly capable for its price. So if incredibly capable for its price. So if you're on the $100 tier or the $20 tier you're on the $100 tier or the $20 tier you're on the $100 tier or the $20 tier for codecs, Terra is probably a much for codecs, Terra is probably a much for codecs, Terra is probably a much better default. If you find yourself at better default. If you find yourself at better default. If you find yourself at the end of a usage window with a bunch the end of a usage window with a bunch the end of a usage window with a bunch of remaining usage, switching up the of remaining usage, switching up the of remaining usage, switching up the Soul makes sense even on the cheaper Soul makes sense even on the cheaper Soul makes sense even on the cheaper tiers. I wrote a bit more about when I tiers. I wrote a bit more about when I tiers. I wrote a bit more about when I think Terra makes sense. It's great for think Terra makes sense. It's great for think Terra makes sense. It's great for reviewing work cuz it'll just read reviewing work cuz it'll just read reviewing work cuz it'll just read through whatever and give you good through whatever and give you good through whatever and give you good feedback on it. It doesn't have quite feedback on it. It doesn't have quite feedback on it. It doesn't have quite the same level of determination that the same level of determination that the same level of determination that Soul does. it will stop more Soul does. it will stop more Soul does. it will stop more aggressively. But that's also kind of a aggressively. But that's also kind of a aggressively. But that's also kind of a good thing if you want to be sitting good thing if you want to be sitting good thing if you want to be sitting there going back and forth with the there going back and forth with the there going back and forth with the model. It's also a pretty solid model. It's also a pretty solid model. It's also a pretty solid workhorse for implementation. If you workhorse for implementation. If you workhorse for implementation. If you have a model like Soul do all of the have a model like Soul do all of the have a model like Soul do all of the deep diving, figuring out what needs to deep diving, figuring out what needs to deep diving, figuring out what needs to be done, having Terra come in and write be done, having Terra come in and write be done, having Terra come in and write the code is not a bad idea. Especially the code is not a bad idea. Especially the code is not a bad idea. Especially because from my limited experience, it because from my limited experience, it because from my limited experience, it seems a little less aggressive about seems a little less aggressive about seems a little less aggressive about overwriting. Doesn't do the too much overwriting. Doesn't do the too much overwriting. Doesn't do the too much code thing that I've experienced with code thing that I've experienced with code thing that I've experienced with Soul. And as I was sending at there, Soul. And as I was sending at there, Soul. And as I was sending at there, having a human in the loop where you're having a human in the loop where you're having a human in the loop where you're sending a prompt, waiting for it to make sending a prompt, waiting for it to make sending a prompt, waiting for it to make a couple changes, and then looking at a couple changes, and then looking at a couple changes, and then looking at it, it's pretty damn good for that. All it, it's pretty damn good for that. All it, it's pretty damn good for that. All that said, if you've been dealing with that said, if you've been dealing with that said, if you've been dealing with Fable at all at its absurdly high price Fable at all at its absurdly high price Fable at all at its absurdly high price and usage limits, Soul's probably going
-
and usage limits, Soul's probably going and usage limits, Soul's probably going to be a good default. And my honest to be a good default. And my honest to be a good default. And my honest recommendation is that you should start recommendation is that you should start recommendation is that you should start with Soul. You should push it to its with Soul. You should push it to its with Soul. You should push it to its limits. Push your usage to its limits, limits. Push your usage to its limits, limits. Push your usage to its limits, too. And once you start getting close to too. And once you start getting close to too. And once you start getting close to hitting those limits and running out of hitting those limits and running out of hitting those limits and running out of usage, bump some of your work down to usage, bump some of your work down to usage, bump some of your work down to Terra and see how it goes. I do want to Terra and see how it goes. I do want to Terra and see how it goes. I do want to talk a little bit about reasoning talk a little bit about reasoning talk a little bit about reasoning efforts. Again, we're not talking about efforts. Again, we're not talking about efforts. Again, we're not talking about ultra yet. My honest advice there is ultra yet. My honest advice there is ultra yet. My honest advice there is don't use it unless you know what you're don't use it unless you know what you're don't use it unless you know what you're doing. Ultra is not just a new reasoning doing. Ultra is not just a new reasoning doing. Ultra is not just a new reasoning level, so be very careful. It will burn level, so be very careful. It will burn level, so be very careful. It will burn your usage. Same with fast mode because your usage. Same with fast mode because your usage. Same with fast mode because this model will go so much longer if you this model will go so much longer if you this model will go so much longer if you get soul in one of those impossible task get soul in one of those impossible task get soul in one of those impossible task loops where it keeps on going. On fast loops where it keeps on going. On fast loops where it keeps on going. On fast mode, you will burn through your weekly mode, you will burn through your weekly mode, you will burn through your weekly usage in a few hours. It's rough. And usage in a few hours. It's rough. And usage in a few hours. It's rough. And again, when we go to benches like deepsw again, when we go to benches like deepsw again, when we go to benches like deepsw SWE, you can see that from high to max SWE, you can see that from high to max SWE, you can see that from high to max is not a big difference in capability, is not a big difference in capability, is not a big difference in capability, but is a massive difference in cost but is a massive difference in cost but is a massive difference in cost where high scored a 69 and max scored a where high scored a 69 and max scored a where high scored a 69 and max scored a 73 with X high only at a 71, like smack 73 with X high only at a 71, like smack 73 with X high only at a 71, like smack dab between the two. But you went from dab between the two. But you went from dab between the two. But you went from $3.40 per task to $4.70 to $8.39.
-
$3.40 per task to $4.70 to $8.39. $3.40 per task to $4.70 to $8.39. It also ends up much slower when you do It also ends up much slower when you do It also ends up much slower when you do that because of the number of output that because of the number of output that because of the number of output tokens it is generating and the tokens it is generating and the tokens it is generating and the additional steps that the agent is additional steps that the agent is additional steps that the agent is taking. So for all of those reasons, I taking. So for all of those reasons, I taking. So for all of those reasons, I think medium and high on soul are think medium and high on soul are think medium and high on soul are really, really good defaults. My current really, really good defaults. My current really, really good defaults. My current default is soul on high and I do expect default is soul on high and I do expect default is soul on high and I do expect to stay on that for a bit. I've dropped to stay on that for a bit. I've dropped to stay on that for a bit. I've dropped down to medium and noticed it missed a down to medium and noticed it missed a down to medium and noticed it missed a few things that it wouldn't have on high few things that it wouldn't have on high few things that it wouldn't have on high and I've went up to X high and noticed and I've went up to X high and noticed and I've went up to X high and noticed that tasks took a lot longer without that tasks took a lot longer without that tasks took a lot longer without getting meaningfully better results. So getting meaningfully better results. So getting meaningfully better results. So to really briefly summarize my to really briefly summarize my to really briefly summarize my recommendations, Soul on High for most recommendations, Soul on High for most recommendations, Soul on High for most things, especially once you get into the things, especially once you get into the things, especially once you get into the habit of orchestrating your work and habit of orchestrating your work and habit of orchestrating your work and doing lots of sub aents, Terra on Medium doing lots of sub aents, Terra on Medium doing lots of sub aents, Terra on Medium is a budget king. It's a very, very good is a budget king. It's a very, very good is a budget king. It's a very, very good option for realworld work while also option for realworld work while also option for realworld work while also being super fast and cheap. And Luna is being super fast and cheap. And Luna is being super fast and cheap. And Luna is a thing your agent should use a hell of a thing your agent should use a hell of a thing your agent should use a hell of a lot more than you do for when they're a lot more than you do for when they're a lot more than you do for when they're looking for specific things, they're looking for specific things, they're looking for specific things, they're breaking down lots of data, or you're breaking down lots of data, or you're breaking down lots of data, or you're just processing things and generating just processing things and generating just processing things and generating simple text outputs. It's pretty cool simple text outputs. It's pretty cool simple text outputs. It's pretty cool having a model that is that capable, having a model that is that capable, having a model that is that capable, that intelligent, that cheap, and also that intelligent, that cheap, and also that intelligent, that cheap, and also good at calling tools. Honestly, Luna good at calling tools. Honestly, Luna good at calling tools. Honestly, Luna kills my use cases for something like kills my use cases for something like kills my use cases for something like Flash from the Gemini series while also Flash from the Gemini series while also Flash from the Gemini series while also being cheaper, more efficient, faster, being cheaper, more efficient, faster, being cheaper, more efficient, faster, and way more reliable. So, uh, yeah, one and way more reliable. So, uh, yeah, one and way more reliable. So, uh, yeah, one way to think of this list is that Luna way to think of this list is that Luna way to think of this list is that Luna is their attempts to kill Flash because is their attempts to kill Flash because is their attempts to kill Flash because Google fumbled Flash. Terra's there Google fumbled Flash. Terra's there Google fumbled Flash. Terra's there attempts to kill Sonnet because attempts to kill Sonnet because attempts to kill Sonnet because Anthropic fumbled Sonnet and Soul is Anthropic fumbled Sonnet and Soul is Anthropic fumbled Sonnet and Soul is their attempts to kill GBD 5.5 because their attempts to kill GBD 5.5 because their attempts to kill GBD 5.5 because they want to have a better, smarter, they want to have a better, smarter, they want to have a better, smarter, more capable model that can do much
-
more capable model that can do much more capable model that can do much longer tasks. This has been a lot to go longer tasks. This has been a lot to go longer tasks. This has been a lot to go over. I know this is a bit much. I was over. I know this is a bit much. I was over. I know this is a bit much. I was hoping that breaking this up into hoping that breaking this up into hoping that breaking this up into multiple videos would keep me from going multiple videos would keep me from going multiple videos would keep me from going off as long as I did, but I wanted to do off as long as I did, but I wanted to do off as long as I did, but I wanted to do a real review here. And my real review a real review here. And my real review a real review here. And my real review is that 56 is a phenomenal model and is that 56 is a phenomenal model and is that 56 is a phenomenal model and it's the default that I go to for most it's the default that I go to for most it's the default that I go to for most things. But does that mean it's my things. But does that mean it's my things. But does that mean it's my favorite model? Does that mean that it's favorite model? Does that mean that it's favorite model? Does that mean that it's better than Fable? This video is already better than Fable? This video is already better than Fable? This video is already too long, so you'll have to wait for the too long, so you'll have to wait for the too long, so you'll have to wait for the next one to learn my answer to that. You next one to learn my answer to that. You next one to learn my answer to that. You might be surprised, though. In fact, might be surprised, though. In fact, might be surprised, though. In fact, you'll probably be surprised. I have a you'll probably be surprised. I have a you'll probably be surprised. I have a feeling that no one quite knows my take feeling that no one quite knows my take feeling that no one quite knows my take on this overall. So, if you want to hear on this overall. So, if you want to hear on this overall. So, if you want to hear my thoughts there, make sure you hit the my thoughts there, make sure you hit the my thoughts there, make sure you hit the subscribe button and that little bell subscribe button and that little bell subscribe button and that little bell next to it so you know when my next next to it so you know when my next next to it so you know when my next videos are coming out. We're going to videos are coming out. We're going to videos are coming out. We're going to have a lot to talk about with both Fable have a lot to talk about with both Fable have a lot to talk about with both Fable and 56 now finally available for and 56 now finally available for and 56 now finally available for everybody, especially with Fable being everybody, especially with Fable being everybody, especially with Fable being taken away from the subscription plans taken away from the subscription plans taken away from the subscription plans in a very, very short amount of time. in a very, very short amount of time. in a very, very short amount of time. I'm going to go get back to prompting. I'm going to go get back to prompting. I'm going to go get back to prompting. So until next time, peace nerds.
Summary
This tech review focuses on the GBT 56 model, comparing its various versions (Sole, Luna, Terra) and reasoning levels. It references benchmark scores, specifically mentioning DeepSWE where "56 soul on max" achieved a high score at a lower cost than "Fable," and also touches on user sentiment and a sponsor, PostHog, which provides product analytics tools for understanding users. The practical takeaway is that while benchmarks are important, GBT 56's performance and cost-effectiveness are notable, and understanding its options is crucial for effective use, with PostHog recommended for user insights.