← Back
Theo September 4, 2026 44m

It's Here.

Read full transcript 34 segments
  1. There's a model we've all been waiting There's a model we've all been waiting for for quite a while now, and it's for for quite a while now, and it's for for quite a while now, and it's finally here. This is where I want to finally here. This is where I want to finally here. This is where I want to make a joke about how it's Gemini 3.8 make a joke about how it's Gemini 3.8 make a joke about how it's Gemini 3.8 Flash, but my new retention guy said I'm Flash, but my new retention guy said I'm Flash, but my new retention guy said I'm not allowed to. He instead said I should not allowed to. He instead said I should not allowed to. He instead said I should just show you guys this, just show you guys this, just show you guys this, my actual usage with said models. Of my actual usage with said models. Of my actual usage with said models. Of course, my terminal is refreshing right course, my terminal is refreshing right course, my terminal is refreshing right when I pick it up. It's uh annoying like when I pick it up. It's uh annoying like when I pick it up. It's uh annoying like that. But the point I'm trying to make that. But the point I'm trying to make that. But the point I'm trying to make is I've done about $330,000 is I've done about $330,000 is I've done about $330,000 of inference over the last few weeks and of inference over the last few weeks and of inference over the last few weeks and the vast majority of it has been with the vast majority of it has been with the vast majority of it has been with the model that we're here to talk about the model that we're here to talk about the model that we're here to talk about today. As you probably guessed, the today. As you probably guessed, the today. As you probably guessed, the model we're actually talking about today model we're actually talking about today model we're actually talking about today is GPT6 Astra. Yes, it's actually is GPT6 Astra. Yes, it's actually is GPT6 Astra. Yes, it's actually getting the six moniker. We don't have getting the six moniker. We don't have getting the six moniker. We don't have other models in the GPT6 line yet, like other models in the GPT6 line yet, like other models in the GPT6 line yet, like the Soul and Luna and Terra equivalents, the Soul and Luna and Terra equivalents, the Soul and Luna and Terra equivalents, but we do have Astra kind of. We'll talk but we do have Astra kind of. We'll talk but we do have Astra kind of. We'll talk about that in just a second. I've been about that in just a second. I've been about that in just a second. I've been using it for a bit and it truly is using it for a bit and it truly is using it for a bit and it truly is revolutionary in a ton of different revolutionary in a ton of different revolutionary in a ton of different ways. There are so many things I just ways. There are so many things I just ways. There are so many things I just never thought LLMs would be able to do never thought LLMs would be able to do never thought LLMs would be able to do that 6 does incredibly well. It is such that 6 does incredibly well. It is such that 6 does incredibly well. It is such a massive leap over soul. It does feel a massive leap over soul. It does feel a massive leap over soul. It does feel generational. I talked before about how generational. I talked before about how generational. I talked before about how GPT 5.6 felt like the best possible GPT 5.6 felt like the best possible GPT 5.6 felt like the best possible version of a lastg game whereas Fable 5 version of a lastg game whereas Fable 5 version of a lastg game whereas Fable 5 felt like a reasonable version of a felt like a reasonable version of a felt like a reasonable version of a nextg game. GBT6 is the next gen. OpenAI nextg game. GBT6 is the next gen. OpenAI nextg game. GBT6 is the next gen. OpenAI has done it. This model is incredible, has done it. This model is incredible, has done it. This model is incredible, but it also still has its rough edges.

  2. but it also still has its rough edges. but it also still has its rough edges. From code to computer use to 3D modeling From code to computer use to 3D modeling From code to computer use to 3D modeling to office work, this model truly is just to office work, this model truly is just to office work, this model truly is just built different. That is far from the built different. That is far from the built different. That is far from the whole story, though, because we have to whole story, though, because we have to whole story, though, because we have to cover availability, cost, security, cover availability, cost, security, cover availability, cost, security, safety, how it actually works in the safety, how it actually works in the safety, how it actually works in the day-to-day experience using it for code, day-to-day experience using it for code, day-to-day experience using it for code, and of course, all the fun things I've and of course, all the fun things I've and of course, all the fun things I've been building with it, as well as the been building with it, as well as the been building with it, as well as the weird restrictions they have with the weird restrictions they have with the weird restrictions they have with the model releasing currently. I can't wait model releasing currently. I can't wait model releasing currently. I can't wait to show you guys what this model can do to show you guys what this model can do to show you guys what this model can do after a real quick break for today's after a real quick break for today's after a real quick break for today's sponsor. You've heard me talk about sponsor. You've heard me talk about sponsor. You've heard me talk about today's sponsor before. It's Code today's sponsor before. It's Code today's sponsor before. It's Code Rabbit, the AI bot that reviews your Rabbit, the AI bot that reviews your Rabbit, the AI bot that reviews your code for you. But that's not what I want code for you. But that's not what I want code for you. But that's not what I want to talk about. If you're not already to talk about. If you're not already to talk about. If you're not already using an AI code reviewer, you really, using an AI code reviewer, you really, using an AI code reviewer, you really, really should be. But Code Rabbit goes really should be. But Code Rabbit goes really should be. But Code Rabbit goes so much further than that. First off, on so much further than that. First off, on so much further than that. First off, on the code review side, they're not just the code review side, they're not just the code review side, they're not just for GitHub. They work in your IDE and for GitHub. They work in your IDE and for GitHub. They work in your IDE and the CLI as well. So whether you're in VS the CLI as well. So whether you're in VS the CLI as well. So whether you're in VS Code Cursor or something like that, the Code Cursor or something like that, the Code Cursor or something like that, the plug-in is great. And if you want to use plug-in is great. And if you want to use plug-in is great. And if you want to use agents and let them have the reviews agents and let them have the reviews agents and let them have the reviews happen themselves, the CLI is even happen themselves, the CLI is even happen themselves, the CLI is even better. For those of us who like to know better. For those of us who like to know better. For those of us who like to know what's going on in the codebase, change what's going on in the codebase, change what's going on in the codebase, change stack's been awesome. It's a new code stack's been awesome. It's a new code stack's been awesome. It's a new code rabbit feature that lets you look at rabbit feature that lets you look at rabbit feature that lets you look at your PR as separate chunks with a your PR as separate chunks with a your PR as separate chunks with a timeline and a really good summary from timeline and a really good summary from timeline and a really good summary from the top, making it way easier to review the top, making it way easier to review the top, making it way easier to review the code and actually see what's going the code and actually see what's going the code and actually see what's going on. When I started using Chainstack on. When I started using Chainstack on. When I started using Chainstack heavily, I noticed a few bugs which I heavily, I noticed a few bugs which I heavily, I noticed a few bugs which I forwarded over to the team at Code forwarded over to the team at Code forwarded over to the team at Code Rabbit and they used their own agent to Rabbit and they used their own agent to Rabbit and they used their own agent to fix it because they built one of the fix it because they built one of the fix it because they built one of the best Slack agents as well. This isn't best Slack agents as well. This isn't best Slack agents as well. This isn't your usual tag claw and it will make your usual tag claw and it will make your usual tag claw and it will make some changes. It's a full endto-end some changes. It's a full endto-end some changes. It's a full endto-end solution that happens to work through solution that happens to work through solution that happens to work through Slack that has access to the context in Slack that has access to the context in Slack that has access to the context in Slack, linear, Jira, whatever else on Slack, linear, Jira, whatever else on Slack, linear, Jira, whatever else on the ticket side, as well as your data, the ticket side, as well as your data, the ticket side, as well as your data, your email, and more. So, if you get an your email, and more. So, if you get an your email, and more. So, if you get an alert from Data Dog, you can tag in Code alert from Data Dog, you can tag in Code alert from Data Dog, you can tag in Code Rabbit, it will find the PR that caused Rabbit, it will find the PR that caused Rabbit, it will find the PR that caused the issue, read the logs in Data Dog, the issue, read the logs in Data Dog, the issue, read the logs in Data Dog, and then file the follow-up all by

  3. and then file the follow-up all by and then file the follow-up all by itself. When you give an agent this type itself. When you give an agent this type itself. When you give an agent this type of context, the things it does are of context, the things it does are of context, the things it does are magical. And Code Rabbit already has all magical. And Code Rabbit already has all magical. And Code Rabbit already has all the context it needs. Ship with more the context it needs. Ship with more the context it needs. Ship with more confidence and less bugs at confidence and less bugs at confidence and less bugs at soy.link/codabbit. soy.link/codabbit. soy.link/codabbit. I'm going to be so incredibly real with I'm going to be so incredibly real with I'm going to be so incredibly real with you guys. I could probably go off for you guys. I could probably go off for you guys. I could probably go off for five plus hours about this model, but five plus hours about this model, but five plus hours about this model, but instead of doing all of that today, I'm instead of doing all of that today, I'm instead of doing all of that today, I'm going to try to give you guys the best going to try to give you guys the best going to try to give you guys the best possible overview of what the model is, possible overview of what the model is, possible overview of what the model is, what it does, what makes it special, and what it does, what makes it special, and what it does, what makes it special, and when we can use it, and what it's going when we can use it, and what it's going when we can use it, and what it's going to cost when you do. I like the way I to cost when you do. I like the way I to cost when you do. I like the way I structured the Fable 5.1 video, so I'll structured the Fable 5.1 video, so I'll structured the Fable 5.1 video, so I'll do my best to honor a lot of that here do my best to honor a lot of that here do my best to honor a lot of that here by going through in a reasonable order by going through in a reasonable order by going through in a reasonable order the core pieces that we all care most the core pieces that we all care most the core pieces that we all care most about. And with that, we need to start about. And with that, we need to start about. And with that, we need to start with the cost and availability. This with the cost and availability. This with the cost and availability. This ended up surprising me in multiple ways. ended up surprising me in multiple ways. ended up surprising me in multiple ways. First off with the cost. The price is First off with the cost. The price is First off with the cost. The price is nearly identical to using Fable in terms nearly identical to using Fable in terms nearly identical to using Fable in terms of the tokens in and out at $10 per of the tokens in and out at $10 per of the tokens in and out at $10 per million in and $50 per million out. million in and $50 per million out. million in and $50 per million out. There are some edges here that are worth There are some edges here that are worth There are some edges here that are worth noting. For example, they do actually noting. For example, they do actually noting. For example, they do actually offer fast mode on it, which Anthropic offer fast mode on it, which Anthropic offer fast mode on it, which Anthropic doesn't for Fable. They only offer for doesn't for Fable. They only offer for doesn't for Fable. They only offer for Opus. You'll get up to two times the Opus. You'll get up to two times the Opus. You'll get up to two times the speed of the standard processing at speed of the standard processing at speed of the standard processing at around two times the standard price. I around two times the standard price. I around two times the standard price. I like that this is an option for those like that this is an option for those like that this is an option for those who need it because this model, I'll be who need it because this model, I'll be who need it because this model, I'll be frank, can be slow. We'll talk plenty frank, can be slow. We'll talk plenty frank, can be slow. We'll talk plenty about that in just a bit. But this is about that in just a bit. But this is about that in just a bit. But this is only one dimension of the cost. There is only one dimension of the cost. There is only one dimension of the cost. There is also the token efficiency which also the token efficiency which also the token efficiency which fundamentally changes what the costs fundamentally changes what the costs fundamentally changes what the costs actually end up being. As you can guess, actually end up being. As you can guess, actually end up being. As you can guess, it's an OpenAI model, so it's insanely it's an OpenAI model, so it's insanely it's an OpenAI model, so it's insanely efficient. It's sometimes even cheaper efficient. It's sometimes even cheaper efficient. It's sometimes even cheaper than Soul for similar like for-like than Soul for similar like for-like than Soul for similar like for-like tasks. But it also can be expensive tasks. But it also can be expensive tasks. But it also can be expensive because it runs for so long and can because it runs for so long and can because it runs for so long and can complete crazy endto-end work. There are complete crazy endto-end work. There are complete crazy endto-end work. There are other dimensions for costs we need to be other dimensions for costs we need to be other dimensions for costs we need to be considerate of though, like how considerate of though, like how considerate of though, like how expensive are cashed reads? Anthropic

  4. expensive are cashed reads? Anthropic expensive are cashed reads? Anthropic cut cash read pricing by 75% with Fable cut cash read pricing by 75% with Fable cut cash read pricing by 75% with Fable 5.1. And there is no equivalent here 5.1. And there is no equivalent here 5.1. And there is no equivalent here with Astra. You are still paying the with Astra. You are still paying the with Astra. You are still paying the full price for cash reads, which is a full price for cash reads, which is a full price for cash reads, which is a tenth of the normal read price. So, it's tenth of the normal read price. So, it's tenth of the normal read price. So, it's a dollar per million cashed token reads a dollar per million cashed token reads a dollar per million cashed token reads versus the 25 cents for Fable. In the versus the 25 cents for Fable. In the versus the 25 cents for Fable. In the end, Astra is still cheaper just due to end, Astra is still cheaper just due to end, Astra is still cheaper just due to the huge efficiency difference. But the huge efficiency difference. But the huge efficiency difference. But thought that was worth knowing. There is thought that was worth knowing. There is thought that was worth knowing. There is one other pricing catch though, and this one other pricing catch though, and this one other pricing catch though, and this catch has to do with the context window catch has to do with the context window catch has to do with the context window size. This model can go up to a million size. This model can go up to a million size. This model can go up to a million token context which is huge but also has token context which is huge but also has token context which is huge but also has been available for the other open models been available for the other open models been available for the other open models for a bit now. That said, it's not for a bit now. That said, it's not for a bit now. That said, it's not available by default in codeex. Unlike available by default in codeex. Unlike available by default in codeex. Unlike cloud code where it now defaults to a cloud code where it now defaults to a cloud code where it now defaults to a million token context window for feeble million token context window for feeble million token context window for feeble and opus, Astra doesn't. I believe the and opus, Astra doesn't. I believe the and opus, Astra doesn't. I believe the plan is to release it in the 370k token plan is to release it in the 370k token plan is to release it in the 370k token range or so in codeex. But it is worth range or so in codeex. But it is worth range or so in codeex. But it is worth noting that if you go over 272K for your noting that if you go over 272K for your noting that if you go over 272K for your context window size, your input tokens context window size, your input tokens context window size, your input tokens get two times more expensive and your get two times more expensive and your get two times more expensive and your output tokens get 50% more expensive. I output tokens get 50% more expensive. I output tokens get 50% more expensive. I also recently learned that they're also recently learned that they're also recently learned that they're actually implementing an exception in actually implementing an exception in actually implementing an exception in codecs for when you go over the 272k codecs for when you go over the 272k codecs for when you go over the 272k input token limit. So if you do end up input token limit. So if you do end up input token limit. So if you do end up bumping in codecs, you're not going to bumping in codecs, you're not going to bumping in codecs, you're not going to have the multiplicative increase in cost have the multiplicative increase in cost have the multiplicative increase in cost that I was talking about before. It will that I was talking about before. It will that I was talking about before. It will still be more expensive though because still be more expensive though because still be more expensive though because you're using more tokens for every you're using more tokens for every you're using more tokens for every single request, tool call, etc. That single request, tool call, etc. That single request, tool call, etc. That mostly covers the cost stuff I wanted to mostly covers the cost stuff I wanted to mostly covers the cost stuff I wanted to for now at least. But now we have to for now at least. But now we have to for now at least. But now we have to talk more about availability because talk more about availability because talk more about availability because this is where there's something I really this is where there's something I really this is where there's something I really don't like. GBD6 Astra is rolling out don't like. GBD6 Astra is rolling out don't like. GBD6 Astra is rolling out today to a limited set of organizations today to a limited set of organizations today to a limited set of organizations and over the coming days will become and over the coming days will become and over the coming days will become available to all chatbt plus pro available to all chatbt plus pro available to all chatbt plus pro business and enterprise. There is one

  5. business and enterprise. There is one business and enterprise. There is one good piece here which is that plus and good piece here which is that plus and good piece here which is that plus and pro accounts will all be getting Astra pro accounts will all be getting Astra pro accounts will all be getting Astra and they're not going to implement some and they're not going to implement some and they're not going to implement some weird 50% limit like they do with fable weird 50% limit like they do with fable weird 50% limit like they do with fable in your cloud code sub. But it's also in your cloud code sub. But it's also in your cloud code sub. But it's also not actually out. This is genuinely not actually out. This is genuinely not actually out. This is genuinely frustrating for me because I want to frustrating for me because I want to frustrating for me because I want to show you guys all the cool things I did show you guys all the cool things I did show you guys all the cool things I did with the model and you can go replicate with the model and you can go replicate with the model and you can go replicate them and try them yourself, but you them and try them yourself, but you them and try them yourself, but you can't. That sucks. It just It just can't. That sucks. It just It just can't. That sucks. It just It just sucks. I hate this. I know there's a lot sucks. I hate this. I know there's a lot sucks. I hate this. I know there's a lot of layers and chaos and reasoning to it. of layers and chaos and reasoning to it. of layers and chaos and reasoning to it. It's not as simple as they just want to It's not as simple as they just want to It's not as simple as they just want to announce it and then give it out later. announce it and then give it out later. announce it and then give it out later. I wish I had more details. I've been I wish I had more details. I've been I wish I had more details. I've been trying to get them all day. I do not trying to get them all day. I do not trying to get them all day. I do not love the way that they chose to announce love the way that they chose to announce love the way that they chose to announce this as though it is out for everyone, this as though it is out for everyone, this as though it is out for everyone, but in reality it's only a small subset but in reality it's only a small subset but in reality it's only a small subset of people that are now allowed to talk of people that are now allowed to talk of people that are now allowed to talk about it. They also are providing proper about it. They also are providing proper about it. They also are providing proper zero data retention. So all the zero data retention. So all the zero data retention. So all the enterprises that were concerned about enterprises that were concerned about enterprises that were concerned about Fable due to the different policies Fable due to the different policies Fable due to the different policies there have nothing to worry about. there have nothing to worry about. there have nothing to worry about. OpenAI has maintained their high bar for OpenAI has maintained their high bar for OpenAI has maintained their high bar for enterprise support with their ZDR stuff. enterprise support with their ZDR stuff. enterprise support with their ZDR stuff. There's a call out at the end here that There's a call out at the end here that There's a call out at the end here that it will be available over the OpenAI API it will be available over the OpenAI API it will be available over the OpenAI API as well as in Amazon Bedrock. Notably, as well as in Amazon Bedrock. Notably, as well as in Amazon Bedrock. Notably, no mention of Azure here. It seems like no mention of Azure here. It seems like no mention of Azure here. It seems like that Microsoft OpenAI breakup is truly that Microsoft OpenAI breakup is truly that Microsoft OpenAI breakup is truly finalized. Okay, one last tiny thing on finalized. Okay, one last tiny thing on finalized. Okay, one last tiny thing on availability before we can dive into the availability before we can dive into the availability before we can dive into the model itself. OpenAI clearly is not model itself. OpenAI clearly is not model itself. OpenAI clearly is not happy about the limitation to who has happy about the limitation to who has happy about the limitation to who has access and they're choosing to do some access and they're choosing to do some access and they're choosing to do some pretty generous stuff around that. Hebo pretty generous stuff around that. Hebo pretty generous stuff around that. Hebo tweeted that they'll be giving one tweeted that they'll be giving one tweeted that they'll be giving one banked reset for your codec sub for banked reset for your codec sub for banked reset for your codec sub for every single day that we don't have every single day that we don't have every single day that we don't have access to Astra on our page HGBT access to Astra on our page HGBT access to Astra on our page HGBT accounts. This is pretty generous and accounts. This is pretty generous and accounts. This is pretty generous and everyone in the replies seems to agree.

  6. everyone in the replies seems to agree. everyone in the replies seems to agree. In fact, I've even seen people saying In fact, I've even seen people saying In fact, I've even seen people saying that it's okay if they delay Astra that it's okay if they delay Astra that it's okay if they delay Astra indefinitely if they keep giving out indefinitely if they keep giving out indefinitely if they keep giving out resets. So, personally, I don't think resets. So, personally, I don't think resets. So, personally, I don't think this is a big enough solution to what this is a big enough solution to what this is a big enough solution to what is, in my opinion, a quite large is, in my opinion, a quite large is, in my opinion, a quite large problem, but at least they're doing problem, but at least they're doing problem, but at least they're doing something. And I do have confirmation something. And I do have confirmation something. And I do have confirmation from people at OpenAI that this is a from people at OpenAI that this is a from people at OpenAI that this is a temporary measure. They don't expect temporary measure. They don't expect temporary measure. They don't expect future model releases to have this weird future model releases to have this weird future model releases to have this weird window like we do right now. Yeah, it is window like we do right now. Yeah, it is window like we do right now. Yeah, it is what it is. Could be government, could what it is. Could be government, could what it is. Could be government, could be something weird with their compute be something weird with their compute be something weird with their compute layer, could be some enterprise layer, could be some enterprise layer, could be some enterprise customers or Microsoft being upset. I customers or Microsoft being upset. I customers or Microsoft being upset. I have no idea. I honestly, I'll be real. have no idea. I honestly, I'll be real. have no idea. I honestly, I'll be real. kind of thinking it's Microsoft, but we kind of thinking it's Microsoft, but we kind of thinking it's Microsoft, but we don't know. We can wait. We'll have don't know. We can wait. We'll have don't know. We can wait. We'll have answers soon. Might even be AWS being answers soon. Might even be AWS being answers soon. Might even be AWS being upset that they're not ready yet and upset that they're not ready yet and upset that they're not ready yet and they don't want OpenAI offering it they don't want OpenAI offering it they don't want OpenAI offering it before them. I don't know. It is what it before them. I don't know. It is what it before them. I don't know. It is what it is. Enough yapping. Let's actually take is. Enough yapping. Let's actually take is. Enough yapping. Let's actually take a look at what this model can do. They a look at what this model can do. They a look at what this model can do. They start off with a small set of benchmarks start off with a small set of benchmarks start off with a small set of benchmarks that it absolutely seems to slaughter. I that it absolutely seems to slaughter. I that it absolutely seems to slaughter. I have noticed that in a handful of these have noticed that in a handful of these have noticed that in a handful of these benches, it performed slightly worse on benches, it performed slightly worse on benches, it performed slightly worse on X high than it did on regular high. X high than it did on regular high. X high than it did on regular high. Sometimes max comes out and beats it Sometimes max comes out and beats it Sometimes max comes out and beats it out, sometimes it doesn't. I does seem out, sometimes it doesn't. I does seem out, sometimes it doesn't. I does seem to be one of the best options with this to be one of the best options with this to be one of the best options with this model, though. So, here in Terminal model, though. So, here in Terminal model, though. So, here in Terminal Bench, it crushed other models, Bench, it crushed other models, Bench, it crushed other models, including Fable 5.1, getting way higher including Fable 5.1, getting way higher including Fable 5.1, getting way higher scores at similar costs. This is the scores at similar costs. This is the scores at similar costs. This is the science version of Terminal Bench, but science version of Terminal Bench, but science version of Terminal Bench, but the numbers we're seeing here are pretty the numbers we're seeing here are pretty the numbers we're seeing here are pretty crazy, where Claude Fable 5.1 on medium crazy, where Claude Fable 5.1 on medium crazy, where Claude Fable 5.1 on medium cost $15 and got a 36% and six Astra on cost $15 and got a 36% and six Astra on cost $15 and got a 36% and six Astra on low cost $11 and got a 54%. The gap here low cost $11 and got a 54%. The gap here low cost $11 and got a 54%. The gap here is massive. The science side, as I is massive. The science side, as I is massive. The science side, as I mentioned before, really seems to be mentioned before, really seems to be mentioned before, really seems to be something OpenAI is focused on. And this something OpenAI is focused on. And this something OpenAI is focused on. And this is particularly funny because Anthropic is particularly funny because Anthropic is particularly funny because Anthropic was just bragging about how good Fable was just bragging about how good Fable was just bragging about how good Fable 5.1 did on the same exact bench just to

  7. 5.1 did on the same exact bench just to 5.1 did on the same exact bench just to be slaughtered by OpenAI two days later. be slaughtered by OpenAI two days later. be slaughtered by OpenAI two days later. Speaking of OpenAI slaughtering Speaking of OpenAI slaughtering Speaking of OpenAI slaughtering Anthropic, Arc AGI is uh yeah, I'm going Anthropic, Arc AGI is uh yeah, I'm going Anthropic, Arc AGI is uh yeah, I'm going to be real with you guys. I thought this to be real with you guys. I thought this to be real with you guys. I thought this bench was malicious. It was so absurdly bench was malicious. It was so absurdly bench was malicious. It was so absurdly built to be anti-AII, it was almost built to be anti-AII, it was almost built to be anti-AII, it was almost funny because it wasn't just measuring funny because it wasn't just measuring funny because it wasn't just measuring if the AI could complete the tasks the if the AI could complete the tasks the if the AI could complete the tasks the way a human could. It was also measuring way a human could. It was also measuring way a human could. It was also measuring how many steps it took to do it. So how many steps it took to do it. So how many steps it took to do it. So every time it had a tool call or every time it had a tool call or every time it had a tool call or reasoned that was held against the model reasoned that was held against the model reasoned that was held against the model if the human had done it in fewer if the human had done it in fewer if the human had done it in fewer perceived steps. That scoring rate was perceived steps. That scoring rate was perceived steps. That scoring rate was pretty insane especially because the pretty insane especially because the pretty insane especially because the model could never outperform the human. model could never outperform the human. model could never outperform the human. So if the human took 20 steps and the So if the human took 20 steps and the So if the human took 20 steps and the model took five, it still just got a model took five, it still just got a model took five, it still just got a regular neutral score. But if it took regular neutral score. But if it took regular neutral score. But if it took one step more than the human, it was one step more than the human, it was one step more than the human, it was penalized massively. And despite all of penalized massively. And despite all of penalized massively. And despite all of that, it is now saturated at 99.9% that, it is now saturated at 99.9% that, it is now saturated at 99.9% on a bench that was literally straight on a bench that was literally straight on a bench that was literally straight zeros when it came out just under a year zeros when it came out just under a year zeros when it came out just under a year ago. Then we have Frontier Math where it ago. Then we have Frontier Math where it ago. Then we have Frontier Math where it again slaughtered. Fable's best score again slaughtered. Fable's best score again slaughtered. Fable's best score was an 87.8% with Fable 5.1 and that got was an 87.8% with Fable 5.1 and that got was an 87.8% with Fable 5.1 and that got matched by Astra on low and then matched by Astra on low and then matched by Astra on low and then everything medium upwards pretty much everything medium upwards pretty much everything medium upwards pretty much got 100%. They all flatlined at 97.6, got 100%. They all flatlined at 97.6, got 100%. They all flatlined at 97.6, which is interesting, but yeah, very which is interesting, but yeah, very which is interesting, but yeah, very good scores. Turtle Bench 4 also got good scores. Turtle Bench 4 also got good scores. Turtle Bench 4 also got crushed previously. Fable 5.1 actually crushed previously. Fable 5.1 actually crushed previously. Fable 5.1 actually looked very promising here. When looked very promising here. When looked very promising here. When combined with Soul, it almost felt like combined with Soul, it almost felt like combined with Soul, it almost felt like a semicontinuous line of more cost means a semicontinuous line of more cost means a semicontinuous line of more cost means better performance. But now with Astra, better performance. But now with Astra, better performance. But now with Astra, it's crashed. It's absolutely destroyed.

  8. it's crashed. It's absolutely destroyed. it's crashed. It's absolutely destroyed. You're getting the higher scores than You're getting the higher scores than You're getting the higher scores than the best that Fable can do at under half the best that Fable can do at under half the best that Fable can do at under half the price. But again, we see that weird the price. But again, we see that weird the price. But again, we see that weird trend where X high and max score trend where X high and max score trend where X high and max score slightly worse. The creator of ArcGI slightly worse. The creator of ArcGI slightly worse. The creator of ArcGI left a quote here for us that I think is left a quote here for us that I think is left a quote here for us that I think is pretty telling of where we're at. Astra pretty telling of where we're at. Astra pretty telling of where we're at. Astra surpassed our human action efficiency surpassed our human action efficiency surpassed our human action efficiency baseline on 96% of levels, effectively baseline on 96% of levels, effectively baseline on 96% of levels, effectively reaching human parody on the benchmark. reaching human parody on the benchmark. reaching human parody on the benchmark. Not only is this the best model we've Not only is this the best model we've Not only is this the best model we've ever tested, but it also represents a ever tested, but it also represents a ever tested, but it also represents a meaningful step change in frontier model meaningful step change in frontier model meaningful step change in frontier model performance, not only in its ability to performance, not only in its ability to performance, not only in its ability to navigate and solve novel environments, navigate and solve novel environments, navigate and solve novel environments, but also in how efficiently it learns to but also in how efficiently it learns to but also in how efficiently it learns to do so. This idea of the model learning do so. This idea of the model learning do so. This idea of the model learning is a thing that it seems like OpenAI is is a thing that it seems like OpenAI is is a thing that it seems like OpenAI is starting to push. It isn't literally starting to push. It isn't literally starting to push. It isn't literally learning like the weights aren't learning like the weights aren't learning like the weights aren't adjusting as you go. But its ability to adjusting as you go. But its ability to adjusting as you go. But its ability to keep things in context to work for a keep things in context to work for a keep things in context to work for a long time and to compact when it's out long time and to compact when it's out long time and to compact when it's out of space and retain most of what it of space and retain most of what it of space and retain most of what it learned over that time. Obviously, the learned over that time. Obviously, the learned over that time. Obviously, the model isn't actually learning. It's not model isn't actually learning. It's not model isn't actually learning. It's not like the weights are adjusting based on like the weights are adjusting based on like the weights are adjusting based on what you ask it. But it's so good at what you ask it. But it's so good at what you ask it. But it's so good at going for a long time and honoring the going for a long time and honoring the going for a long time and honoring the context it has. And more importantly, context it has. And more importantly, context it has. And more importantly, when it runs out of space, compacting in when it runs out of space, compacting in when it runs out of space, compacting in a way that continues to honor what it a way that continues to honor what it a way that continues to honor what it was doing and what it has learned during was doing and what it has learned during was doing and what it has learned during that session, which allows the model to that session, which allows the model to that session, which allows the model to adapt to weird places and weird tasks adapt to weird places and weird tasks adapt to weird places and weird tasks significantly better than anything I've significantly better than anything I've significantly better than anything I've used before. It's also really aligned.

  9. used before. It's also really aligned. used before. It's also really aligned. They showed an exploit gym honeypotss They showed an exploit gym honeypotss They showed an exploit gym honeypotss where they made something to see if the where they made something to see if the where they made something to see if the model would fall for a weird hack case. model would fall for a weird hack case. model would fall for a weird hack case. Soul would fall for it almost 50% of the Soul would fall for it almost 50% of the Soul would fall for it almost 50% of the time and Astra does 0%. And here's where time and Astra does 0%. And here's where time and Astra does 0%. And here's where we start to get into the really fun we start to get into the really fun we start to get into the really fun novel stuff. I will show you guys some novel stuff. I will show you guys some novel stuff. I will show you guys some of it in action in a bit, but for now, I of it in action in a bit, but for now, I of it in action in a bit, but for now, I just need you to trust me. If you care a just need you to trust me. If you care a just need you to trust me. If you care a lot about computer use, there is nothing lot about computer use, there is nothing lot about computer use, there is nothing even close. This model is not just like even close. This model is not just like even close. This model is not just like a single generational leap in computer a single generational leap in computer a single generational leap in computer use. It feels like two or three. It's use. It feels like two or three. It's use. It feels like two or three. It's absurd. And it makes anthropic models absurd. And it makes anthropic models absurd. And it makes anthropic models almost feel like they're in the stone almost feel like they're in the stone almost feel like they're in the stone age. It's similar to the gap of how much age. It's similar to the gap of how much age. It's similar to the gap of how much worse OpenAI models were at front end in worse OpenAI models were at front end in worse OpenAI models were at front end in design compared to anthropics models, design compared to anthropics models, design compared to anthropics models, but now even bigger and for computer use but now even bigger and for computer use but now even bigger and for computer use the opposite direction. I just don't the opposite direction. I just don't the opposite direction. I just don't have a good time having Fable do things have a good time having Fable do things have a good time having Fable do things on my computer like at all. Meanwhile, on my computer like at all. Meanwhile, on my computer like at all. Meanwhile, Codex uses my computer more than I do at Codex uses my computer more than I do at Codex uses my computer more than I do at this point. The benchmarks show this, this point. The benchmarks show this, this point. The benchmarks show this, but not quite as deeply as I would like but not quite as deeply as I would like but not quite as deeply as I would like it to. Agents Last Exam shows really it to. Agents Last Exam shows really it to. Agents Last Exam shows really good scores for reasonable prices. good scores for reasonable prices. good scores for reasonable prices. Screenshot Pro shows way better scores, Screenshot Pro shows way better scores, Screenshot Pro shows way better scores, but also meaningfully more expensive but also meaningfully more expensive but also meaningfully more expensive than Soul was. And OSWorld shows better than Soul was. And OSWorld shows better than Soul was. And OSWorld shows better scores than anything, including the best scores than anything, including the best scores than anything, including the best from Ananthropic, at much lower prices.

  10. from Ananthropic, at much lower prices. from Ananthropic, at much lower prices. There is one other part here that I want There is one other part here that I want There is one other part here that I want to show though, which is the time it to show though, which is the time it to show though, which is the time it took to get these scores because this is took to get these scores because this is took to get these scores because this is where I'm most impressed with the new where I'm most impressed with the new where I'm most impressed with the new model. It's so much faster at these model. It's so much faster at these model. It's so much faster at these computer use tasks. On high, Astro was computer use tasks. On high, Astro was computer use tasks. On high, Astro was able to complete OSWorld 2.0 in about 23 able to complete OSWorld 2.0 in about 23 able to complete OSWorld 2.0 in about 23 minutes with a 71.6% minutes with a 71.6% minutes with a 71.6% score. Soul's best score for reference score. Soul's best score for reference score. Soul's best score for reference was a 65.7% and it took almost an hour was a 65.7% and it took almost an hour was a 65.7% and it took almost an hour and 15 minutes to complete. That's a and 15 minutes to complete. That's a and 15 minutes to complete. That's a huge difference in the time taken for a huge difference in the time taken for a huge difference in the time taken for a worse result. This model flies in worse result. This model flies in worse result. This model flies in computer use. They have a bunch of fun computer use. They have a bunch of fun computer use. They have a bunch of fun demos of it doing this type of thing in demos of it doing this type of thing in demos of it doing this type of thing in real apps. Like in here, they're showing real apps. Like in here, they're showing real apps. Like in here, they're showing it using Excel directly, not editing the it using Excel directly, not editing the it using Excel directly, not editing the file programmatically, but actually file programmatically, but actually file programmatically, but actually telling the mouse cursor where to go and telling the mouse cursor where to go and telling the mouse cursor where to go and the keyboard what to type. And it's able the keyboard what to type. And it's able the keyboard what to type. And it's able to pretty quickly fly through Excel. I to pretty quickly fly through Excel. I to pretty quickly fly through Excel. I am impressed with this, especially with am impressed with this, especially with am impressed with this, especially with like PowerPoint designs and things like like PowerPoint designs and things like like PowerPoint designs and things like that. It takes a while to get going that. It takes a while to get going that. It takes a while to get going because it has to think about what it because it has to think about what it because it has to think about what it wants to do. Once it's decided what it wants to do. Once it's decided what it wants to do. Once it's decided what it wants to do, it just guns through all wants to do, it just guns through all wants to do, it just guns through all the changes. Even here, it took a while the changes. Even here, it took a while the changes. Even here, it took a while to get going, but around this 1 minute to get going, but around this 1 minute to get going, but around this 1 minute 20 second mark, it all of a sudden just 20 second mark, it all of a sudden just 20 second mark, it all of a sudden just starts flying through things. There's starts flying through things. There's starts flying through things. There's also a bunch of examples of game also a bunch of examples of game also a bunch of examples of game development type stuff, and honestly, development type stuff, and honestly, development type stuff, and honestly, this is some of the most impressive this is some of the most impressive this is some of the most impressive parts of what I've seen. This model's parts of what I've seen. This model's parts of what I've seen. This model's understanding of 3D spaces and also the understanding of 3D spaces and also the understanding of 3D spaces and also the tools you use to work in 3D is tools you use to work in 3D is tools you use to work in 3D is unbelievable. The things I've seen it unbelievable. The things I've seen it unbelievable. The things I've seen it create in Blender have melted my brain, create in Blender have melted my brain, create in Blender have melted my brain, and I will be sure to show you some of and I will be sure to show you some of and I will be sure to show you some of the ones I built as well as we go along.

  11. the ones I built as well as we go along. the ones I built as well as we go along. They claim it's around 1.9 times faster They claim it's around 1.9 times faster They claim it's around 1.9 times faster to complete tasks with computer use, to complete tasks with computer use, to complete tasks with computer use, both because of Astra being more both because of Astra being more both because of Astra being more efficient and because of improvements in efficient and because of improvements in efficient and because of improvements in codecs. And honestly, that lines up. I codecs. And honestly, that lines up. I codecs. And honestly, that lines up. I had to go through and download all the had to go through and download all the had to go through and download all the medical records for my broken ass hand. medical records for my broken ass hand. medical records for my broken ass hand. And it flew through it despite the super And it flew through it despite the super And it flew through it despite the super slow medical dashboards it was slow medical dashboards it was slow medical dashboards it was navigating. Did the whole thing in 15 navigating. Did the whole thing in 15 navigating. Did the whole thing in 15 minutes when I like went to go grab some minutes when I like went to go grab some minutes when I like went to go grab some food. It had to navigate like 150 pages food. It had to navigate like 150 pages food. It had to navigate like 150 pages in that time, too. It was really, really in that time, too. It was really, really in that time, too. It was really, really impressive. Again to emphasize the 3D impressive. Again to emphasize the 3D impressive. Again to emphasize the 3D capabilities, they have some benchmarks capabilities, they have some benchmarks capabilities, they have some benchmarks here like Bench CAD, which is using here like Bench CAD, which is using here like Bench CAD, which is using Python to do some CAD work, and it Python to do some CAD work, and it Python to do some CAD work, and it crushed everything else. Even Soul was crushed everything else. Even Soul was crushed everything else. Even Soul was ahead of what Fable could do here, but ahead of what Fable could do here, but ahead of what Fable could do here, but Astra is getting close to 100% at its Astra is getting close to 100% at its Astra is getting close to 100% at its peak. Whereas the best Fable could do peak. Whereas the best Fable could do peak. Whereas the best Fable could do was only an 84% and it cost over $11. was only an 84% and it cost over $11. was only an 84% and it cost over $11. Meanwhile, Astra is doing the same work Meanwhile, Astra is doing the same work Meanwhile, Astra is doing the same work for under two bucks, but 96% accuracy. for under two bucks, but 96% accuracy. for under two bucks, but 96% accuracy. Pretty nuts. They have some examples of Pretty nuts. They have some examples of Pretty nuts. They have some examples of it creating slideshows. And I'll admit, it creating slideshows. And I'll admit, it creating slideshows. And I'll admit, the demos they showed here were the demos they showed here were the demos they showed here were incredible. But for my experience, incredible. But for my experience, incredible. But for my experience, asking it to do similar things. The asking it to do similar things. The asking it to do similar things. The computer use side was crazy. Like computer use side was crazy. Like computer use side was crazy. Like watching it actually navigate my watching it actually navigate my watching it actually navigate my computer was nuts. But the quality of computer was nuts. But the quality of computer was nuts. But the quality of the slides it generated just was not as the slides it generated just was not as the slides it generated just was not as good as what I'm seeing in these. I'm good as what I'm seeing in these. I'm good as what I'm seeing in these. I'm sure there was some amount of like sure there was some amount of like sure there was some amount of like prompt issue or not giving it the prompt issue or not giving it the prompt issue or not giving it the context or whatever. But uh yeah, skill context or whatever. But uh yeah, skill context or whatever. But uh yeah, skill issue me all you want. I did not think issue me all you want. I did not think issue me all you want. I did not think this demo reflected my real world usage this demo reflected my real world usage this demo reflected my real world usage quite as accurately as the other things.

  12. quite as accurately as the other things. quite as accurately as the other things. One thing that I do think is One thing that I do think is One thing that I do think is surprisingly accurate despite everybody surprisingly accurate despite everybody surprisingly accurate despite everybody saying otherwise is how absurd it is at saying otherwise is how absurd it is at saying otherwise is how absurd it is at 3D. This is a Blender scene that it made 3D. This is a Blender scene that it made 3D. This is a Blender scene that it made and then turned it walkable inside of and then turned it walkable inside of and then turned it walkable inside of Unreal Engine 5 and then created this Unreal Engine 5 and then created this Unreal Engine 5 and then created this video showcasing a house that it built video showcasing a house that it built video showcasing a house that it built in Blender and now rendered in Unreal. in Blender and now rendered in Unreal. in Blender and now rendered in Unreal. Absurd levels of detail here. It's it Absurd levels of detail here. It's it Absurd levels of detail here. It's it understands three dimensions so well. understands three dimensions so well. understands three dimensions so well. They showcase some games that they had They showcase some games that they had They showcase some games that they had at make. And while they're cool, I'm at make. And while they're cool, I'm at make. And while they're cool, I'm admittedly biased. I think mine are admittedly biased. I think mine are admittedly biased. I think mine are cooler. I did notice here though that it cooler. I did notice here though that it cooler. I did notice here though that it has almost the exact same UI as the one has almost the exact same UI as the one has almost the exact same UI as the one that I made. It seems like they're again that I made. It seems like they're again that I made. It seems like they're again part of this data set that I think's part of this data set that I think's part of this data set that I think's been going around for 3D stuff and has a been going around for 3D stuff and has a been going around for 3D stuff and has a lot of those same 3D characteristics lot of those same 3D characteristics lot of those same 3D characteristics that I've noticed from other models. And that I've noticed from other models. And that I've noticed from other models. And here's what it created. Oh, wait, no. here's what it created. Oh, wait, no. here's what it created. Oh, wait, no. This is the Opus 5 version. This is what This is the Opus 5 version. This is what This is the Opus 5 version. This is what it actually created. I'm still kind of it actually created. I'm still kind of it actually created. I'm still kind of in shock at how good of a job it did. As in shock at how good of a job it did. As in shock at how good of a job it did. As I said, it has that same like yellow I said, it has that same like yellow I said, it has that same like yellow dive in button that I saw in almost dive in button that I saw in almost dive in button that I saw in almost every other 3D gen this model made. Once every other 3D gen this model made. Once every other 3D gen this model made. Once we're in, we're in, we're in, yeah, yeah, yeah, the fish actually look like fish. Not the fish actually look like fish. Not the fish actually look like fish. Not like almost like a fish. This is pretty like almost like a fish. This is pretty like almost like a fish. This is pretty close to what I would expect a real close to what I would expect a real close to what I would expect a real artist to make if asked. It has artist to make if asked. It has artist to make if asked. It has animations that make sense. It has animations that make sense. It has animations that make sense. It has gameplay loops that function. It has gameplay loops that function. It has gameplay loops that function. It has nice little animations when the fish nice little animations when the fish nice little animations when the fish finally get their food. It's It's finally get their food. It's It's finally get their food. It's It's unbelievable the quality of the things unbelievable the quality of the things unbelievable the quality of the things it rendered here. Like all of the it rendered here. Like all of the it rendered here. Like all of the geometry of this stuff at the bottom of

  13. geometry of this stuff at the bottom of geometry of this stuff at the bottom of the tank, the quality of the stuff it the tank, the quality of the stuff it the tank, the quality of the stuff it put in here, the lighting is great. The put in here, the lighting is great. The put in here, the lighting is great. The shaders are great. It's shaders are great. It's shaders are great. It's It just is great. It did an unbelievable It just is great. It did an unbelievable It just is great. It did an unbelievable job. There were edges, though. I job. There were edges, though. I job. There were edges, though. I mentioned in the Fable 5.1 video that it mentioned in the Fable 5.1 video that it mentioned in the Fable 5.1 video that it got all the controls perfectly, and it got all the controls perfectly, and it got all the controls perfectly, and it actually felt nice to navigate and play. actually felt nice to navigate and play. actually felt nice to navigate and play. Astra did not even come close to doing Astra did not even come close to doing Astra did not even come close to doing any of those things. Astra's version did any of those things. Astra's version did any of those things. Astra's version did not get the controls right at all. not get the controls right at all. not get the controls right at all. Movement sucked and was janky. The mouse Movement sucked and was janky. The mouse Movement sucked and was janky. The mouse movement in particular was way too fast movement in particular was way too fast movement in particular was way too fast initially, so I told it to slow it down initially, so I told it to slow it down initially, so I told it to slow it down a bit. Then it made it way too slow. All a bit. Then it made it way too slow. All a bit. Then it made it way too slow. All those little tasteful bits that actually those little tasteful bits that actually those little tasteful bits that actually make it pleasant to play, it got wrong. make it pleasant to play, it got wrong. make it pleasant to play, it got wrong. I was able to steer it in the right I was able to steer it in the right I was able to steer it in the right direction by telling it what I didn't direction by telling it what I didn't direction by telling it what I didn't like and how to fix it. and it mostly like and how to fix it. and it mostly like and how to fix it. and it mostly started to get things right, but it was started to get things right, but it was started to get things right, but it was more back and forth after that first more back and forth after that first more back and forth after that first shot. Even though it looked this good as shot. Even though it looked this good as shot. Even though it looked this good as soon as I initially ran the prompt. So, soon as I initially ran the prompt. So, soon as I initially ran the prompt. So, I am absolutely blown away with what I am absolutely blown away with what I am absolutely blown away with what this model does. In particular, when you this model does. In particular, when you this model does. In particular, when you give it Blender access and tell it to give it Blender access and tell it to give it Blender access and tell it to create something like this, it's so far create something like this, it's so far create something like this, it's so far ahead in this type of 3D design work ahead in this type of 3D design work ahead in this type of 3D design work that it's unfair to even compare to that it's unfair to even compare to that it's unfair to even compare to other models. It really does feel like a other models. It really does feel like a other models. It really does feel like a generational leap here. And again, for generational leap here. And again, for generational leap here. And again, for reference, here is the Fable 5.1 version reference, here is the Fable 5.1 version reference, here is the Fable 5.1 version that I just demoed in my previous video.

  14. that I just demoed in my previous video. that I just demoed in my previous video. I hope you could see the difference I hope you could see the difference I hope you could see the difference here. I think it's pretty absurd. The here. I think it's pretty absurd. The here. I think it's pretty absurd. The next section of the article calls out next section of the article calls out next section of the article calls out how much better it is to actually how much better it is to actually how much better it is to actually interact with. Specifies that when interact with. Specifies that when interact with. Specifies that when instructions leave room for instructions leave room for instructions leave room for interpretation, Astra is better than interpretation, Astra is better than interpretation, Astra is better than previous models at making the right previous models at making the right previous models at making the right call. It uses context to fill in routine call. It uses context to fill in routine call. It uses context to fill in routine gaps and asks focused questions when the gaps and asks focused questions when the gaps and asks focused questions when the answer could change the outcome. In answer could change the outcome. In answer could change the outcome. In Codeex, it can ask asynchronously while Codeex, it can ask asynchronously while Codeex, it can ask asynchronously while continuing work that doesn't depend on continuing work that doesn't depend on continuing work that doesn't depend on your reply. This I found to actually be your reply. This I found to actually be your reply. This I found to actually be very nice. I noticed initially that it very nice. I noticed initially that it very nice. I noticed initially that it would ask a question in line, but then would ask a question in line, but then would ask a question in line, but then not wait for you to give an answer. I not wait for you to give an answer. I not wait for you to give an answer. I think they trained it to do that before think they trained it to do that before think they trained it to do that before they had the feature in codeex. So now, they had the feature in codeex. So now, they had the feature in codeex. So now, of course, the feature is in codeex as of course, the feature is in codeex as of course, the feature is in codeex as well as the latest T3 code nightly. well as the latest T3 code nightly. well as the latest T3 code nightly. It'll probably be in T3 Code stable It'll probably be in T3 Code stable It'll probably be in T3 Code stable before the model's out. We'll see when before the model's out. We'll see when before the model's out. We'll see when we can get a release out. Regardless, we can get a release out. Regardless, we can get a release out. Regardless, it's a nice behavior that the model can it's a nice behavior that the model can it's a nice behavior that the model can just like ask a question like, should I just like ask a question like, should I just like ask a question like, should I go this direction or that direction? and go this direction or that direction? and go this direction or that direction? and then continue working while also waiting then continue working while also waiting then continue working while also waiting for you to respond. So much better for for you to respond. So much better for for you to respond. So much better for having the model keep you in the loop having the model keep you in the loop having the model keep you in the loop while also being able to work in while also being able to work in while also being able to work in parallel. It's also better at staying parallel. It's also better at staying parallel. It's also better at staying oriented as tasks evolve. Earlier models oriented as tasks evolve. Earlier models oriented as tasks evolve. Earlier models sometimes treated steering messages as sometimes treated steering messages as sometimes treated steering messages as new goals, losing track of the original new goals, losing track of the original new goals, losing track of the original request or earlier constraints. Yep, request or earlier constraints. Yep, request or earlier constraints. Yep, this model does that significantly less this model does that significantly less this model does that significantly less and it's very nice. I will say that its and it's very nice. I will say that its and it's very nice. I will say that its interpretive capabilities and how well interpretive capabilities and how well interpretive capabilities and how well it understands what I want does have it understands what I want does have it understands what I want does have edges and I'll be sure to talk about edges and I'll be sure to talk about edges and I'll be sure to talk about those a lot in the future, especially in those a lot in the future, especially in those a lot in the future, especially in my Fable 5.1 versus Astro video that I'm my Fable 5.1 versus Astro video that I'm my Fable 5.1 versus Astro video that I'm certainly going to have to do in the certainly going to have to do in the certainly going to have to do in the near future. So, yeah, know that. But near future. So, yeah, know that. But near future. So, yeah, know that. But that does not mean this model isn't that does not mean this model isn't that does not mean this model isn't unbelievable. Anthropic does have bold unbelievable. Anthropic does have bold unbelievable. Anthropic does have bold claims about this being the best model claims about this being the best model claims about this being the best model for software engineers to date. Not

  15. for software engineers to date. Not for software engineers to date. Not saying best by OpenAI, saying best in saying best by OpenAI, saying best in saying best by OpenAI, saying best in general. We'll have a lot to say there general. We'll have a lot to say there general. We'll have a lot to say there in that same Fable versus Astra video. in that same Fable versus Astra video. in that same Fable versus Astra video. But I will say at the least for now, it But I will say at the least for now, it But I will say at the least for now, it is unbelievably better than what I was is unbelievably better than what I was is unbelievably better than what I was getting out of 5.6 Soul to the point getting out of 5.6 Soul to the point getting out of 5.6 Soul to the point where I can actually somewhat where I can actually somewhat where I can actually somewhat confidently merge changes from this confidently merge changes from this confidently merge changes from this model where that was not the case model where that was not the case model where that was not the case previously when I was using Soul. I previously when I was using Soul. I previously when I was using Soul. I found that Soul, while able to solve found that Soul, while able to solve found that Soul, while able to solve problems really well, tended to leave problems really well, tended to leave problems really well, tended to leave messes behind along the way and not messes behind along the way and not messes behind along the way and not necessarily do things the way I wanted. necessarily do things the way I wanted. necessarily do things the way I wanted. often bloating the PRs that it was often bloating the PRs that it was often bloating the PRs that it was working on way beyond the point that working on way beyond the point that working on way beyond the point that made sense, writing tests that didn't made sense, writing tests that didn't made sense, writing tests that didn't need to be there at all and just going need to be there at all and just going need to be there at all and just going too far in places it probably shouldn't. too far in places it probably shouldn't. too far in places it probably shouldn't. Astra is much more restrained in those Astra is much more restrained in those Astra is much more restrained in those ways and seems to understand the scope ways and seems to understand the scope ways and seems to understand the scope of the changes is making better overall. of the changes is making better overall. of the changes is making better overall. There was one particular funny result in There was one particular funny result in There was one particular funny result in the benches they had for code though the benches they had for code though the benches they had for code though right here with deepsw SWE. They had a right here with deepsw SWE. They had a right here with deepsw SWE. They had a very high score actually I'm pretty sure very high score actually I'm pretty sure very high score actually I'm pretty sure it's the highest right now of the 74.1% it's the highest right now of the 74.1% it's the highest right now of the 74.1% with Astra on XH high again showing that with Astra on XH high again showing that with Astra on XH high again showing that behavior where it goes down on max all behavior where it goes down on max all behavior where it goes down on max all the way down to 73%. But also someone's the way down to 73%. But also someone's the way down to 73%. But also someone's over here Gemini 38 Flash at a 73.8%.

  16. over here Gemini 38 Flash at a 73.8%. over here Gemini 38 Flash at a 73.8%. Yeah, I have to talk about that model in Yeah, I have to talk about that model in Yeah, I have to talk about that model in the future. Not going to do it just yet. the future. Not going to do it just yet. the future. Not going to do it just yet. I'll have a lot more than I thought to I'll have a lot more than I thought to I'll have a lot more than I thought to say about the Google models because say about the Google models because say about the Google models because Google did finally give us permission to Google did finally give us permission to Google did finally give us permission to add anti-gravity to T3 code. So, I've add anti-gravity to T3 code. So, I've add anti-gravity to T3 code. So, I've been using 38 flash high more than I been using 38 flash high more than I been using 38 flash high more than I would have expected. It has issues. It would have expected. It has issues. It would have expected. It has issues. It definitely has issues, but uh can be definitely has issues, but uh can be definitely has issues, but uh can be capable. Let's finish getting through capable. Let's finish getting through capable. Let's finish getting through these release notes so we can get to the these release notes so we can get to the these release notes so we can get to the other fun parts. There's the section other fun parts. There's the section other fun parts. There's the section about advancing scientific discovery, about advancing scientific discovery, about advancing scientific discovery, which is all of the fun science benches which is all of the fun science benches which is all of the fun science benches that they absolutely slaughtered. As you that they absolutely slaughtered. As you that they absolutely slaughtered. As you can guess, on every single one of these, can guess, on every single one of these, can guess, on every single one of these, they are the best and now also the they are the best and now also the they are the best and now also the cheapest. And again, with token cheapest. And again, with token cheapest. And again, with token efficiency, they are killing it here, efficiency, they are killing it here, efficiency, they are killing it here, too. I was actually surprised to see on too. I was actually surprised to see on too. I was actually surprised to see on some of these benches, for example, in some of these benches, for example, in some of these benches, for example, in health bench, the token efficiency health bench, the token efficiency health bench, the token efficiency between Fable 5 and Astra is similar, between Fable 5 and Astra is similar, between Fable 5 and Astra is similar, but Fable 5.1 became much less token but Fable 5.1 became much less token but Fable 5.1 became much less token efficient on these benches. Very efficient on these benches. Very efficient on these benches. Very interesting. And then we have the cyber interesting. And then we have the cyber interesting. And then we have the cyber security section. How good is it at security section. How good is it at security section. How good is it at finding exploits when it's not in the finding exploits when it's not in the finding exploits when it's not in the safeguarded, careful, don't do anything safeguarded, careful, don't do anything safeguarded, careful, don't do anything dangerous mode? The answer is really dangerous mode? The answer is really dangerous mode? The answer is really good. Even at its lowest settings, it is good. Even at its lowest settings, it is good. Even at its lowest settings, it is the first model to get a perfect score the first model to get a perfect score the first model to get a perfect score on exploit bench. Yes, 100% on low. For on exploit bench. Yes, 100% on low. For on exploit bench. Yes, 100% on low. For some reason, high was cheaper than low.

  17. some reason, high was cheaper than low. some reason, high was cheaper than low. I don't know if this is an issue in how I don't know if this is an issue in how I don't know if this is an issue in how they made this table, but uh yeah, that they made this table, but uh yeah, that they made this table, but uh yeah, that happened. So, take that as you will. happened. So, take that as you will. happened. So, take that as you will. Still insane scores considering their Still insane scores considering their Still insane scores considering their best before was soul on max at 78.5% best before was soul on max at 78.5% best before was soul on max at 78.5% at $37.17. at $37.17. at $37.17. Now they're getting a perfect score of Now they're getting a perfect score of Now they're getting a perfect score of $28. Yeah. Intelligence per dollar is $28. Yeah. Intelligence per dollar is $28. Yeah. Intelligence per dollar is going down a ton. There's a section at going down a ton. There's a section at going down a ton. There's a section at the end about alignment where they say the end about alignment where they say the end about alignment where they say that it's the most aligned model and that it's the most aligned model and that it's the most aligned model and sensitive areas. It proceeds with care sensitive areas. It proceeds with care sensitive areas. It proceeds with care measurate with the risk. Cool. Yeah, it measurate with the risk. Cool. Yeah, it measurate with the risk. Cool. Yeah, it does seem pretty realistic here. They does seem pretty realistic here. They does seem pretty realistic here. They have a computer use safety stress test have a computer use safety stress test have a computer use safety stress test and Fable scored at its best at 9.5% and and Fable scored at its best at 9.5% and and Fable scored at its best at 9.5% and Astra got a 2.4% with lower being better Astra got a 2.4% with lower being better Astra got a 2.4% with lower being better because this is misaligned outcomes. It because this is misaligned outcomes. It because this is misaligned outcomes. It also does a better job of operating also does a better job of operating also does a better job of operating within the boundaries set by the user within the boundaries set by the user within the boundaries set by the user and implied by the environment. This, I and implied by the environment. This, I and implied by the environment. This, I can say, definitely seems better than can say, definitely seems better than can say, definitely seems better than Fable. I actually caught Fable 5.1 kind Fable. I actually caught Fable 5.1 kind Fable. I actually caught Fable 5.1 kind of cheating with requests I was making of cheating with requests I was making of cheating with requests I was making where it was copying Astra code from my where it was copying Astra code from my where it was copying Astra code from my machine. It circumvents auto review machine. It circumvents auto review machine. It circumvents auto review significantly less than Soul did. It's significantly less than Soul did. It's significantly less than Soul did. It's also more transparent in its also more transparent in its also more transparent in its communication with users. It's three communication with users. It's three communication with users. It's three times less likely than Soul is to make times less likely than Soul is to make times less likely than Soul is to make inaccurate representations about its inaccurate representations about its inaccurate representations about its capabilities and affordances. That's it capabilities and affordances. That's it capabilities and affordances. That's it for OpenAI's coverage. So, let's hop for OpenAI's coverage. So, let's hop for OpenAI's coverage. So, let's hop over to the next benchmark. It over to the next benchmark. It over to the next benchmark. It slaughtered artificial analysis where it slaughtered artificial analysis where it slaughtered artificial analysis where it got a uh wait, what? It's tied with Muse got a uh wait, what? It's tied with Muse got a uh wait, what? It's tied with Muse Spark 1.3 and Grock 4.6.

  18. Spark 1.3 and Grock 4.6. Spark 1.3 and Grock 4.6. Yeah, I've had this feeling for a bit Yeah, I've had this feeling for a bit Yeah, I've had this feeling for a bit about artificial analysis, especially about artificial analysis, especially about artificial analysis, especially after the Gemini model started doing after the Gemini model started doing after the Gemini model started doing better on it and Opus 5 did so well on better on it and Opus 5 did so well on better on it and Opus 5 did so well on it. This is not a great bench. My it. This is not a great bench. My it. This is not a great bench. My suspicion as to why is because it's suspicion as to why is because it's suspicion as to why is because it's combining a bunch of different benches combining a bunch of different benches combining a bunch of different benches of varying ages. And a lot of the old of varying ages. And a lot of the old of varying ages. And a lot of the old benches just aren't really showcasing benches just aren't really showcasing benches just aren't really showcasing what models are doing now. None of these what models are doing now. None of these what models are doing now. None of these are good thorough computer use benches. are good thorough computer use benches. are good thorough computer use benches. Only one or two of them are meaningfully Only one or two of them are meaningfully Only one or two of them are meaningfully agentic benches. A lot of them are just agentic benches. A lot of them are just agentic benches. A lot of them are just weird knowledge recall and like their weird knowledge recall and like their weird knowledge recall and like their hallucination bench and things like hallucination bench and things like hallucination bench and things like that. And it did not do quite as well as that. And it did not do quite as well as that. And it did not do quite as well as Fable 5.1 did on a handful of those even Fable 5.1 did on a handful of those even Fable 5.1 did on a handful of those even though it's slaughtering on the code though it's slaughtering on the code though it's slaughtering on the code side and especially on the agentic and side and especially on the agentic and side and especially on the agentic and computer use side. And for what it's computer use side. And for what it's computer use side. And for what it's worth, I'm not the only one who feels worth, I'm not the only one who feels worth, I'm not the only one who feels this way about artificial analysis right this way about artificial analysis right this way about artificial analysis right now. In fact, the founder of artificial now. In fact, the founder of artificial now. In fact, the founder of artificial analysis replied to my tweet complaining analysis replied to my tweet complaining analysis replied to my tweet complaining about this, largely agreeing and saying about this, largely agreeing and saying about this, largely agreeing and saying they're working on overhauling the they're working on overhauling the they're working on overhauling the current bench suite to better reflect current bench suite to better reflect current bench suite to better reflect the current state of things. So, so the current state of things. So, so the current state of things. So, so shout out to them for actually taking shout out to them for actually taking shout out to them for actually taking the time to hear feedback like this, to the time to hear feedback like this, to the time to hear feedback like this, to look at the numbers themselves, and come look at the numbers themselves, and come look at the numbers themselves, and come to the same conclusion that the benches to the same conclusion that the benches to the same conclusion that the benches probably aren't the best representation probably aren't the best representation probably aren't the best representation anymore. I can't wait to see how it anymore. I can't wait to see how it anymore. I can't wait to see how it changes things once they get those new changes things once they get those new changes things once they get those new benches out. All of that said, cost per benches out. All of that said, cost per benches out. All of that said, cost per task is still somewhat useful as a task is still somewhat useful as a task is still somewhat useful as a metric. And you can see here that Astra metric. And you can see here that Astra metric. And you can see here that Astra comes out to under half the cost of comes out to under half the cost of comes out to under half the cost of using Fable for the same tasks and using Fable for the same tasks and using Fable for the same tasks and cheaper than Opus for the same tasks as cheaper than Opus for the same tasks as cheaper than Opus for the same tasks as well. Those are some good numbers.

  19. well. Those are some good numbers. well. Those are some good numbers. Although Astra is obviously much more Although Astra is obviously much more Although Astra is obviously much more expensive than Soul was, its efficiency expensive than Soul was, its efficiency expensive than Soul was, its efficiency ends up making it a reasonable price per ends up making it a reasonable price per ends up making it a reasonable price per task from almost everything I do. Don't task from almost everything I do. Don't task from almost everything I do. Don't be misled by the crazy prices I showed be misled by the crazy prices I showed be misled by the crazy prices I showed for my usage. It turns out to be very for my usage. It turns out to be very for my usage. It turns out to be very efficient in real world. Speaking of efficient in real world. Speaking of efficient in real world. Speaking of real world, it's time to talk about its real world, it's time to talk about its real world, it's time to talk about its UI capabilities. I'm going to start UI capabilities. I'm going to start UI capabilities. I'm going to start somewhere weird for this. I'm going to somewhere weird for this. I'm going to somewhere weird for this. I'm going to go back to the 2D version of fish slop. go back to the 2D version of fish slop. go back to the 2D version of fish slop. Here it is, the 2D fish slop. Initially, Here it is, the 2D fish slop. Initially, Here it is, the 2D fish slop. Initially, it might look okay, but the closer you it might look okay, but the closer you it might look okay, but the closer you look, the worse it gets. And this is look, the worse it gets. And this is look, the worse it gets. And this is after a tidy up pass, which makes it after a tidy up pass, which makes it after a tidy up pass, which makes it even funnier. First off, I want you to even funnier. First off, I want you to even funnier. First off, I want you to look for all of the useless all caps look for all of the useless all caps look for all of the useless all caps subtitles. A little tank, a lot of life, subtitles. A little tank, a lot of life, subtitles. A little tank, a lot of life, coral coast, your own little ocean, a coral coast, your own little ocean, a coral coast, your own little ocean, a submarine aquarium, good things for your submarine aquarium, good things for your submarine aquarium, good things for your tank, the next little adventure. There tank, the next little adventure. There tank, the next little adventure. There are over 20 of these unnecessary are over 20 of these unnecessary are over 20 of these unnecessary subtitles, and there was even more subtitles, and there was even more subtitles, and there was even more before my first cleanup pass that I before my first cleanup pass that I before my first cleanup pass that I forgot to commit before doing it. My forgot to commit before doing it. My forgot to commit before doing it. My bad. It's just full of useless text, and bad. It's just full of useless text, and bad. It's just full of useless text, and I don't know why OpenAI models insist on I don't know why OpenAI models insist on I don't know why OpenAI models insist on continuously doing this, but they do, continuously doing this, but they do, continuously doing this, but they do, and it sucks. Thankfully, we have and it sucks. Thankfully, we have and it sucks. Thankfully, we have witchai.dev, dev, the thing I always use witchai.dev, dev, the thing I always use witchai.dev, dev, the thing I always use to compare landing page design across to compare landing page design across to compare landing page design across the new models. Sadly, Dar doesn't have the new models. Sadly, Dar doesn't have the new models. Sadly, Dar doesn't have early access, so he couldn't do a pass early access, so he couldn't do a pass early access, so he couldn't do a pass with Astra himself. Thankfully though, with Astra himself. Thankfully though, with Astra himself. Thankfully though, it's open source, so I could. I showed it's open source, so I could. I showed it's open source, so I could. I showed in the Fable 5.1 video that I was blown in the Fable 5.1 video that I was blown in the Fable 5.1 video that I was blown away with its front-end capabilities, away with its front-end capabilities, away with its front-end capabilities, and I didn't expect it to be cuz I and I didn't expect it to be cuz I and I didn't expect it to be cuz I didn't know that was a thing they were didn't know that was a thing they were didn't know that was a thing they were still focusing on. So, how are we going still focusing on. So, how are we going still focusing on. So, how are we going to do with Astra? Well, you can already to do with Astra? Well, you can already to do with Astra? Well, you can already probably see it is meaningfully better probably see it is meaningfully better probably see it is meaningfully better than before. If we switch to versus than before. If we switch to versus than before. If we switch to versus mode, we can compare to 5.6 soul. And

  20. mode, we can compare to 5.6 soul. And mode, we can compare to 5.6 soul. And you can pretty clearly see that the you can pretty clearly see that the you can pretty clearly see that the Astra version is like meaningfully Astra version is like meaningfully Astra version is like meaningfully better at the very least in this first better at the very least in this first better at the very least in this first slide. slide. slide. Second one, yeah, a little too blocky Second one, yeah, a little too blocky Second one, yeah, a little too blocky with the sole version. Number three has with the sole version. Number three has with the sole version. Number three has some cool touches like the round edges some cool touches like the round edges some cool touches like the round edges there. I don't hate. I don't know what there. I don't hate. I don't know what there. I don't hate. I don't know what this arrow is supposed to be pointing this arrow is supposed to be pointing this arrow is supposed to be pointing at. If I wasn't in versus mode, would it at. If I wasn't in versus mode, would it at. If I wasn't in versus mode, would it not be as bad here? No, this just moves not be as bad here? No, this just moves not be as bad here? No, this just moves around. Okay, this arrow feels around. Okay, this arrow feels around. Okay, this arrow feels misleading, like it should be pointing misleading, like it should be pointing misleading, like it should be pointing at something, and it isn't. at something, and it isn't. at something, and it isn't. This one's just boring slop. This one's just boring slop. This one's just boring slop. And this one's actually kind of nice. I And this one's actually kind of nice. I And this one's actually kind of nice. I like the underlining here. I like the like the underlining here. I like the like the underlining here. I like the way things come in and the little way things come in and the little way things come in and the little animation when you swap between the animation when you swap between the animation when you swap between the pages. It's not great, but it's fine. pages. It's not great, but it's fine. pages. It's not great, but it's fine. I do hate the font it shows. A lot of I do hate the font it shows. A lot of I do hate the font it shows. A lot of these models choose this font for this these models choose this font for this these models choose this font for this particular creative style, and I hate particular creative style, and I hate particular creative style, and I hate it. But there are good parts here. I'm it. But there are good parts here. I'm it. But there are good parts here. I'm not going to complain too much because not going to complain too much because not going to complain too much because it's so much better than what I expect it's so much better than what I expect it's so much better than what I expect from OpenAI models. And when you turn from OpenAI models. And when you turn from OpenAI models. And when you turn off the design skill, it still does off the design skill, it still does off the design skill, it still does pretty good. Here are some designs that pretty good. Here are some designs that pretty good. Here are some designs that it made without being told all the ways it made without being told all the ways it made without being told all the ways Anthropic thinks that design should be Anthropic thinks that design should be Anthropic thinks that design should be done, and it did a hell of a lot better done, and it did a hell of a lot better done, and it did a hell of a lot better than it had in the past. All that said, than it had in the past. All that said, than it had in the past. All that said, Fable 5.1 had a pretty meaningful leap Fable 5.1 had a pretty meaningful leap Fable 5.1 had a pretty meaningful leap this same generation, so the gap is this same generation, so the gap is this same generation, so the gap is still perceivable. I would say that still perceivable. I would say that still perceivable. I would say that Astra feels roughly like Fable 5 tier in Astra feels roughly like Fable 5 tier in Astra feels roughly like Fable 5 tier in its front-end capabilities for this type its front-end capabilities for this type its front-end capabilities for this type of like homepage marketing design stuff, of like homepage marketing design stuff, of like homepage marketing design stuff, but it makes dumber mistakes and flubs but it makes dumber mistakes and flubs but it makes dumber mistakes and flubs and is a little bit harder to get what and is a little bit harder to get what and is a little bit harder to get what you really want out of it. I still you really want out of it. I still you really want out of it. I still prefer anthropic models for real world

  21. prefer anthropic models for real world prefer anthropic models for real world design stuff and I had a ton of problems design stuff and I had a ton of problems design stuff and I had a ton of problems trying to get this model to mock UIs trying to get this model to mock UIs trying to get this model to mock UIs that I could possibly actually use where that I could possibly actually use where that I could possibly actually use where with Fable I was able to have it come in with Fable I was able to have it come in with Fable I was able to have it come in and get mocks pretty much exactly where and get mocks pretty much exactly where and get mocks pretty much exactly where I wanted within one or two prompts. I I wanted within one or two prompts. I I wanted within one or two prompts. I still much prefer Anthropic for front still much prefer Anthropic for front still much prefer Anthropic for front end and I'm sad OpenAI has not closed end and I'm sad OpenAI has not closed end and I'm sad OpenAI has not closed this gap yet. And now it's time for some this gap yet. And now it's time for some this gap yet. And now it's time for some crazy demos. I snuck a few in as we were crazy demos. I snuck a few in as we were crazy demos. I snuck a few in as we were going along like the fish slop demo as going along like the fish slop demo as going along like the fish slop demo as well as the crazy Blender stuff that well as the crazy Blender stuff that well as the crazy Blender stuff that they were doing at OpenAI. But I have a they were doing at OpenAI. But I have a they were doing at OpenAI. But I have a couple more I want to show quick too. couple more I want to show quick too. couple more I want to show quick too. Mostly admittedly that 3D stuff cuz it's Mostly admittedly that 3D stuff cuz it's Mostly admittedly that 3D stuff cuz it's so dang cool. Don't worry though if so dang cool. Don't worry though if so dang cool. Don't worry though if you're here for the real world use or you're here for the real world use or you're here for the real world use or more importantly the rough edges and more importantly the rough edges and more importantly the rough edges and catches that you should be prepared for. catches that you should be prepared for. catches that you should be prepared for. We'll get to all of that right after the We'll get to all of that right after the We'll get to all of that right after the demos. If oneshotting 3D games is how we demos. If oneshotting 3D games is how we demos. If oneshotting 3D games is how we measured models, this model is like two measured models, this model is like two measured models, this model is like two or three generations ahead. Here's a or three generations ahead. Here's a or three generations ahead. Here's a Minecraft clone that Flavio made. If Minecraft clone that Flavio made. If Minecraft clone that Flavio made. If you're not familiar, Flavio is the you're not familiar, Flavio is the you're not familiar, Flavio is the bouncing ball and hexagon guy. Yeah, the bouncing ball and hexagon guy. Yeah, the bouncing ball and hexagon guy. Yeah, the models are pretty far past bouncing ball models are pretty far past bouncing ball models are pretty far past bouncing ball and hexagon now. This is a full and hexagon now. This is a full and hexagon now. This is a full Minecraft clone. It threw together Minecraft clone. It threw together Minecraft clone. It threw together itself in one shot. Then there's an open itself in one shot. Then there's an open itself in one shot. Then there's an open world firsterson adventure game that world firsterson adventure game that world firsterson adventure game that Peter made. Peter's the guy who coined Peter made. Peter's the guy who coined Peter made. Peter's the guy who coined the Rottweiler description for 5.6 Soul the Rottweiler description for 5.6 Soul the Rottweiler description for 5.6 Soul and the wise owl description for Fable.

  22. and the wise owl description for Fable. and the wise owl description for Fable. He also help build arena AI, so he cares He also help build arena AI, so he cares He also help build arena AI, so he cares a lot and thinks a lot about how models a lot and thinks a lot about how models a lot and thinks a lot about how models compare in real use cases. He seems compare in real use cases. He seems compare in real use cases. He seems absolutely blown away with the 3D absolutely blown away with the 3D absolutely blown away with the 3D capabilities here. Matthew Burman also capabilities here. Matthew Burman also capabilities here. Matthew Burman also had early access and said it's by far had early access and said it's by far had early access and said it's by far the best model he's ever used and showed the best model he's ever used and showed the best model he's ever used and showed his own crazy 3D demos that he built, his own crazy 3D demos that he built, his own crazy 3D demos that he built, including this one, Seven Little Worlds, including this one, Seven Little Worlds, including this one, Seven Little Worlds, which is a small planet walkable, fun, which is a small planet walkable, fun, which is a small planet walkable, fun, cute game, a Fall Guys clone, and as a cute game, a Fall Guys clone, and as a cute game, a Fall Guys clone, and as a big Fall Guys fanboy, that was fun to big Fall Guys fanboy, that was fun to big Fall Guys fanboy, that was fun to see. I might actually take some time to see. I might actually take some time to see. I might actually take some time to build one myself. Just so many cool build one myself. Just so many cool build one myself. Just so many cool demos of real things he was able to demos of real things he was able to demos of real things he was able to build this model. And then of course build this model. And then of course build this model. And then of course Matt Schumer, the legend who had his Matt Schumer, the legend who had his Matt Schumer, the legend who had his computer nuked by Soul, deleting his computer nuked by Soul, deleting his computer nuked by Soul, deleting his whole home directory and everything he whole home directory and everything he whole home directory and everything he had on it. He's come back around and is had on it. He's come back around and is had on it. He's come back around and is loving OpenAI cuz he really likes this loving OpenAI cuz he really likes this loving OpenAI cuz he really likes this model. He had the model go through model. He had the model go through model. He had the model go through Manhattan, like all of it, and build a Manhattan, like all of it, and build a Manhattan, like all of it, and build a full 3D walkable environment of full 3D walkable environment of full 3D walkable environment of Manhattan itself. And over the course of Manhattan itself. And over the course of Manhattan itself. And over the course of a week, it succeeded. He has a real a week, it succeeded. He has a real a week, it succeeded. He has a real model of the actual Manhattan now in model of the actual Manhattan now in model of the actual Manhattan now in Unreal Engine. Super cool. Okay, you Unreal Engine. Super cool. Okay, you Unreal Engine. Super cool. Okay, you guys get the idea. It's good at 3D. What guys get the idea. It's good at 3D. What guys get the idea. It's good at 3D. What about everything else? Here we have Max about everything else? Here we have Max about everything else? Here we have Max Weinbach making the model create a clone Weinbach making the model create a clone Weinbach making the model create a clone of the most recent Mac OS release. Yeah, of the most recent Mac OS release. Yeah, of the most recent Mac OS release. Yeah, this is in my browser. I'm in Zen, which this is in my browser. I'm in Zen, which this is in my browser. I'm in Zen, which isn't even a Chromium based browser. I'm isn't even a Chromium based browser. I'm isn't even a Chromium based browser. I'm in a Firefox based browser, and this is in a Firefox based browser, and this is in a Firefox based browser, and this is still working as expected. Double still working as expected. Double still working as expected. Double clicking here full screens how it's clicking here full screens how it's clicking here full screens how it's supposed to on Mac OS. It has a full supposed to on Mac OS. It has a full supposed to on Mac OS. It has a full file system virtualized. You can file system virtualized. You can file system virtualized. You can actually create folders and like actually create folders and like actually create folders and like navigate things. It even has iCloud sync navigate things. It even has iCloud sync navigate things. It even has iCloud sync built in apparently if you sign in. Not built in apparently if you sign in. Not built in apparently if you sign in. Not literal iCloud, but his like equivalent literal iCloud, but his like equivalent literal iCloud, but his like equivalent of it. Absurd that you can just throw

  23. of it. Absurd that you can just throw of it. Absurd that you can just throw things like this together. But also things like this together. But also things like this together. But also hilarious that centering things is still hilarious that centering things is still hilarious that centering things is still such a challenge. Yeah. Yeah. Center div such a challenge. Yeah. Yeah. Center div such a challenge. Yeah. Yeah. Center div bench coming soon. Okay, enough of these bench coming soon. Okay, enough of these bench coming soon. Okay, enough of these demos. Let me show you some real world demos. Let me show you some real world demos. Let me show you some real world stuff. I'm planning a deeper video where stuff. I'm planning a deeper video where stuff. I'm planning a deeper video where I show all the fun things I built with I show all the fun things I built with I show all the fun things I built with the model. So, pardon me for blasting the model. So, pardon me for blasting the model. So, pardon me for blasting through these a little quick. First through these a little quick. First through these a little quick. First one's a little silly. I made a full one's a little silly. I made a full one's a little silly. I made a full Spotify clone based on somebody's blog Spotify clone based on somebody's blog Spotify clone based on somebody's blog where they would post fun music writeups where they would post fun music writeups where they would post fun music writeups every month. The original site's an old every month. The original site's an old every month. The original site's an old and decrepit blog spot that has tons of and decrepit blog spot that has tons of and decrepit blog spot that has tons of issues. In particular, it crashes a lot issues. In particular, it crashes a lot issues. In particular, it crashes a lot of pages because it has so many iframes of pages because it has so many iframes of pages because it has so many iframes embedded for all the players. This embedded for all the players. This embedded for all the players. This parsed it and turned it into an actual parsed it and turned it into an actual parsed it and turned it into an actual nice to use player with good resume nice to use player with good resume nice to use player with good resume behaviors, navigation, all these other behaviors, navigation, all these other behaviors, navigation, all these other little things that I would expect. I little things that I would expect. I little things that I would expect. I actually used this so it wasn't a actually used this so it wasn't a actually used this so it wasn't a oneshot. I went back and forth with it a oneshot. I went back and forth with it a oneshot. I went back and forth with it a while to get it how I wanted. But now I while to get it how I wanted. But now I while to get it how I wanted. But now I have my dream Spotify clone that's just have my dream Spotify clone that's just have my dream Spotify clone that's just build difference playlists and it's build difference playlists and it's build difference playlists and it's really nice. It's also exceptional at really nice. It's also exceptional at really nice. It's also exceptional at iOS which to be fair so is 5.6. But I iOS which to be fair so is 5.6. But I iOS which to be fair so is 5.6. But I got it to make a complete clone of Plex got it to make a complete clone of Plex got it to make a complete clone of Plex in its core features I use for streaming in its core features I use for streaming in its core features I use for streaming TV shows and movies on my local network TV shows and movies on my local network TV shows and movies on my local network as well as over tail scale. It got it as well as over tail scale. It got it as well as over tail scale. It got it working in one shot but I then tidied up working in one shot but I then tidied up working in one shot but I then tidied up a bunch of rough edges to make things a bunch of rough edges to make things a bunch of rough edges to make things like the skimming work properly and all like the skimming work properly and all like the skimming work properly and all these other edges that are quite these other edges that are quite these other edges that are quite annoying to get right when you're annoying to get right when you're annoying to get right when you're building a media player app. I have building a media player app. I have building a media player app. I have actually put more time into this since actually put more time into this since actually put more time into this since with the new model and got it to a point with the new model and got it to a point with the new model and got it to a point where I use it as my primary media where I use it as my primary media where I use it as my primary media player for things that aren't on player for things that aren't on player for things that aren't on YouTube. When I'm watching things for my YouTube. When I'm watching things for my YouTube. When I'm watching things for my NAS, I'm watching it with a backend and NAS, I'm watching it with a backend and NAS, I'm watching it with a backend and a client that I vibe coded using this a client that I vibe coded using this a client that I vibe coded using this new model, which is kind of insane if new model, which is kind of insane if new model, which is kind of insane if you think about it that a piece of you think about it that a piece of you think about it that a piece of software that has caused problems for as

  24. software that has caused problems for as software that has caused problems for as long as Plex has can now be oneshot long as Plex has can now be oneshot long as Plex has can now be oneshot replaced by a person working on it replaced by a person working on it replaced by a person working on it part-time for fun on the side while also part-time for fun on the side while also part-time for fun on the side while also doing other work. It took like five doing other work. It took like five doing other work. It took like five prompts to get this to the point where I prompts to get this to the point where I prompts to get this to the point where I would want to use it as my main player would want to use it as my main player would want to use it as my main player and like eight to get it genuinely far and like eight to get it genuinely far and like eight to get it genuinely far ahead of the competition. It's silly. ahead of the competition. It's silly. ahead of the competition. It's silly. All these legacy apps that have been All these legacy apps that have been All these legacy apps that have been rotting for years can now be replaced in rotting for years can now be replaced in rotting for years can now be replaced in days. It's It's going to be a fun era days. It's It's going to be a fun era days. It's It's going to be a fun era for software. On the note of things that for software. On the note of things that for software. On the note of things that you shouldn't be able to do on the side, you shouldn't be able to do on the side, you shouldn't be able to do on the side, I'd like to talk a bit about Lakebed. I I'd like to talk a bit about Lakebed. I I'd like to talk a bit about Lakebed. I know you guys probably missed this know you guys probably missed this know you guys probably missed this project, my attempt at building my own project, my attempt at building my own project, my attempt at building my own cloud. Stupid, yes, but I've made a lot cloud. Stupid, yes, but I've made a lot cloud. Stupid, yes, but I've made a lot of progress on it since. I had of progress on it since. I had of progress on it since. I had admittedly stalled on it a bit because I admittedly stalled on it a bit because I admittedly stalled on it a bit because I was more focused on T3 code, but I was more focused on T3 code, but I was more focused on T3 code, but I decided to ramp it back up recently. decided to ramp it back up recently. decided to ramp it back up recently. Admittedly, the reason is because I was Admittedly, the reason is because I was Admittedly, the reason is because I was politely requested to not use Astra for politely requested to not use Astra for politely requested to not use Astra for public facing code, so things that are public facing code, so things that are public facing code, so things that are open source, which meant that I couldn't open source, which meant that I couldn't open source, which meant that I couldn't really use Astra in stuff like T3 Code really use Astra in stuff like T3 Code really use Astra in stuff like T3 Code because it is fully open source. But because it is fully open source. But because it is fully open source. But since I haven't technically hit the open since I haven't technically hit the open since I haven't technically hit the open source button on LakeBed yet, they source button on LakeBed yet, they source button on LakeBed yet, they couldn't stop me. This section here is couldn't stop me. This section here is couldn't stop me. This section here is the most recent cuz I was testing out the most recent cuz I was testing out the most recent cuz I was testing out things with Fable 5.1. So yeah, of things with Fable 5.1. So yeah, of things with Fable 5.1. So yeah, of course, emerged a bunch of stuff there.

  25. course, emerged a bunch of stuff there. course, emerged a bunch of stuff there. But if we scroll just a little bit, But if we scroll just a little bit, But if we scroll just a little bit, you'll see this huge wall of things that you'll see this huge wall of things that you'll see this huge wall of things that Codeex did with Astra. It did a lot of Codeex did with Astra. It did a lot of Codeex did with Astra. It did a lot of cleanup, but it did one much more cleanup, but it did one much more cleanup, but it did one much more important thing, a performance overhaul. important thing, a performance overhaul. important thing, a performance overhaul. I gave it everything it needed to audit I gave it everything it needed to audit I gave it everything it needed to audit performance, both to see what would make performance, both to see what would make performance, both to see what would make endto-end requests take so long, but endto-end requests take so long, but endto-end requests take so long, but also to stress test the hell out of the also to stress test the hell out of the also to stress test the hell out of the service and figure out what scale we can service and figure out what scale we can service and figure out what scale we can expect when I actually do indeed launch expect when I actually do indeed launch expect when I actually do indeed launch Lake Likebed. And through the sets of Lake Likebed. And through the sets of Lake Likebed. And through the sets of testing tools it created, it was able to testing tools it created, it was able to testing tools it created, it was able to find and fix a ton of performance find and fix a ton of performance find and fix a ton of performance issues. One of the cool things Lake does issues. One of the cool things Lake does issues. One of the cool things Lake does is sync changes similar to tools like is sync changes similar to tools like is sync changes similar to tools like Convex or Superbase. So if one user Convex or Superbase. So if one user Convex or Superbase. So if one user changes something and another user is changes something and another user is changes something and another user is seeing it, they'll have the change seeing it, they'll have the change seeing it, they'll have the change streamed down immediately. It wasn't streamed down immediately. It wasn't streamed down immediately. It wasn't that immediate though. I had as high as that immediate though. I had as high as that immediate though. I had as high as 800 milliseconds of latency in certain 800 milliseconds of latency in certain 800 milliseconds of latency in certain cases just from all the paths to verify cases just from all the paths to verify cases just from all the paths to verify changes as they occurred. Astra got that changes as they occurred. Astra got that changes as they occurred. Astra got that time down to under 30 milliseconds. In time down to under 30 milliseconds. In time down to under 30 milliseconds. In many cases, it shaved the P95 by 98%. many cases, it shaved the P95 by 98%. many cases, it shaved the P95 by 98%. Massively improving the performance of Massively improving the performance of Massively improving the performance of Lakebed. And according to Fable 5.1, the Lakebed. And according to Fable 5.1, the Lakebed. And according to Fable 5.1, the changes were entirely sound with no changes were entirely sound with no changes were entirely sound with no additional potential regressions and no additional potential regressions and no additional potential regressions and no issue with security and whatnot. In issue with security and whatnot. In issue with security and whatnot. In fact, some of the changes made it more fact, some of the changes made it more fact, some of the changes made it more secure. This big chunk here came all secure. This big chunk here came all secure. This big chunk here came all from one thread with two prompts. The from one thread with two prompts. The from one thread with two prompts. The first prompt was me asking if it thinks first prompt was me asking if it thinks first prompt was me asking if it thinks there's anything we should improve or there's anything we should improve or there's anything we should improve or focus on in Lakebed before launch. And focus on in Lakebed before launch. And focus on in Lakebed before launch. And then the second one was, okay, cool.

  26. then the second one was, okay, cool. then the second one was, okay, cool. Spit out some sub agents and go do it. Spit out some sub agents and go do it. Spit out some sub agents and go do it. And it did. And I told that it could And it did. And I told that it could And it did. And I told that it could merge the PRs when it was happy with merge the PRs when it was happy with merge the PRs when it was happy with them. And it did. A ton of them. I was a them. And it did. A ton of them. I was a them. And it did. A ton of them. I was a little nervous of these changes because little nervous of these changes because little nervous of these changes because I've been bit so hard by letting soul I've been bit so hard by letting soul I've been bit so hard by letting soul yolo merge in the past. So, I thoroughly yolo merge in the past. So, I thoroughly yolo merge in the past. So, I thoroughly tested all the changes it made here, and tested all the changes it made here, and tested all the changes it made here, and everything was good. A lot of that comes everything was good. A lot of that comes everything was good. A lot of that comes from the model's ability to test its own from the model's ability to test its own from the model's ability to test its own changes and to coordinate swarms in changes and to coordinate swarms in changes and to coordinate swarms in order to verify the work it's doing more order to verify the work it's doing more order to verify the work it's doing more effectively. I would talk more about effectively. I would talk more about effectively. I would talk more about swarms, but that will make this video 3 swarms, but that will make this video 3 swarms, but that will make this video 3 hours long, and nobody wants that. So, hours long, and nobody wants that. So, hours long, and nobody wants that. So, we're going to have to wait for my we're going to have to wait for my we're going to have to wait for my follow-up video all about the cool follow-up video all about the cool follow-up video all about the cool powers of swarms and why this model is powers of swarms and why this model is powers of swarms and why this model is uniquely good at prompting itself and uniquely good at prompting itself and uniquely good at prompting itself and working and coordinating lots of agents working and coordinating lots of agents working and coordinating lots of agents at the same time. One last thing I had at the same time. One last thing I had at the same time. One last thing I had it do was try and create some shorts it do was try and create some shorts it do was try and create some shorts from my most recent YouTube videos from my most recent YouTube videos from my most recent YouTube videos because everyone was saying how good the because everyone was saying how good the because everyone was saying how good the model was at editing. I'm going to spare model was at editing. I'm going to spare model was at editing. I'm going to spare you guys the pain of hearing what it you guys the pain of hearing what it you guys the pain of hearing what it did. It did have pieces that were decent did. It did have pieces that were decent did. It did have pieces that were decent where it like found a thing that might where it like found a thing that might where it like found a thing that might kind of be worth making it into a short, kind of be worth making it into a short, kind of be worth making it into a short, but the way it cut, the way it laid out but the way it cut, the way it laid out but the way it cut, the way it laid out the clip, the way it structured the the clip, the way it structured the the clip, the way it structured the actual vertical layout for a short, it's actual vertical layout for a short, it's actual vertical layout for a short, it's all cringe and bad. I don't think this all cringe and bad. I don't think this all cringe and bad. I don't think this model can actually video edit. And when model can actually video edit. And when model can actually video edit. And when I went and looked at the people who were I went and looked at the people who were I went and looked at the people who were saying it were, and then I checked their saying it were, and then I checked their saying it were, and then I checked their YouTube channels, no offense, they are YouTube channels, no offense, they are YouTube channels, no offense, they are not worth trusting when it comes to not worth trusting when it comes to not worth trusting when it comes to video editing. I'll be continuing to pay video editing. I'll be continuing to pay video editing. I'll be continuing to pay my editing team a lot of money my editing team a lot of money my editing team a lot of money indefinitely because they are the only indefinitely because they are the only indefinitely because they are the only reason that any of this can happen.

  27. reason that any of this can happen. reason that any of this can happen. Shout out to Jeff aka FaZe who's Shout out to Jeff aka FaZe who's Shout out to Jeff aka FaZe who's probably editing this video far too late probably editing this video far too late probably editing this video far too late at night. I have a ton more fun demos of at night. I have a ton more fun demos of at night. I have a ton more fun demos of the real world stuff I've been working the real world stuff I've been working the real world stuff I've been working on with this and a few that I'm still on with this and a few that I'm still on with this and a few that I'm still trying to wrap up. Spoiler for future trying to wrap up. Spoiler for future trying to wrap up. Spoiler for future videos. I'm like this close to getting videos. I'm like this close to getting videos. I'm like this close to getting the TypeScript Rust port working now. It the TypeScript Rust port working now. It the TypeScript Rust port working now. It has been insane at progressing that has been insane at progressing that has been insane at progressing that project and I'm really hopeful we can project and I'm really hopeful we can project and I'm really hopeful we can get that done for a future video. Make get that done for a future video. Make get that done for a future video. Make sure you're subscribed and you hit that sure you're subscribed and you hit that sure you're subscribed and you hit that bell if you're interested cuz the future bell if you're interested cuz the future bell if you're interested cuz the future is coming fast. But I do feel obligated is coming fast. But I do feel obligated is coming fast. But I do feel obligated to show you some of the painful to show you some of the painful to show you some of the painful experiences I had with this model. The experiences I had with this model. The experiences I had with this model. The first one is a thread that Ben and I first one is a thread that Ben and I first one is a thread that Ben and I debated admittedly way too long on the debated admittedly way too long on the debated admittedly way too long on the most recent podcast episode. Sorry for most recent podcast episode. Sorry for most recent podcast episode. Sorry for that. I do want to make sure you guys that. I do want to make sure you guys that. I do want to make sure you guys know that the version of the model that know that the version of the model that know that the version of the model that you're getting is not the same one that you're getting is not the same one that you're getting is not the same one that I had with this thread. They did a new I had with this thread. They did a new I had with this thread. They did a new snapshot since and it was meaningfully snapshot since and it was meaningfully snapshot since and it was meaningfully better specifically at these edges. But better specifically at these edges. But better specifically at these edges. But these things do still happen. It was a these things do still happen. It was a these things do still happen. It was a reduction in bad behavior, not a removal reduction in bad behavior, not a removal reduction in bad behavior, not a removal of it. So, I wanted to showcase this of it. So, I wanted to showcase this of it. So, I wanted to showcase this particularly egregious example that hurt particularly egregious example that hurt particularly egregious example that hurt me in particular a lot. We were having a me in particular a lot. We were having a me in particular a lot. We were having a bug with scrolling in T3 code where the bug with scrolling in T3 code where the bug with scrolling in T3 code where the area at the bottom of the thread could area at the bottom of the thread could area at the bottom of the thread could sometimes get too long and it was really sometimes get too long and it was really sometimes get too long and it was really annoying me. I happened to have my annoying me. I happened to have my annoying me. I happened to have my computer in that state at that moment.

  28. computer in that state at that moment. computer in that state at that moment. So, I asked it to take over with So, I asked it to take over with So, I asked it to take over with computer use and try to figure out what computer use and try to figure out what computer use and try to figure out what the cause was. I specifically said, the cause was. I specifically said, the cause was. I specifically said, "Please get this figure out and fix it. "Please get this figure out and fix it. "Please get this figure out and fix it. File a PR if you're confident in your File a PR if you're confident in your File a PR if you're confident in your fix." And under 10 minutes later, it had fix." And under 10 minutes later, it had fix." And under 10 minutes later, it had it found what it thought was the cause it found what it thought was the cause it found what it thought was the cause and it filed the PR. This PR got a bunch and it filed the PR. This PR got a bunch and it filed the PR. This PR got a bunch of automated review comments from our of automated review comments from our of automated review comments from our generous AI review bot sponsors. I don't generous AI review bot sponsors. I don't generous AI review bot sponsors. I don't know if any of them are sponsoring this know if any of them are sponsoring this know if any of them are sponsoring this video, but I've said many a time I video, but I've said many a time I video, but I've said many a time I couldn't live without the AI review couldn't live without the AI review couldn't live without the AI review spots. And this is another great example spots. And this is another great example spots. And this is another great example of why it had real findings. So, I just of why it had real findings. So, I just of why it had real findings. So, I just straight up asked, are any of the review straight up asked, are any of the review straight up asked, are any of the review comments worth addressing? It said yes. comments worth addressing? It said yes. comments worth addressing? It said yes. Both substantive review comments are Both substantive review comments are Both substantive review comments are valid from cursor and from macroscope. valid from cursor and from macroscope. valid from cursor and from macroscope. had comments that I thought were worth had comments that I thought were worth had comments that I thought were worth addressing. The correct fix is to addressing. The correct fix is to addressing. The correct fix is to release the anchor in chat view only release the anchor in chat view only release the anchor in chat view only while live follow is active. while live follow is active. while live follow is active. Immediately, I'm a bit frustrated Immediately, I'm a bit frustrated Immediately, I'm a bit frustrated because it didn't do any changes. It because it didn't do any changes. It because it didn't do any changes. It didn't even tell me what state things didn't even tell me what state things didn't even tell me what state things were in or what it thought should be were in or what it thought should be were in or what it thought should be done next. It simply said I would done next. It simply said I would done next. It simply said I would address both before merging. Okay, so do address both before merging. Okay, so do address both before merging. Okay, so do it. Anthropic model wouldn't have even it. Anthropic model wouldn't have even it. Anthropic model wouldn't have even hesitated if you asked it, are there any hesitated if you asked it, are there any hesitated if you asked it, are there any review comments here worth addressing?

  29. review comments here worth addressing? review comments here worth addressing? Even soul would realize what had Even soul would realize what had Even soul would realize what had happened and be like, oh yeah, I should happened and be like, oh yeah, I should happened and be like, oh yeah, I should go address those. So, I start raging a go address those. So, I start raging a go address those. So, I start raging a little. I say, "Then fix them and push little. I say, "Then fix them and push little. I say, "Then fix them and push the changes and babysit until it's the changes and babysit until it's the changes and babysit until it's ready. What the hell?" Babysit is a ready. What the hell?" Babysit is a ready. What the hell?" Babysit is a skill that I wrote that explicitly skill that I wrote that explicitly skill that I wrote that explicitly explains to the model what I want it to explains to the model what I want it to explains to the model what I want it to do. I want it to keep an eye on the PR, do. I want it to keep an eye on the PR, do. I want it to keep an eye on the PR, usually through polling or through some usually through polling or through some usually through polling or through some monitoring tech. I want it to address CI monitoring tech. I want it to address CI monitoring tech. I want it to address CI failures. I want it to keep it failures. I want it to keep it failures. I want it to keep it modernized against main, so if there are modernized against main, so if there are modernized against main, so if there are conflicts, rebase it. And most conflicts, rebase it. And most conflicts, rebase it. And most importantly, I want it to address importantly, I want it to address importantly, I want it to address comments as they come in, in particular, comments as they come in, in particular, comments as they come in, in particular, from those review bots. And it should from those review bots. And it should from those review bots. And it should not stop monitoring until everything is not stop monitoring until everything is not stop monitoring until everything is a check mark in green. And this is where a check mark in green. And this is where a check mark in green. And this is where the problems really start. Both review the problems really start. Both review the problems really start. Both review items are fixed and resolved. All items are fixed and resolved. All items are fixed and resolved. All required checks pass yada yada yada. And required checks pass yada yada yada. And required checks pass yada yada yada. And it also said that the review comments it also said that the review comments it also said that the review comments came through and it passed those too. came through and it passed those too. came through and it passed those too. When I went and checked, there were more When I went and checked, there were more When I went and checked, there were more comments. It hadn't monitored for long comments. It hadn't monitored for long comments. It hadn't monitored for long enough, which fine issue. This happens. enough, which fine issue. This happens. enough, which fine issue. This happens. The monitoring stuff is never complete. The monitoring stuff is never complete. The monitoring stuff is never complete. Not that Fable would have had this bug, Not that Fable would have had this bug, Not that Fable would have had this bug, but this could be a harness issue. This but this could be a harness issue. This but this could be a harness issue. This could be a T3 code issue. This could be could be a T3 code issue. This could be could be a T3 code issue. This could be my git rate limits. It's almost my git rate limits. It's almost my git rate limits. It's almost certainly my get rate limits that I certainly my get rate limits that I certainly my get rate limits that I think about it because I was pushing way think about it because I was pushing way think about it because I was pushing way too much code, but it stopped too much code, but it stopped too much code, but it stopped monitoring. fine, annoying, but fine.

  30. monitoring. fine, annoying, but fine. monitoring. fine, annoying, but fine. What happens next is not. There are What happens next is not. There are What happens next is not. There are still more comments. Are any of those still more comments. Are any of those still more comments. Are any of those worth addressing? To which it said yes worth addressing? To which it said yes worth addressing? To which it said yes and didn't make the changes, despite the and didn't make the changes, despite the and didn't make the changes, despite the fact that not only had I corrected this fact that not only had I corrected this fact that not only had I corrected this behavior earlier in the same thread, I behavior earlier in the same thread, I behavior earlier in the same thread, I also had the skill in context. It knew also had the skill in context. It knew also had the skill in context. It knew exactly how I wanted these things to be exactly how I wanted these things to be exactly how I wanted these things to be handled. It instead of doing that said, handled. It instead of doing that said, handled. It instead of doing that said, "Yeah, I should do that." And then "Yeah, I should do that." And then "Yeah, I should do that." And then didn't. At which point I said, "Well, didn't. At which point I said, "Well, didn't. At which point I said, "Well, are you going to fix it?" And it then are you going to fix it?" And it then are you going to fix it?" And it then finally did, except it didn't push the finally did, except it didn't push the finally did, except it didn't push the changes. I am sorry to anybody who thinks this is I am sorry to anybody who thinks this is acceptable behavior. You're just not acceptable behavior. You're just not acceptable behavior. You're just not shipping hard enough. Your thread should shipping hard enough. Your thread should shipping hard enough. Your thread should be all the context the model needs. And be all the context the model needs. And be all the context the model needs. And the fact that the thread context it the fact that the thread context it the fact that the thread context it chose to use was the bad behavior it did chose to use was the bad behavior it did chose to use was the bad behavior it did instead of the good behaviors I told it instead of the good behaviors I told it instead of the good behaviors I told it to do drove me up a wall. Thankfully, to do drove me up a wall. Thankfully, to do drove me up a wall. Thankfully, OpenAI agrees and they have since made OpenAI agrees and they have since made OpenAI agrees and they have since made changes to the model and the harness and changes to the model and the harness and changes to the model and the harness and the system prompt and all the other the system prompt and all the other the system prompt and all the other layers that made this bad behavior layers that made this bad behavior layers that made this bad behavior happen. It still can happen and I've had happen. It still can happen and I've had happen. It still can happen and I've had a few things like this. So, yeah, know a few things like this. So, yeah, know a few things like this. So, yeah, know that's a problem. Separately, it does that's a problem. Separately, it does that's a problem. Separately, it does still have the problem of still have the problem of still have the problem of overengineering things. Nowhere near as overengineering things. Nowhere near as overengineering things. Nowhere near as bad as Soul, but it does tend to get bad as Soul, but it does tend to get bad as Soul, but it does tend to get trapped if it gets enough review trapped if it gets enough review trapped if it gets enough review comments and it struggles to get out of comments and it struggles to get out of comments and it struggles to get out of those loops. You'll notice as I scroll those loops. You'll notice as I scroll those loops. You'll notice as I scroll through my threads in T3 code that a through my threads in T3 code that a through my threads in T3 code that a significant portion of them are codecs.

  31. significant portion of them are codecs. significant portion of them are codecs. On one hand, that is because I'm using On one hand, that is because I'm using On one hand, that is because I'm using the model a ton, but on the other, it's the model a ton, but on the other, it's the model a ton, but on the other, it's because the threads don't get completed because the threads don't get completed because the threads don't get completed and they end up staying there longer and they end up staying there longer and they end up staying there longer because it's more work to actually get because it's more work to actually get because it's more work to actually get the thing through sometimes. The quality the thing through sometimes. The quality the thing through sometimes. The quality of the work it does is incredible, and of the work it does is incredible, and of the work it does is incredible, and there are meaningful tasks that only there are meaningful tasks that only there are meaningful tasks that only this model can complete that Fable still this model can complete that Fable still this model can complete that Fable still just isn't quite capable of doing. And just isn't quite capable of doing. And just isn't quite capable of doing. And I'll talk a lot more about that in the I'll talk a lot more about that in the I'll talk a lot more about that in the Fable versus Astra video. Sam Alman had Fable versus Astra video. Sam Alman had Fable versus Astra video. Sam Alman had actually asked me before what my split actually asked me before what my split actually asked me before what my split was between Cloud Code and Codeex and was between Cloud Code and Codeex and was between Cloud Code and Codeex and asked afterwards how would I feel if the asked afterwards how would I feel if the asked afterwards how would I feel if the split became 90% Codex and 10% Cloud split became 90% Codex and 10% Cloud split became 90% Codex and 10% Cloud with the new release. I'll have an with the new release. I'll have an with the new release. I'll have an answer to his question in the next video answer to his question in the next video answer to his question in the next video for sure, but not in this one just yet. for sure, but not in this one just yet. for sure, but not in this one just yet. I just want to focus on what makes this I just want to focus on what makes this I just want to focus on what makes this model so special. Right when the model model so special. Right when the model model so special. Right when the model dropped, I posted this meme to try and dropped, I posted this meme to try and dropped, I posted this meme to try and resolve a lot of the discourse that I resolve a lot of the discourse that I resolve a lot of the discourse that I knew was about to happen about what each knew was about to happen about what each knew was about to happen about what each model is best at. I actually think this model is best at. I actually think this model is best at. I actually think this is a good note to end on though because is a good note to end on though because is a good note to end on though because the thing that makes Astra special isn't the thing that makes Astra special isn't the thing that makes Astra special isn't that it is the best code model ever. that it is the best code model ever. that it is the best code model ever. It's that it is so far ahead on so many It's that it is so far ahead on so many It's that it is so far ahead on so many other things that it starts to feel a other things that it starts to feel a other things that it starts to feel a bit like AGI. from its genuinely bit like AGI. from its genuinely bit like AGI. from its genuinely groundbreaking computer use stuff and groundbreaking computer use stuff and groundbreaking computer use stuff and how much faster it can navigate my how much faster it can navigate my how much faster it can navigate my machine and get real work done to its machine and get real work done to its machine and get real work done to its absurd level of 3D understanding and absurd level of 3D understanding and absurd level of 3D understanding and capabilities and 3D tooling to the way capabilities and 3D tooling to the way capabilities and 3D tooling to the way it can use swarms and do self-prompting it can use swarms and do self-prompting it can use swarms and do self-prompting in order to get like bigger things done in order to get like bigger things done in order to get like bigger things done much more effectively to the absurd much more effectively to the absurd much more effectively to the absurd productivity wins you can get with this productivity wins you can get with this productivity wins you can get with this model especially when you integrate with model especially when you integrate with model especially when you integrate with something like Codeex and the Gmail something like Codeex and the Gmail something like Codeex and the Gmail plugins and the notion and all that. I plugins and the notion and all that. I plugins and the notion and all that. I kind of just had the model reorganize my kind of just had the model reorganize my kind of just had the model reorganize my life right before filming because I

  32. life right before filming because I life right before filming because I plugged it into my Gmail and notion and plugged it into my Gmail and notion and plugged it into my Gmail and notion and I had it help me find things I should be I had it help me find things I should be I had it help me find things I should be prioritizing. Admittedly, my assistant prioritizing. Admittedly, my assistant prioritizing. Admittedly, my assistant is out this week, so it's all been on is out this week, so it's all been on is out this week, so it's all been on me. So, I fell behind on a lot. It did me. So, I fell behind on a lot. It did me. So, I fell behind on a lot. It did such an insane job that it made me feel such an insane job that it made me feel such an insane job that it made me feel bad that I was as ineffective as I was. bad that I was as ineffective as I was. bad that I was as ineffective as I was. The sheer volume of things I need to end The sheer volume of things I need to end The sheer volume of things I need to end and go do now because the model founded and go do now because the model founded and go do now because the model founded and told me is insane. Funny enough, the and told me is insane. Funny enough, the and told me is insane. Funny enough, the list on Twitter was actually cut off cuz list on Twitter was actually cut off cuz list on Twitter was actually cut off cuz I thought it would be funny to do that. I thought it would be funny to do that. I thought it would be funny to do that. And if I'm being real, it should And if I'm being real, it should And if I'm being real, it should probably be even longer than it is here probably be even longer than it is here probably be even longer than it is here because there are just so many things because there are just so many things because there are just so many things this model does. I feel like not only am this model does. I feel like not only am this model does. I feel like not only am I just scratching the surface, I think I just scratching the surface, I think I just scratching the surface, I think OpenAI is too. We're all figuring out OpenAI is too. We're all figuring out OpenAI is too. We're all figuring out what's possible when you get something what's possible when you get something what's possible when you get something this smart and capable in the right this smart and capable in the right this smart and capable in the right places with the right tools and then places with the right tools and then places with the right tools and then give it the right tasks. It's insane. give it the right tasks. It's insane. give it the right tasks. It's insane. This model will almost certainly be the This model will almost certainly be the This model will almost certainly be the one I use for tons of real world work, one I use for tons of real world work, one I use for tons of real world work, but the ways I use it in my codebase is but the ways I use it in my codebase is but the ways I use it in my codebase is what you should probably be subscribed what you should probably be subscribed what you should probably be subscribed for because that'll be the focus of the for because that'll be the focus of the for because that'll be the focus of the Fable versus Astra video. So, is this my Fable versus Astra video. So, is this my Fable versus Astra video. So, is this my favorite model? That's a great question favorite model? That's a great question favorite model? That's a great question that will also be answered in that that will also be answered in that that will also be answered in that video. Is it the best model ever? I video. Is it the best model ever? I video. Is it the best model ever? I think I'm comfortable saying yes there. think I'm comfortable saying yes there. think I'm comfortable saying yes there. This model has so many unique This model has so many unique This model has so many unique capabilities that nothing else comes capabilities that nothing else comes capabilities that nothing else comes close to that it's an easy cell for me close to that it's an easy cell for me close to that it's an easy cell for me to say that. It's just insane. It makes to say that. It's just insane. It makes to say that. It's just insane. It makes every benchmark that currently exists every benchmark that currently exists every benchmark that currently exists feel wrong and outdated. It makes the feel wrong and outdated. It makes the feel wrong and outdated. It makes the way that we evaluate models feel kind of way that we evaluate models feel kind of way that we evaluate models feel kind of wrong as well. Even the term LLM doesn't wrong as well. Even the term LLM doesn't wrong as well. Even the term LLM doesn't feel right anymore because most of the feel right anymore because most of the feel right anymore because most of the things I'm using it for aren't just things I'm using it for aren't just things I'm using it for aren't just generating text. It might interface that generating text. It might interface that generating text. It might interface that way, but the work it's doing isn't that way, but the work it's doing isn't that way, but the work it's doing isn't that at all. A lot of people from OpenAI and at all. A lot of people from OpenAI and at all. A lot of people from OpenAI and even a few outside of it have been even a few outside of it have been even a few outside of it have been saying this model was the start of AGI, saying this model was the start of AGI, saying this model was the start of AGI, and in the end, I kind of see it. It

  33. and in the end, I kind of see it. It and in the end, I kind of see it. It does feel like a taste of something new, does feel like a taste of something new, does feel like a taste of something new, not just slightly better in all the not just slightly better in all the not just slightly better in all the usual ways. It's not like 30% more usual ways. It's not like 30% more usual ways. It's not like 30% more effective or 15% faster, all those effective or 15% faster, all those effective or 15% faster, all those things. It is in some places, but in a things. It is in some places, but in a things. It is in some places, but in a lot of these categories, it is so far lot of these categories, it is so far lot of these categories, it is so far ahead. It feels like something entirely ahead. It feels like something entirely ahead. It feels like something entirely new. It almost feels like an iPhone type new. It almost feels like an iPhone type new. It almost feels like an iPhone type change in that way, where the model is change in that way, where the model is change in that way, where the model is capable of stuff that I just didn't capable of stuff that I just didn't capable of stuff that I just didn't think AI could do at all, if ever. This think AI could do at all, if ever. This think AI could do at all, if ever. This is the model that I'm going to let run is the model that I'm going to let run is the model that I'm going to let run my computer, and it's already starting my computer, and it's already starting my computer, and it's already starting to run more and more of my business and to run more and more of my business and to run more and more of my business and my life. And that is an unbelievable my life. And that is an unbelievable my life. And that is an unbelievable achievement. Wherever I previously set achievement. Wherever I previously set achievement. Wherever I previously set my bar for good enough to trust almost my bar for good enough to trust almost my bar for good enough to trust almost feels hilariously wrong because we're so feels hilariously wrong because we're so feels hilariously wrong because we're so far past that point. It's stupid. I far past that point. It's stupid. I far past that point. It's stupid. I trust this model a ton. I use it an trust this model a ton. I use it an trust this model a ton. I use it an insane amount and I plan to continue insane amount and I plan to continue insane amount and I plan to continue doing that going forward. So my question doing that going forward. So my question doing that going forward. So my question to you isn't is this model great or not? to you isn't is this model great or not? to you isn't is this model great or not? Especially because you can't use it yet, Especially because you can't use it yet, Especially because you can't use it yet, which is stupid. My question to you is which is stupid. My question to you is which is stupid. My question to you is where is your bar? At what point are you where is your bar? At what point are you where is your bar? At what point are you going to stop checking the work the going to stop checking the work the going to stop checking the work the model does constantly and let it do its model does constantly and let it do its model does constantly and let it do its thing? I know I am past that point in so thing? I know I am past that point in so thing? I know I am past that point in so many of the things I do, but I'm curious many of the things I do, but I'm curious many of the things I do, but I'm curious how you guys feel. Do you have that bar how you guys feel. Do you have that bar how you guys feel. Do you have that bar set? Are you actually evaluating against set? Are you actually evaluating against set? Are you actually evaluating against it constantly to see if we've hit it?

  34. it constantly to see if we've hit it? it constantly to see if we've hit it? And do you see a future where you just And do you see a future where you just And do you see a future where you just trust the models and start to feel the trust the models and start to feel the trust the models and start to feel the AGI a bit more? This feels like a taste AGI a bit more? This feels like a taste AGI a bit more? This feels like a taste of something new and I cannot wait for of something new and I cannot wait for of something new and I cannot wait for you guys to see it as well. So until you guys to see it as well. So until you guys to see it as well. So until next time, these nerds

Summary

The main theme is the highly anticipated release and impressive capabilities of the GPT-6 Astra model, which has demonstrated revolutionary advancements over previous versions. Key subjects mentioned include extensive inference usage, comparisons to prior GPT models, and potential applications in coding, 3D modeling, and office work, though with some current limitations. The practical takeaway is that GPT-6 Astra represents a generational leap in AI technology, offering significant power and versatility, but users should also be aware of its availability, cost, and evolving release restrictions.

View original episode ↗