Oh no (the new Grok model is good)
Read full transcript 20 segments
-
A new model just dropped and its A new model just dropped and its creators are making some very bold creators are making some very bold creators are making some very bold claims. Specifically, they're saying claims. Specifically, they're saying claims. Specifically, they're saying that for dev work, it should compare to that for dev work, it should compare to that for dev work, it should compare to models like Fable 5 at a fraction of the models like Fable 5 at a fraction of the models like Fable 5 at a fraction of the cost. The model is Gro 5 from Space XAI, cost. The model is Gro 5 from Space XAI, cost. The model is Gro 5 from Space XAI, now partnered with Cursor. And I was a now partnered with Cursor. And I was a now partnered with Cursor. And I was a little skeptical after hearing this, but little skeptical after hearing this, but little skeptical after hearing this, but it turns out I was actually testing it. it turns out I was actually testing it. it turns out I was actually testing it. Over the last 24 hours, Cursor gave me Over the last 24 hours, Cursor gave me Over the last 24 hours, Cursor gave me early access to a new model that I early access to a new model that I early access to a new model that I thought was going to be a new composer, thought was going to be a new composer, thought was going to be a new composer, and I was pretty impressed with it. I and I was pretty impressed with it. I and I was pretty impressed with it. I learned today that that model was learned today that that model was learned today that that model was actually Grock 4.5. And now that I'm actually Grock 4.5. And now that I'm actually Grock 4.5. And now that I'm seeing the benchmarks, yeah, it's a seeing the benchmarks, yeah, it's a seeing the benchmarks, yeah, it's a pretty damn good model. This is from the pretty damn good model. This is from the pretty damn good model. This is from the artificial analysis code index, which artificial analysis code index, which artificial analysis code index, which combines a handful of benches that I combines a handful of benches that I combines a handful of benches that I actually like and trust. And according actually like and trust. And according actually like and trust. And according to this bench, Gro 45 is neckandneck to this bench, Gro 45 is neckandneck to this bench, Gro 45 is neckandneck with GPT55 and just barely below Fable with GPT55 and just barely below Fable with GPT55 and just barely below Fable while also beating out Opus 48. While I while also beating out Opus 48. While I while also beating out Opus 48. While I think these numbers might be a little think these numbers might be a little think these numbers might be a little bold based on my experience using it, bold based on my experience using it, bold based on my experience using it, they're not that far off. Gro 45 has they're not that far off. Gro 45 has they're not that far off. Gro 45 has been genuinely impressive. And it has a been genuinely impressive. And it has a been genuinely impressive. And it has a few things that are truly novel to it few things that are truly novel to it few things that are truly novel to it that I never would have guessed. And at that I never would have guessed. And at that I never would have guessed. And at its current price, it's a steal, its current price, it's a steal, its current price, it's a steal, especially with the 50% discount they're especially with the 50% discount they're especially with the 50% discount they're currently offering for people using it currently offering for people using it currently offering for people using it through tools like Cursor. I want to through tools like Cursor. I want to through tools like Cursor. I want to break down all of the good, the bad, and break down all of the good, the bad, and break down all of the good, the bad, and the ugly with this model. But first, a the ugly with this model. But first, a the ugly with this model. But first, a quick word from today's sponsor.
-
quick word from today's sponsor. quick word from today's sponsor. Normally, my ads are showcasing all the Normally, my ads are showcasing all the Normally, my ads are showcasing all the cool things AI can do. I'm going to do cool things AI can do. I'm going to do cool things AI can do. I'm going to do something a little different here. I something a little different here. I something a little different here. I want to show you something AI failed want to show you something AI failed want to show you something AI failed hard at with me. LG, notoriously hard at with me. LG, notoriously hard at with me. LG, notoriously wonderful at naming, announced a monitor wonderful at naming, announced a monitor wonderful at naming, announced a monitor I've been really excited about since I've been really excited about since I've been really excited about since January that still hasn't come out yet. January that still hasn't come out yet. January that still hasn't come out yet. This monitor has a bunch of cool tech This monitor has a bunch of cool tech This monitor has a bunch of cool tech that I'm excited about. That's not what that I'm excited about. That's not what that I'm excited about. That's not what I'm here to talk about, though. I'm here I'm here to talk about, though. I'm here I'm here to talk about, though. I'm here to talk about my attempts to get it. to talk about my attempts to get it. to talk about my attempts to get it. When you click the notify me button, it When you click the notify me button, it When you click the notify me button, it doesn't notify you. It launches a broken doesn't notify you. It launches a broken doesn't notify you. It launches a broken JavaScript thing, which in other JavaScript thing, which in other JavaScript thing, which in other browsers lets you sign up for browsers lets you sign up for browsers lets you sign up for notifications, and it doesn't even have notifications, and it doesn't even have notifications, and it doesn't even have the monitor I want in the options. So, I the monitor I want in the options. So, I the monitor I want in the options. So, I have no way of knowing when this monitor have no way of knowing when this monitor have no way of knowing when this monitor comes out. So, I did what any nerd would comes out. So, I did what any nerd would comes out. So, I did what any nerd would do. I asked my Hermes agent to monitor do. I asked my Hermes agent to monitor do. I asked my Hermes agent to monitor it and let me know when it comes out. it and let me know when it comes out. it and let me know when it comes out. And it did every single day because it And it did every single day because it And it did every single day because it was waiting for content on the page to was waiting for content on the page to was waiting for content on the page to change. And the content was changing, change. And the content was changing, change. And the content was changing, but never changing in the ways that but never changing in the ways that but never changing in the ways that mattered. And my AI was not smart enough mattered. And my AI was not smart enough mattered. And my AI was not smart enough to realize that the reads it was getting to realize that the reads it was getting to realize that the reads it was getting were not actual changed content on the were not actual changed content on the were not actual changed content on the page. All I wanted to know is when this page. All I wanted to know is when this page. All I wanted to know is when this monitor came out for sale. Today's monitor came out for sale. Today's monitor came out for sale. Today's sponsor is Firecrawl and they make it sponsor is Firecrawl and they make it sponsor is Firecrawl and they make it way easier for your agents to scrape the way easier for your agents to scrape the way easier for your agents to scrape the web. That in and of itself would have web. That in and of itself would have web. That in and of itself would have been really useful for this type of been really useful for this type of been really useful for this type of thing. If it translates the page to thing. If it translates the page to thing. If it translates the page to markdown, it's more likely to notice markdown, it's more likely to notice markdown, it's more likely to notice when things have changed because it can when things have changed because it can when things have changed because it can get the page in markdown. It can get a get the page in markdown. It can get a get the page in markdown. It can get a screenshot of the page or it can get the screenshot of the page or it can get the screenshot of the page or it can get the plain text. Super useful. But what it plain text. Super useful. But what it plain text. Super useful. But what it can also do is a new feature they just can also do is a new feature they just can also do is a new feature they just added called monitoring where you added called monitoring where you added called monitoring where you schedule a recurring check to detect schedule a recurring check to detect schedule a recurring check to detect changes on a page. This makes this exact changes on a page. This makes this exact changes on a page. This makes this exact task trivial. And not only is it easy, task trivial. And not only is it easy, task trivial. And not only is it easy, it falls entirely under the free tier.
-
it falls entirely under the free tier. it falls entirely under the free tier. And even if it didn't fall under the And even if it didn't fall under the And even if it didn't fall under the free tier, I could have just forked it free tier, I could have just forked it free tier, I could have just forked it and ran it myself because they're open and ran it myself because they're open and ran it myself because they're open source. Buyer crawl has pretty much source. Buyer crawl has pretty much source. Buyer crawl has pretty much everything your agents need to scrape everything your agents need to scrape everything your agents need to scrape the web. From search to URL specific the web. From search to URL specific the web. From search to URL specific readouts to interactions to monitoring readouts to interactions to monitoring readouts to interactions to monitoring to an MCP server to make it way easier to an MCP server to make it way easier to an MCP server to make it way easier to connect that doesn't even need an API to connect that doesn't even need an API to connect that doesn't even need an API key, so it's trivial to set up. I can key, so it's trivial to set up. I can key, so it's trivial to set up. I can see why these guys are hiring right now. see why these guys are hiring right now. see why these guys are hiring right now. They're clearly doing really well. All They're clearly doing really well. All They're clearly doing really well. All of this is so useful. And if you want to of this is so useful. And if you want to of this is so useful. And if you want to see it yourself, check them out now at see it yourself, check them out now at see it yourself, check them out now at swive.link/firecrawl. swive.link/firecrawl. swive.link/firecrawl. As I mentioned before, I've been using As I mentioned before, I've been using As I mentioned before, I've been using this model for a ton of real work, as this model for a ton of real work, as this model for a ton of real work, as well as some silly demos, like making 3D well as some silly demos, like making 3D well as some silly demos, like making 3D games and whatnot. I've been going at it games and whatnot. I've been going at it games and whatnot. I've been going at it non-stop for the last 24-ish hours, and non-stop for the last 24-ish hours, and non-stop for the last 24-ish hours, and I have a lot of thoughts. But first, I I have a lot of thoughts. But first, I I have a lot of thoughts. But first, I want to start with the official want to start with the official want to start with the official reporting. Introducing Grock 4.5. It's reporting. Introducing Grock 4.5. It's reporting. Introducing Grock 4.5. It's SpaceX's smartest model built for SpaceX's smartest model built for SpaceX's smartest model built for coding, agentic tasks, and knowledge coding, agentic tasks, and knowledge coding, agentic tasks, and knowledge work. Let's talk about what we care work. Let's talk about what we care work. Let's talk about what we care about. Real world engineering about. Real world engineering about. Real world engineering excellence. Grock 45 was trained on data excellence. Grock 45 was trained on data excellence. Grock 45 was trained on data sets spanning knowledge in coding, sets spanning knowledge in coding, sets spanning knowledge in coding, science, engineering, and math. With science, engineering, and math. With science, engineering, and math. With both intelligent and efficient both intelligent and efficient both intelligent and efficient reasoning, Grock 4.5 excels at real reasoning, Grock 4.5 excels at real reasoning, Grock 4.5 excels at real engineering tasks and it exceeds engineering tasks and it exceeds engineering tasks and it exceeds comparable leading models at many of comparable leading models at many of comparable leading models at many of these tasks. So, Deepsw SWE, which is these tasks. So, Deepsw SWE, which is these tasks. So, Deepsw SWE, which is the software bench that I've been the software bench that I've been the software bench that I've been repping pretty hard. I find it to be a repping pretty hard. I find it to be a repping pretty hard. I find it to be a very reasonable benchmark measuring how very reasonable benchmark measuring how very reasonable benchmark measuring how work actually happens with agents.
-
work actually happens with agents. work actually happens with agents. Obviously, Fable is still number one. Obviously, Fable is still number one. Obviously, Fable is still number one. DBT55 is still number two. We don't have DBT55 is still number two. We don't have DBT55 is still number two. We don't have numbers for 5.6 yet. Excited to see numbers for 5.6 yet. Excited to see numbers for 5.6 yet. Excited to see those. But for now, third place is those. But for now, third place is those. But for now, third place is Grock. Massively beating out any Google Grock. Massively beating out any Google Grock. Massively beating out any Google model. I don't think any are even model. I don't think any are even model. I don't think any are even referenced here. Opus 48 is falling a referenced here. Opus 48 is falling a referenced here. Opus 48 is falling a decent bit behind and then 47 is way decent bit behind and then 47 is way decent bit behind and then 47 is way lower. This is a new like frontier tier lower. This is a new like frontier tier lower. This is a new like frontier tier that we are seeing happen and Grock 45 that we are seeing happen and Grock 45 that we are seeing happen and Grock 45 is on the line between last gen and this is on the line between last gen and this is on the line between last gen and this gen in a lot of ways. 4.5 was trained gen in a lot of ways. 4.5 was trained gen in a lot of ways. 4.5 was trained across tens of thousands of GB300 GPUs across tens of thousands of GB300 GPUs across tens of thousands of GB300 GPUs which is the newest technology from which is the newest technology from which is the newest technology from Nvidia. Training and stability Nvidia. Training and stability Nvidia. Training and stability techniques designed for large scale runs techniques designed for large scale runs techniques designed for large scale runs beyond raw token volume. We invested beyond raw token volume. We invested beyond raw token volume. We invested heavily in data filtering and curation, heavily in data filtering and curation, heavily in data filtering and curation, dduplication, quality scoring and domain dduplication, quality scoring and domain dduplication, quality scoring and domain focus selection. So the data mixture focus selection. So the data mixture focus selection. So the data mixture stayed high coverage and high signal. stayed high coverage and high signal. stayed high coverage and high signal. That is interesting. This model does That is interesting. This model does That is interesting. This model does seem to be a full new base like seem to be a full new base like seem to be a full new base like pre-training rather than an adjustment pre-training rather than an adjustment pre-training rather than an adjustment on previous Grock models. The biggest on previous Grock models. The biggest on previous Grock models. The biggest indicator of that is that they mention indicator of that is that they mention indicator of that is that they mention it's a 1.5 trill per model where it's a 1.5 trill per model where it's a 1.5 trill per model where previous ones were only 500 bill per. So previous ones were only 500 bill per. So previous ones were only 500 bill per. So clearly whole new model here. Their RL clearly whole new model here. Their RL clearly whole new model here. Their RL training covered hundreds of thousands training covered hundreds of thousands training covered hundreds of thousands of tasks centered on multi-step software of tasks centered on multi-step software of tasks centered on multi-step software engineering and other technical work engineering and other technical work engineering and other technical work with automated and modelbased grading.
-
with automated and modelbased grading. with automated and modelbased grading. Our stack is built for highly Our stack is built for highly Our stack is built for highly asynchronous training, so agentic asynchronous training, so agentic asynchronous training, so agentic rollouts can run for many hours while rollouts can run for many hours while rollouts can run for many hours while learning continues across tens of learning continues across tens of learning continues across tens of thousands of GPUs. The result is more thousands of GPUs. The result is more thousands of GPUs. The result is more intelligent and efficient reasoning on intelligent and efficient reasoning on intelligent and efficient reasoning on real engineering and agentic tasks. They real engineering and agentic tasks. They real engineering and agentic tasks. They give some examples of one-shotted tasks give some examples of one-shotted tasks give some examples of one-shotted tasks that they have the model make, and it's that they have the model make, and it's that they have the model make, and it's impressive. It does 3D particularly impressive. It does 3D particularly impressive. It does 3D particularly well, which I will be sure to talk about well, which I will be sure to talk about well, which I will be sure to talk about in a bit. It's also quite fast, both in a bit. It's also quite fast, both in a bit. It's also quite fast, both because it uses not too many tokens and because it uses not too many tokens and because it uses not too many tokens and because it is running faster in general because it is running faster in general because it is running faster in general on the info they're serving it on at on the info they're serving it on at on the info they're serving it on at roughly 80 tokens per second. Cursor, roughly 80 tokens per second. Cursor, roughly 80 tokens per second. Cursor, who is now part of SpaceX, has their own who is now part of SpaceX, has their own who is now part of SpaceX, has their own article about this drop and it has a article about this drop and it has a article about this drop and it has a couple more interesting details I wanted couple more interesting details I wanted couple more interesting details I wanted to jump on. First off, they have more to jump on. First off, they have more to jump on. First off, they have more benchmarks, including Terminal Bench benchmarks, including Terminal Bench benchmarks, including Terminal Bench 2.1, where it performed just barely 2.1, where it performed just barely 2.1, where it performed just barely below GPT55 and pretty close to Fable below GPT55 and pretty close to Fable below GPT55 and pretty close to Fable with an 83.3 versus 83.4 and 84.3, with an 83.3 versus 83.4 and 84.3, with an 83.3 versus 83.4 and 84.3, respectively. SWE Bench, which I don't respectively. SWE Bench, which I don't respectively. SWE Bench, which I don't care about. It's a bad benchmark, so care about. It's a bad benchmark, so care about. It's a bad benchmark, so I'll skip it. And then Deepswe, which I'll skip it. And then Deepswe, which I'll skip it. And then Deepswe, which you mentioned before, it's doing very you mentioned before, it's doing very you mentioned before, it's doing very well on. SWB Pro, which again, bad well on. SWB Pro, which again, bad well on. SWB Pro, which again, bad benchmark. I don't care. Cursor benchmark. I don't care. Cursor benchmark. I don't care. Cursor subscription plans for individuals and subscription plans for individuals and subscription plans for individuals and teams include significant usage of the teams include significant usage of the teams include significant usage of the model with double usage for the first model with double usage for the first model with double usage for the first week. This is the other exciting thing.
-
week. This is the other exciting thing. week. This is the other exciting thing. One of the problems Cursor has as a One of the problems Cursor has as a One of the problems Cursor has as a business is that they struggle to business is that they struggle to business is that they struggle to compete with the labs just doing this compete with the labs just doing this compete with the labs just doing this crazy subsidization because you can pay crazy subsidization because you can pay crazy subsidization because you can pay OpenAI 200 bucks and then get up to OpenAI 200 bucks and then get up to OpenAI 200 bucks and then get up to $14,000 of inference not even counting $14,000 of inference not even counting $14,000 of inference not even counting resets. That's hard for them to compete resets. That's hard for them to compete resets. That's hard for them to compete on. They've won on enterprise still on. They've won on enterprise still on. They've won on enterprise still because the API rates are what the because the API rates are what the because the API rates are what the enterprises often have to pay. They enterprises often have to pay. They enterprises often have to pay. They don't get that crazy subsidization from don't get that crazy subsidization from don't get that crazy subsidization from the labs. But winning on individuals has the labs. But winning on individuals has the labs. But winning on individuals has been tough and this is a huge win for been tough and this is a huge win for been tough and this is a huge win for them there because they finally have them there because they finally have them there because they finally have models they can kind of subsidize. Not models they can kind of subsidize. Not models they can kind of subsidize. Not that they have to though because the that they have to though because the that they have to though because the price is really good. We haven't got price is really good. We haven't got price is really good. We haven't got there yet, but we will in just a bit. As there yet, but we will in just a bit. As there yet, but we will in just a bit. As they mentioned, Grock 45 is a mixture of they mentioned, Grock 45 is a mixture of they mentioned, Grock 45 is a mixture of experts model. They trained jointly with experts model. They trained jointly with experts model. They trained jointly with SpaceX. So, this is a model that was SpaceX. So, this is a model that was SpaceX. So, this is a model that was still a Grock model, still SpaceX still a Grock model, still SpaceX still a Grock model, still SpaceX focused, but they came in and trained it focused, but they came in and trained it focused, but they came in and trained it jointly, bringing data as well as their jointly, bringing data as well as their jointly, bringing data as well as their own processes. Training included own processes. Training included own processes. Training included trillions of tokens of cursor data, trillions of tokens of cursor data, trillions of tokens of cursor data, which capture a wide range of user which capture a wide range of user which capture a wide range of user interactions with code bases and interactions with code bases and interactions with code bases and software tools. This data set lets the software tools. This data set lets the software tools. This data set lets the model learn both from existing software model learn both from existing software model learn both from existing software as well as developer agent interactions, as well as developer agent interactions, as well as developer agent interactions, capturing how developers work and how capturing how developers work and how capturing how developers work and how agents interact with their environments.
-
agents interact with their environments. agents interact with their environments. I've seen a lot of good examples here I've seen a lot of good examples here I've seen a lot of good examples here which we will definitely showcase. When which we will definitely showcase. When which we will definitely showcase. When they trained Composer 25 to be a coding they trained Composer 25 to be a coding they trained Composer 25 to be a coding specialist, Grock 45 kept the training specialist, Grock 45 kept the training specialist, Grock 45 kept the training data intentionally mixed and more broad. data intentionally mixed and more broad. data intentionally mixed and more broad. This involved drawing on highquality This involved drawing on highquality This involved drawing on highquality STEM tasks, research papers, and other STEM tasks, research papers, and other STEM tasks, research papers, and other knowledge work so that the model gained knowledge work so that the model gained knowledge work so that the model gained proficiency across a wide range of proficiency across a wide range of proficiency across a wide range of domains. They made a bunch of difficult domains. They made a bunch of difficult domains. They made a bunch of difficult RL problems for the model because even a RL problems for the model because even a RL problems for the model because even a lot of the traditionally hard problems lot of the traditionally hard problems lot of the traditionally hard problems are now trivial for models and in RL are now trivial for models and in RL are now trivial for models and in RL they want to have really really they want to have really really they want to have really really difficult stuff and I think that's why difficult stuff and I think that's why difficult stuff and I think that's why it's benching so well. It can get it can it's benching so well. It can get it can it's benching so well. It can get it can just go on those types of big bold hard just go on those types of big bold hard just go on those types of big bold hard tasks. You might have noticed one bench tasks. You might have noticed one bench tasks. You might have noticed one bench missing from here though. Cursor bench. missing from here though. Cursor bench. missing from here though. Cursor bench. That is certainly not because it That is certainly not because it That is certainly not because it performed poorly. As we see here the top performed poorly. As we see here the top performed poorly. As we see here the top right being the cheapest and the best. right being the cheapest and the best. right being the cheapest and the best. It performed comparable to Fable 5 High It performed comparable to Fable 5 High It performed comparable to Fable 5 High for a significantly lower price where for a significantly lower price where for a significantly lower price where Fable 5 High cost $8.77 per task and Fable 5 High cost $8.77 per task and Fable 5 High cost $8.77 per task and Grock 4.5 cost $151 for a slightly Grock 4.5 cost $151 for a slightly Grock 4.5 cost $151 for a slightly higher score at that tier. Obviously, higher score at that tier. Obviously, higher score at that tier. Obviously, the best is still Fable on Max, even the best is still Fable on Max, even the best is still Fable on Max, even though it's double the price of Fable on though it's double the price of Fable on though it's double the price of Fable on High. But yeah, Grock 4.5 High is High. But yeah, Grock 4.5 High is High. But yeah, Grock 4.5 High is looking insane by this chart, even looking insane by this chart, even looking insane by this chart, even crushing out GPT. But how? Like it it crushing out GPT. But how? Like it it crushing out GPT. But how? Like it it can't possibly be that good, right? If can't possibly be that good, right? If can't possibly be that good, right? If you scroll a little, you see why they you scroll a little, you see why they you scroll a little, you see why they did not include this information. Grock did not include this information. Grock did not include this information. Grock 4.5 has an advantage on cursor bench. An 4.5 has an advantage on cursor bench. An 4.5 has an advantage on cursor bench. An earlier snapshot of the cursor codebase earlier snapshot of the cursor codebase earlier snapshot of the cursor codebase was unintentionally included in was unintentionally included in was unintentionally included in training. The exact score impact is training. The exact score impact is training. The exact score impact is unclear. The data has been removed from
-
unclear. The data has been removed from unclear. The data has been removed from future models. For a rundown of third future models. For a rundown of third future models. For a rundown of third party benchmark scores, see the Grock 45 party benchmark scores, see the Grock 45 party benchmark scores, see the Grock 45 launch blog. Yep, they accidentally put launch blog. Yep, they accidentally put launch blog. Yep, they accidentally put cursors actual code in the training cursors actual code in the training cursors actual code in the training data. And since cursor bench is based on data. And since cursor bench is based on data. And since cursor bench is based on real problems that they have in cursor real problems that they have in cursor real problems that they have in cursor and working on cursor, this bench is now and working on cursor, this bench is now and working on cursor, this bench is now kind of tainted at the very least in the kind of tainted at the very least in the kind of tainted at the very least in the Grock world. It's a shame because I Grock world. It's a shame because I Grock world. It's a shame because I liked this bench, but uh mistakes liked this bench, but uh mistakes liked this bench, but uh mistakes happen. I'm happy they were transparent happen. I'm happy they were transparent happen. I'm happy they were transparent about it and that they're not about it and that they're not about it and that they're not advertising Grock 45 via this benchmark advertising Grock 45 via this benchmark advertising Grock 45 via this benchmark publicly, which other labs may have done publicly, which other labs may have done publicly, which other labs may have done similar things to. They're being similar things to. They're being similar things to. They're being straightforward and transparent with it. straightforward and transparent with it. straightforward and transparent with it. I appreciate them for that. Now, I want I appreciate them for that. Now, I want I appreciate them for that. Now, I want to talk about the price. Gro 4.5 is $2 to talk about the price. Gro 4.5 is $2 to talk about the price. Gro 4.5 is $2 per million tokens in and $6 per million per million tokens in and $6 per million per million tokens in and $6 per million tokens out. That makes it comically tokens out. That makes it comically tokens out. That makes it comically cheaper than a lot of competing models. cheaper than a lot of competing models. cheaper than a lot of competing models. For example, Fable is $10 per million For example, Fable is $10 per million For example, Fable is $10 per million tokens in and 50 per million tokens out. tokens in and 50 per million tokens out. tokens in and 50 per million tokens out. Between five and almost 10x the cost. Between five and almost 10x the cost. Between five and almost 10x the cost. That ignores the fact that Fable is That ignores the fact that Fable is That ignores the fact that Fable is relatively token hungry and this model relatively token hungry and this model relatively token hungry and this model seems to be less. So, hard to know for seems to be less. So, hard to know for seems to be less. So, hard to know for sure until we've really put it through sure until we've really put it through sure until we've really put it through its paces, but based on all the benches its paces, but based on all the benches its paces, but based on all the benches I've seen and all the work I've done I've seen and all the work I've done I've seen and all the work I've done with it, it's relatively efficient. It with it, it's relatively efficient. It with it, it's relatively efficient. It is also worth noting that this price is also worth noting that this price is also worth noting that this price only applies under 200,000 tokens of only applies under 200,000 tokens of only applies under 200,000 tokens of context. If you go over, the price is context. If you go over, the price is context. If you go over, the price is double to $4 per million tokens in and double to $4 per million tokens in and double to $4 per million tokens in and 12 per million tokens out. Still way 12 per million tokens out. Still way 12 per million tokens out. Still way cheaper than any other model at this cheaper than any other model at this cheaper than any other model at this tier, but it only goes up to 500k tier, but it only goes up to 500k tier, but it only goes up to 500k tokens. It's kind of weird to have a tokens. It's kind of weird to have a tokens. It's kind of weird to have a model that can go over 200k and charges
-
model that can go over 200k and charges model that can go over 200k and charges more but is still under a mill more but is still under a mill more but is still under a mill specifically like 200k to 500k is not specifically like 200k to 500k is not specifically like 200k to 500k is not that much more context and to build that much more context and to build that much more context and to build twice as much for it especially to build twice as much for it especially to build twice as much for it especially to build twice as much on output feels a little twice as much on output feels a little twice as much on output feels a little much to me. Feel like they're reaching a much to me. Feel like they're reaching a much to me. Feel like they're reaching a little here. My guess is the reason they little here. My guess is the reason they little here. My guess is the reason they did that is they wanted to get the input did that is they wanted to get the input did that is they wanted to get the input and output token costs for the base tier and output token costs for the base tier and output token costs for the base tier 200k version as cheap as possible. And 200k version as cheap as possible. And 200k version as cheap as possible. And the GPUs that this is running on are the GPUs that this is running on are the GPUs that this is running on are also being resold to companies like also being resold to companies like also being resold to companies like Anthropic and Google with massive Anthropic and Google with massive Anthropic and Google with massive markups. They have to make sure that markups. They have to make sure that markups. They have to make sure that it's priced in a way where they're not it's priced in a way where they're not it's priced in a way where they're not losing too much money that could have losing too much money that could have losing too much money that could have been made from reselling GPUs, but at been made from reselling GPUs, but at been made from reselling GPUs, but at the same time is priced cheaply enough the same time is priced cheaply enough the same time is priced cheaply enough to actually compete with those labs. to actually compete with those labs. to actually compete with those labs. SpaceX has put themselves in a weird SpaceX has put themselves in a weird SpaceX has put themselves in a weird spot here, but I think they navigated spot here, but I think they navigated spot here, but I think they navigated okay according to everything I've been okay according to everything I've been okay according to everything I've been seeing and all the use I've been having seeing and all the use I've been having seeing and all the use I've been having so far. Let's go over the artificial so far. Let's go over the artificial so far. Let's go over the artificial analysis numbers and then I will dive analysis numbers and then I will dive analysis numbers and then I will dive into my experience. SpaceX AAI's Gro 45 into my experience. SpaceX AAI's Gro 45 into my experience. SpaceX AAI's Gro 45 scored a 54 which places it fourth in scored a 54 which places it fourth in scored a 54 which places it fourth in the artificial analysis intelligence the artificial analysis intelligence the artificial analysis intelligence index. Kind of wild to see a different index. Kind of wild to see a different index. Kind of wild to see a different color in the top again. It's been a very color in the top again. It's been a very color in the top again. It's been a very long time since I saw purple all the way long time since I saw purple all the way long time since I saw purple all the way up here. It is right behind GPT55 and up here. It is right behind GPT55 and up here. It is right behind GPT55 and just ahead of Sonnet 5 based on the just ahead of Sonnet 5 based on the just ahead of Sonnet 5 based on the intelligence index crushing GLM52 which intelligence index crushing GLM52 which intelligence index crushing GLM52 which is even more interesting when you is even more interesting when you is even more interesting when you realize how expensive GLM52 is to run.
-
realize how expensive GLM52 is to run. realize how expensive GLM52 is to run. GLM52's base price on a lot of providers GLM52's base price on a lot of providers GLM52's base price on a lot of providers is a $110 in roughly and $4.40 out is a $110 in roughly and $4.40 out is a $110 in roughly and $4.40 out roughly. There are some providers that roughly. There are some providers that roughly. There are some providers that offer quite a bit cheaper now, like offer quite a bit cheaper now, like offer quite a bit cheaper now, like Novita has a temporary 60% off. Deep Novita has a temporary 60% off. Deep Novita has a temporary 60% off. Deep Infra has it at like $3ish per mill out. Infra has it at like $3ish per mill out. Infra has it at like $3ish per mill out. When you remember how much more token When you remember how much more token When you remember how much more token hungry GLM52 is, you realize it's kind hungry GLM52 is, you realize it's kind hungry GLM52 is, you realize it's kind of been crushed by Grock 45. Here we can of been crushed by Grock 45. Here we can of been crushed by Grock 45. Here we can see the actual costs incurred per task see the actual costs incurred per task see the actual costs incurred per task average across the entire suite. And average across the entire suite. And average across the entire suite. And Kimmy K26 was about 35 cents per task. Kimmy K26 was about 35 cents per task. Kimmy K26 was about 35 cents per task. GLM52 is about 37 cents per task. And GLM52 is about 37 cents per task. And GLM52 is about 37 cents per task. And then Gro5, which scored way higher than then Gro5, which scored way higher than then Gro5, which scored way higher than those other models, was only 31 per those other models, was only 31 per those other models, was only 31 per task. And for reference, Fable 5 was task. And for reference, Fable 5 was task. And for reference, Fable 5 was $2.75 $2.75 $2.75 for the same work. Yeah, it's an for the same work. Yeah, it's an for the same work. Yeah, it's an efficient model. And if you look at the efficient model. And if you look at the efficient model. And if you look at the intelligence versus cost chart, you see intelligence versus cost chart, you see intelligence versus cost chart, you see it is very well positioned. It is just it is very well positioned. It is just it is very well positioned. It is just on the edge of the green box, you know, on the edge of the green box, you know, on the edge of the green box, you know, the the good spot that almost nothing is the the good spot that almost nothing is the the good spot that almost nothing is in. Both Gro 45 and Gemini 31 Pro score in. Both Gro 45 and Gemini 31 Pro score in. Both Gro 45 and Gemini 31 Pro score very well here. And Grock 45 is very well here. And Grock 45 is very well here. And Grock 45 is meaningfully more intelligent while also meaningfully more intelligent while also meaningfully more intelligent while also being in this price range. This gets being in this price range. This gets being in this price range. This gets much crazier with the coding focused much crazier with the coding focused much crazier with the coding focused benches, though, because Grock came out benches, though, because Grock came out benches, though, because Grock came out swinging here, crushing every Google swinging here, crushing every Google swinging here, crushing every Google model by a large margin. And if we look model by a large margin. And if we look model by a large margin. And if we look at the token usage per task for the code at the token usage per task for the code at the token usage per task for the code work, you'll see Opus and Fable both work, you'll see Opus and Fable both work, you'll see Opus and Fable both massive token hogs at 7.2 mil for Fable massive token hogs at 7.2 mil for Fable massive token hogs at 7.2 mil for Fable and 9.2 mil for Opus. 55 on X high is and 9.2 mil for Opus. 55 on X high is and 9.2 mil for Opus. 55 on X high is still in the 6 mil range or so. 55 on
-
still in the 6 mil range or so. 55 on still in the 6 mil range or so. 55 on medium was only 3.5 mil tokens, but all medium was only 3.5 mil tokens, but all medium was only 3.5 mil tokens, but all the way at the end here, Grock build the way at the end here, Grock build the way at the end here, Grock build with Grock 4.5 only 2 mil tokens to get with Grock 4.5 only 2 mil tokens to get with Grock 4.5 only 2 mil tokens to get that high of a score. That is an insane that high of a score. That is an insane that high of a score. That is an insane level of token efficiency. I never would level of token efficiency. I never would level of token efficiency. I never would have guessed that this model would be so have guessed that this model would be so have guessed that this model would be so efficient, but it really is. And the efficient, but it really is. And the efficient, but it really is. And the result is that it feels way faster to result is that it feels way faster to result is that it feels way faster to use and the bill ends up being cheaper use and the bill ends up being cheaper use and the bill ends up being cheaper than you might have expected. It kind of than you might have expected. It kind of than you might have expected. It kind of makes this a great go-to default code makes this a great go-to default code makes this a great go-to default code model that you bring other things in to model that you bring other things in to model that you bring other things in to clean up after if you use it and it clean up after if you use it and it clean up after if you use it and it doesn't do quite what you want. doesn't do quite what you want. doesn't do quite what you want. Artificial analysis also calls out that Artificial analysis also calls out that Artificial analysis also calls out that it does very well on agentic tasks, it does very well on agentic tasks, it does very well on agentic tasks, things that are multi-step where it has things that are multi-step where it has things that are multi-step where it has to call tools and synthesize to call tools and synthesize to call tools and synthesize information. One of the most information. One of the most information. One of the most costefficient models to run for near costefficient models to run for near costefficient models to run for near frontier intelligence. Yep, it is frontier intelligence. Yep, it is frontier intelligence. Yep, it is insanely cheap. The token efficiency insanely cheap. The token efficiency insanely cheap. The token efficiency combined with the low price is what combined with the low price is what combined with the low price is what makes it so compelling. As a coding makes it so compelling. As a coding makes it so compelling. As a coding agent, Grock 455 and Grock build is on agent, Grock 455 and Grock build is on agent, Grock 455 and Grock build is on par with 55 and offers efficiency par with 55 and offers efficiency par with 55 and offers efficiency benefits. Still insane because 55 was benefits. Still insane because 55 was benefits. Still insane because 55 was such an efficient model. Getting more such an efficient model. Getting more such an efficient model. Getting more efficient is just unbelievable. And when efficient is just unbelievable. And when efficient is just unbelievable. And when you see how big of a jump this is for you see how big of a jump this is for you see how big of a jump this is for them, it's kind of insane. They went them, it's kind of insane. They went them, it's kind of insane. They went from in the thousands on the human from in the thousands on the human from in the thousands on the human baseline to 1543 leaprogging across baseline to 1543 leaprogging across baseline to 1543 leaprogging across multiple labs entire like last decade.
-
multiple labs entire like last decade. multiple labs entire like last decade. It's crazy how far they have jumped on It's crazy how far they have jumped on It's crazy how far they have jumped on all of these things. and benches like all of these things. and benches like all of these things. and benches like how to they are now number one in the how to they are now number one in the how to they are now number one in the world in just crazy crazy leap and it world in just crazy crazy leap and it world in just crazy crazy leap and it really shows the benefit that cursor really shows the benefit that cursor really shows the benefit that cursor brings XAI in being part of the business brings XAI in being part of the business brings XAI in being part of the business it seems like cursor's combination of it seems like cursor's combination of it seems like cursor's combination of like data and RL process has been like data and RL process has been like data and RL process has been incredibly beneficial to X so far that incredibly beneficial to X so far that incredibly beneficial to X so far that does not mean it scores well in does not mean it scores well in does not mean it scores well in everything though for example in skate everything though for example in skate everything though for example in skate bench it ended up being quite expensive bench it ended up being quite expensive bench it ended up being quite expensive because it kept thinking and reasoning because it kept thinking and reasoning because it kept thinking and reasoning trying to figure out the tricks and it trying to figure out the tricks and it trying to figure out the tricks and it still only got a 76% which is the lowest still only got a 76% which is the lowest still only got a 76% which is the lowest score from a Frontier Lab on the max score from a Frontier Lab on the max score from a Frontier Lab on the max reasoning settings while also being reasoning settings while also being reasoning settings while also being relatively expensive at 1.3 cents per relatively expensive at 1.3 cents per relatively expensive at 1.3 cents per run. Not as bad as something like Sonnet run. Not as bad as something like Sonnet run. Not as bad as something like Sonnet 5, which cost way, way, way, way more 5, which cost way, way, way, way more 5, which cost way, way, way, way more than it should have at 15 cents per run. than it should have at 15 cents per run. than it should have at 15 cents per run. Literally 10x the cost for a lower Literally 10x the cost for a lower Literally 10x the cost for a lower score, but you compare it to something score, but you compare it to something score, but you compare it to something like Gemini 31 Pro Preview and ended up like Gemini 31 Pro Preview and ended up like Gemini 31 Pro Preview and ended up being little over half the price, but a being little over half the price, but a being little over half the price, but a meaningfully lower score. Yeah, I was meaningfully lower score. Yeah, I was meaningfully lower score. Yeah, I was hoping it would be a little cheaper hoping it would be a little cheaper hoping it would be a little cheaper here, but it does seem very determined here, but it does seem very determined here, but it does seem very determined to get answers as it averaged at 2,100 to get answers as it averaged at 2,100 to get answers as it averaged at 2,100 tokens per response, making it one of tokens per response, making it one of tokens per response, making it one of the most heavy reasoning models I have the most heavy reasoning models I have the most heavy reasoning models I have used here where it really thought before used here where it really thought before used here where it really thought before giving an answer. All of this said, giving an answer. All of this said, giving an answer. All of this said, benchmarks are benchmarks and a lot of benchmarks are benchmarks and a lot of benchmarks are benchmarks and a lot of them don't measure how it feels to use them don't measure how it feels to use them don't measure how it feels to use the model in the real world. For the model in the real world. For the model in the real world. For example, a lot of the Google models example, a lot of the Google models example, a lot of the Google models score great and when I use them for score great and when I use them for score great and when I use them for code, they just don't feel particularly code, they just don't feel particularly code, they just don't feel particularly good to use. So, how has Grock 45 been good to use. So, how has Grock 45 been good to use. So, how has Grock 45 been in my usage? I'll be frank, I'm
-
in my usage? I'll be frank, I'm in my usage? I'll be frank, I'm impressed. I'm working hard on my new impressed. I'm working hard on my new impressed. I'm working hard on my new cloud product, Lake Bed, and I'm really cloud product, Lake Bed, and I'm really cloud product, Lake Bed, and I'm really close to shipping. I wanted to spend a close to shipping. I wanted to spend a close to shipping. I wanted to spend a lot of time the last few days doing a lot of time the last few days doing a lot of time the last few days doing a big pass, auditing it, finding any big pass, auditing it, finding any big pass, auditing it, finding any potential issues, whether it's security, potential issues, whether it's security, potential issues, whether it's security, maintenance problems, things that aren't maintenance problems, things that aren't maintenance problems, things that aren't great to have in an open source project, great to have in an open source project, great to have in an open source project, that type of stuff. I had it do an audit that type of stuff. I had it do an audit that type of stuff. I had it do an audit and it did a pretty good job. It found and it did a pretty good job. It found and it did a pretty good job. It found most of the things that Fable and GPT56 most of the things that Fable and GPT56 most of the things that Fable and GPT56 found. However, it did a great job when found. However, it did a great job when found. However, it did a great job when I asked it to start fixing the things I asked it to start fixing the things I asked it to start fixing the things and when I noticed how well it was and when I noticed how well it was and when I noticed how well it was doing, I started to push it a little doing, I started to push it a little doing, I started to push it a little hard. I opened with working on hardening hard. I opened with working on hardening hard. I opened with working on hardening lake bed for its first public release. I lake bed for its first public release. I lake bed for its first public release. I have the following PR up which addresses have the following PR up which addresses have the following PR up which addresses the majority of the remaining issues the majority of the remaining issues the majority of the remaining issues with the link to the PR. Here's the with the link to the PR. Here's the with the link to the PR. Here's the report for the majority of those issues report for the majority of those issues report for the majority of those issues with a post plan link. Have we resolved with a post plan link. Have we resolved with a post plan link. Have we resolved them? Are there other things worth them? Are there other things worth them? Are there other things worth solving before launch? I had an issue solving before launch? I had an issue solving before launch? I had an issue with the cloud environments. I didn't with the cloud environments. I didn't with the cloud environments. I didn't even set up yet. I was doing this all even set up yet. I was doing this all even set up yet. I was doing this all remote. I think I was actually doing it remote. I think I was actually doing it remote. I think I was actually doing it from my phone if I recall. So, I set all from my phone if I recall. So, I set all from my phone if I recall. So, I set all of that. It went through the PR. It said of that. It went through the PR. It said of that. It went through the PR. It said that PR92 does not close the report. It that PR92 does not close the report. It that PR92 does not close the report. It closes two hard stops cleanly and one closes two hard stops cleanly and one closes two hard stops cleanly and one only partially. Most of the launch gates only partially. Most of the launch gates only partially. Most of the launch gates are still open. It said explicitly what are still open. It said explicitly what are still open. It said explicitly what is resolved. Gave some caveats and then is resolved. Gave some caveats and then is resolved. Gave some caveats and then gave me a list of things that were still gave me a list of things that were still gave me a list of things that were still open and should not be treated as done.
-
open and should not be treated as done. open and should not be treated as done. All of this is great. This is a really All of this is great. This is a really All of this is great. This is a really good way to process and synthesize and good way to process and synthesize and good way to process and synthesize and give me this info. It feels very opacy give me this info. It feels very opacy give me this info. It feels very opacy is how I would put it. So I end up is how I would put it. So I end up is how I would put it. So I end up having questions because it was a lot of having questions because it was a lot of having questions because it was a lot of text. So, I went through it all. I text. So, I went through it all. I text. So, I went through it all. I called out a couple different pieces. I called out a couple different pieces. I called out a couple different pieces. I grabbed this section and I only cared grabbed this section and I only cared grabbed this section and I only cared about issues three and five because I about issues three and five because I about issues three and five because I had questions about what it meant by had questions about what it meant by had questions about what it meant by these things. And then after it had a these things. And then after it had a these things. And then after it had a different section with a lot more stuff. different section with a lot more stuff. different section with a lot more stuff. These types of things could be confusing These types of things could be confusing These types of things could be confusing for models because there was two lists for models because there was two lists for models because there was two lists that had a number three and a number that had a number three and a number that had a number three and a number five in them. So, I give a tiny prefix five in them. So, I give a tiny prefix five in them. So, I give a tiny prefix of which list I'm referring to. I of which list I'm referring to. I of which list I'm referring to. I thought this might trip it up because thought this might trip it up because thought this might trip it up because that's just a lot of context to get that's just a lot of context to get that's just a lot of context to get through. And I'll be real, a lot of like through. And I'll be real, a lot of like through. And I'll be real, a lot of like the openw weight and frankly not the openw weight and frankly not the openw weight and frankly not Frontier stuff tends to struggle once Frontier stuff tends to struggle once Frontier stuff tends to struggle once you get it to this point. And it did a you get it to this point. And it did a you get it to this point. And it did a great job. It addressed all of my great job. It addressed all of my great job. It addressed all of my concerns very well and very directly. It concerns very well and very directly. It concerns very well and very directly. It called out the split release thing and called out the split release thing and called out the split release thing and described what it meant. Called out what described what it meant. Called out what described what it meant. Called out what it thinks I should put on security MD. it thinks I should put on security MD. it thinks I should put on security MD. And then it went through all of these And then it went through all of these And then it went through all of these different issues that I had questions different issues that I had questions different issues that I had questions about and helped me prioritize them. I about and helped me prioritize them. I about and helped me prioritize them. I gave it a little bit more feedback on my gave it a little bit more feedback on my gave it a little bit more feedback on my thoughts there. But specifically here is thoughts there. But specifically here is thoughts there. But specifically here is where it gets fun. I called it I don't where it gets fun. I called it I don't where it gets fun. I called it I don't care about the token and URL issue but care about the token and URL issue but care about the token and URL issue but issue two which was bound public work. I issue two which was bound public work. I issue two which was bound public work. I said I like their ideas for it. Make a said I like their ideas for it. Make a said I like their ideas for it. Make a separate PR that addresses all of the separate PR that addresses all of the separate PR that addresses all of the concerns it raised. Then I asked it to concerns it raised. Then I asked it to concerns it raised. Then I asked it to do an investigation for this part and do an investigation for this part and do an investigation for this part and then I asked it to hold on to other then I asked it to hold on to other then I asked it to hold on to other things. It made two PRs because there things. It made two PRs because there things. It made two PRs because there was two issues it was told to address was two issues it was told to address was two issues it was told to address and it addressed both of them and did a and it addressed both of them and did a and it addressed both of them and did a great job. It also answered my great job. It also answered my great job. It also answered my questions. Remember this this is a lot
-
questions. Remember this this is a lot questions. Remember this this is a lot of different pieces I asked for here. as of different pieces I asked for here. as of different pieces I asked for here. as it do a PR for splitting up the actions it do a PR for splitting up the actions it do a PR for splitting up the actions earlier and a separate PR here for the earlier and a separate PR here for the earlier and a separate PR here for the public work bindings that it was public work bindings that it was public work bindings that it was mentioning and it was able to in just mentioning and it was able to in just mentioning and it was able to in just one run make both of those PRs and also one run make both of those PRs and also one run make both of those PRs and also address my other questions and give me a address my other questions and give me a address my other questions and give me a to-do list on all the things I have to to-do list on all the things I have to to-do list on all the things I have to go do before we go live. I noticed there go do before we go live. I noticed there go do before we go live. I noticed there were some comments on the PR so I asked were some comments on the PR so I asked were some comments on the PR so I asked it to look I just said both PRs have it to look I just said both PRs have it to look I just said both PRs have review comments to address and it review comments to address and it review comments to address and it addressed them all. I then used cursor's addressed them all. I then used cursor's addressed them all. I then used cursor's built-in babysit skill to monitor both built-in babysit skill to monitor both built-in babysit skill to monitor both of the PRs and continue addressing of the PRs and continue addressing of the PRs and continue addressing things that come up on them, and it things that come up on them, and it things that come up on them, and it succeeded. I then asked another agent succeeded. I then asked another agent succeeded. I then asked another agent about these PRs, and they said a couple about these PRs, and they said a couple about these PRs, and they said a couple different things should be fixed. So, I different things should be fixed. So, I different things should be fixed. So, I literally just screenshotted it, both to literally just screenshotted it, both to literally just screenshotted it, both to test its visual capabilities, but also test its visual capabilities, but also test its visual capabilities, but also out of laziness. Pasted the screenshot out of laziness. Pasted the screenshot out of laziness. Pasted the screenshot and said, "Can you make these changes to and said, "Can you make these changes to and said, "Can you make these changes to 93?" And it did. This is great. This is 93?" And it did. This is great. This is 93?" And it did. This is great. This is the type of back and forth and complex the type of back and forth and complex the type of back and forth and complex multi-target multi-target multi-target tasks and work that you could make most tasks and work that you could make most tasks and work that you could make most models do if you break them up really models do if you break them up really models do if you break them up really carefully and cleanly and try to not carefully and cleanly and try to not carefully and cleanly and try to not bloat the context. Even a model like 55 bloat the context. Even a model like 55 bloat the context. Even a model like 55 in my experience would get confused at in my experience would get confused at in my experience would get confused at some point during this and get too some point during this and get too some point during this and get too fixated on something from three messages fixated on something from three messages fixated on something from three messages ago instead of doing what I'm asking ago instead of doing what I'm asking ago instead of doing what I'm asking now. I had no ergonomic issues with four now. I had no ergonomic issues with four now. I had no ergonomic issues with four five here at all. It was actually very five here at all. It was actually very five here at all. It was actually very pleasant to work with. I kept making my pleasant to work with. I kept making my pleasant to work with. I kept making my responses like worse and worse almost responses like worse and worse almost responses like worse and worse almost intentionally just to see where it would intentionally just to see where it would intentionally just to see where it would stumble and it didn't. Did it find stumble and it didn't. Did it find stumble and it didn't. Did it find everything Fable found? No. Was its code everything Fable found? No. Was its code everything Fable found? No. Was its code as thorough as 56? No. Was it able to go
-
as thorough as 56? No. Was it able to go as thorough as 56? No. Was it able to go back and forth with me on really big, back and forth with me on really big, back and forth with me on really big, heavy tasks and be pleasant to use while heavy tasks and be pleasant to use while heavy tasks and be pleasant to use while also being very fast and cheap? also being very fast and cheap? also being very fast and cheap? Absolutely. Yeah. In a lot of ways, I Absolutely. Yeah. In a lot of ways, I Absolutely. Yeah. In a lot of ways, I see this model as a good alternative to see this model as a good alternative to see this model as a good alternative to something like Opus 48. and it kind of something like Opus 48. and it kind of something like Opus 48. and it kind of shits all over something like I don't shits all over something like I don't shits all over something like I don't know GLM52. I have no interest in that know GLM52. I have no interest in that know GLM52. I have no interest in that model anymore other than the fact that model anymore other than the fact that model anymore other than the fact that it's openweight which is genuinely it's openweight which is genuinely it's openweight which is genuinely really cool and I appreciate them for really cool and I appreciate them for really cool and I appreciate them for that. But Grock 45 is a weirdly good that. But Grock 45 is a weirdly good that. But Grock 45 is a weirdly good default code model. So I wanted to push default code model. So I wanted to push default code model. So I wanted to push further and I did. I took a game I built further and I did. I took a game I built further and I did. I took a game I built previously, my little fish slop thing, previously, my little fish slop thing, previously, my little fish slop thing, and I asked it to make the game in 3D. and I asked it to make the game in 3D. and I asked it to make the game in 3D. It did and it has some issues like this It did and it has some issues like this It did and it has some issues like this layout is broken because of the model. I layout is broken because of the model. I layout is broken because of the model. I didn't choose for it to be panned hard didn't choose for it to be panned hard didn't choose for it to be panned hard to the side like this. Wait till you see to the side like this. Wait till you see to the side like this. Wait till you see what happens when we open the game. It what happens when we open the game. It what happens when we open the game. It successfully made a full 3D environment, successfully made a full 3D environment, successfully made a full 3D environment, modeled all of the different creatures modeled all of the different creatures modeled all of the different creatures itself, as well as all of the geometry itself, as well as all of the geometry itself, as well as all of the geometry of the things on the bottom of the tank. of the things on the bottom of the tank. of the things on the bottom of the tank. And it did a way better job with that And it did a way better job with that And it did a way better job with that than any other model I've used. Even than any other model I've used. Even than any other model I've used. Even models like Fable and 56, which I have models like Fable and 56, which I have models like Fable and 56, which I have given access to 3D modeling tools as given access to 3D modeling tools as given access to 3D modeling tools as well. This crushed all of them.
-
well. This crushed all of them. well. This crushed all of them. Obviously, the models are far from Obviously, the models are far from Obviously, the models are far from great, especially the creature models great, especially the creature models great, especially the creature models and the model for the submarine. And it and the model for the submarine. And it and the model for the submarine. And it also got the controls super wrong where also got the controls super wrong where also got the controls super wrong where A goes right and D goes left for some A goes right and D goes left for some A goes right and D goes left for some reason. They Yeah, they screwed these reason. They Yeah, they screwed these reason. They Yeah, they screwed these things up. And it also just doesn't things up. And it also just doesn't things up. And it also just doesn't control very cleanly. I gave it a control very cleanly. I gave it a control very cleanly. I gave it a follow-up telling it to fix it, and it follow-up telling it to fix it, and it follow-up telling it to fix it, and it fixed the pointer and the clicking, but fixed the pointer and the clicking, but fixed the pointer and the clicking, but it didn't fix a lot of the rest. it didn't fix a lot of the rest. it didn't fix a lot of the rest. One of the models I was really impressed One of the models I was really impressed One of the models I was really impressed with in here though was the enemy model with in here though was the enemy model with in here though was the enemy model for the aliens that invade the tank. It for the aliens that invade the tank. It for the aliens that invade the tank. It did a very good job with that. But yeah, did a very good job with that. But yeah, did a very good job with that. But yeah, I did not expect this to have I did not expect this to have I did not expect this to have three-dimensional taste. And I'll say three-dimensional taste. And I'll say three-dimensional taste. And I'll say outright, a lot of these, while far from outright, a lot of these, while far from outright, a lot of these, while far from like I would chip this confidently tier, like I would chip this confidently tier, like I would chip this confidently tier, are so far ahead anything I've gotten are so far ahead anything I've gotten are so far ahead anything I've gotten from any of the other models that I can from any of the other models that I can from any of the other models that I can confidently say this is the first model confidently say this is the first model confidently say this is the first model to be almost decent at 3D modeling in to be almost decent at 3D modeling in to be almost decent at 3D modeling in game engines like 3JS. Take it as you game engines like 3JS. Take it as you game engines like 3JS. Take it as you will, this is not bad. I I am more will, this is not bad. I I am more will, this is not bad. I I am more impressed than I expected to be impressed than I expected to be impressed than I expected to be regarding a Grock model's 3D regarding a Grock model's 3D regarding a Grock model's 3D capabilities. I do have one last thought capabilities. I do have one last thought capabilities. I do have one last thought I want to discuss here, which is this I want to discuss here, which is this I want to discuss here, which is this idea of model generations. I do really idea of model generations. I do really idea of model generations. I do really feel like Fable's a new generation of feel like Fable's a new generation of feel like Fable's a new generation of model and I have a lot of thoughts on model and I have a lot of thoughts on model and I have a lot of thoughts on where GBD56 falls here. Video probably where GBD56 falls here. Video probably where GBD56 falls here. Video probably coming tomorrow depending on a lot.
-
coming tomorrow depending on a lot. coming tomorrow depending on a lot. We'll see. The reason I bring this up is We'll see. The reason I bring this up is We'll see. The reason I bring this up is because Fable and 56 in many ways feel because Fable and 56 in many ways feel because Fable and 56 in many ways feel very different. And as silly as it is, very different. And as silly as it is, very different. And as silly as it is, even something like Sonet 5 has a bit of even something like Sonet 5 has a bit of even something like Sonet 5 has a bit of that feeling, too. The difference is in that feeling, too. The difference is in that feeling, too. The difference is in the model's ability to orchestrate. They the model's ability to orchestrate. They the model's ability to orchestrate. They can step up a level and prompt sub can step up a level and prompt sub can step up a level and prompt sub agents and orchestrate big work into agents and orchestrate big work into agents and orchestrate big work into lots of smaller chunks to go longer, do lots of smaller chunks to go longer, do lots of smaller chunks to go longer, do more, and complete more difficult tasks. more, and complete more difficult tasks. more, and complete more difficult tasks. This is why I got so addicted both to This is why I got so addicted both to This is why I got so addicted both to Fable and to GBT56. I don't necessarily Fable and to GBT56. I don't necessarily Fable and to GBT56. I don't necessarily see that capability in Gro 4.5. My see that capability in Gro 4.5. My see that capability in Gro 4.5. My attempts to break up sub agents with it attempts to break up sub agents with it attempts to break up sub agents with it were admittedly limited, but I was not were admittedly limited, but I was not were admittedly limited, but I was not super impressed. It didn't seem to have super impressed. It didn't seem to have super impressed. It didn't seem to have the same nuance in how it would break the same nuance in how it would break the same nuance in how it would break work up and delegate it and it would work up and delegate it and it would work up and delegate it and it would often get stuck as a result of a certain often get stuck as a result of a certain often get stuck as a result of a certain process it ran hanging and then not process it ran hanging and then not process it ran hanging and then not knowing how to clean up after. Some of knowing how to clean up after. Some of knowing how to clean up after. Some of this is harness specific, some of this this is harness specific, some of this this is harness specific, some of this is model specific, some of this is just is model specific, some of this is just is model specific, some of this is just being behind. But I would say in many being behind. But I would say in many being behind. But I would say in many ways what they built with Gro 4.5 is ways what they built with Gro 4.5 is ways what they built with Gro 4.5 is less a comparable model to this new less a comparable model to this new less a comparable model to this new generation with things like Fable and generation with things like Fable and generation with things like Fable and GPD56. More it's really really GPD56. More it's really really GPD56. More it's really really impressive what they did with the impressive what they did with the impressive what they did with the previous technology. If you were to previous technology. If you were to previous technology. If you were to think of this in gaming, for example, think of this in gaming, for example, think of this in gaming, for example, they just put out the best PS2 game they just put out the best PS2 game they just put out the best PS2 game ever, but the PS3 has been out for two ever, but the PS3 has been out for two ever, but the PS3 has been out for two months. And as such, I'm very excited to months. And as such, I'm very excited to months. And as such, I'm very excited to see if they can level up to a PS3 see if they can level up to a PS3 see if they can level up to a PS3 generation game. And I think they can.
-
generation game. And I think they can. generation game. And I think they can. They have all of the pieces. And the They have all of the pieces. And the They have all of the pieces. And the amount that they just jumped in such a amount that they just jumped in such a amount that they just jumped in such a short window is truly a sight to behold. short window is truly a sight to behold. short window is truly a sight to behold. I don't think any lab has had a jump I don't think any lab has had a jump I don't think any lab has had a jump like this ever, other than maybe like this ever, other than maybe like this ever, other than maybe arguably like deepseek. going from arguably like deepseek. going from arguably like deepseek. going from forgotten in the conversation and just forgotten in the conversation and just forgotten in the conversation and just reselling GPUs to beating out reselling GPUs to beating out reselling GPUs to beating out practically everyone else in the space practically everyone else in the space practically everyone else in the space at the tier you're competing in for at the tier you're competing in for at the tier you're competing in for cheaper and more efficient models. It's cheaper and more efficient models. It's cheaper and more efficient models. It's impressive and we shouldn't be sleeping impressive and we shouldn't be sleeping impressive and we shouldn't be sleeping on SpaceX anymore. I posted in April on SpaceX anymore. I posted in April on SpaceX anymore. I posted in April that I legitimately believed XAI could that I legitimately believed XAI could that I legitimately believed XAI could have a crazy comeback. I thought it have a crazy comeback. I thought it have a crazy comeback. I thought it would take 6 months to a year. I didn't would take 6 months to a year. I didn't would take 6 months to a year. I didn't think it would take two and a half. I am think it would take two and a half. I am think it would take two and a half. I am blown away. These guys are cooking blown away. These guys are cooking blown away. These guys are cooking again. I have no idea where this will again. I have no idea where this will again. I have no idea where this will all end up, but I'm thankful somebody all end up, but I'm thankful somebody all end up, but I'm thankful somebody other than Google is actually competing other than Google is actually competing other than Google is actually competing now. This is the first real player that now. This is the first real player that now. This is the first real player that Anthropic and OpenAI have had to be Anthropic and OpenAI have had to be Anthropic and OpenAI have had to be scared of in quite a while. And I hope scared of in quite a while. And I hope scared of in quite a while. And I hope they are. I hope the result of Gro 45 is they are. I hope the result of Gro 45 is they are. I hope the result of Gro 45 is faster, cheaper, and smarter models for faster, cheaper, and smarter models for faster, cheaper, and smarter models for everyone because that's what we want in everyone because that's what we want in everyone because that's what we want in the end. So, congratulations to SpaceX the end. So, congratulations to SpaceX the end. So, congratulations to SpaceX AAI for catching up. I didn't think you AAI for catching up. I didn't think you AAI for catching up. I didn't think you had it in you, but clearly you do. I had it in you, but clearly you do. I had it in you, but clearly you do. I can't wait to see what comes out next.
-
can't wait to see what comes out next. can't wait to see what comes out next. And until next time, peace nerds.
Summary
The transcript highlights the impressive performance of Space XAI's Grok 4.5 model, particularly in developer work, comparing favorably to established models like Fable 5 and GPT-55. Benchmarks from the artificial analysis code index suggest it is a highly capable and cost-effective AI tool, especially with current discounts. The speaker also humorously contrasts AI's power with a personal failure of LG's website to notify about a product release.