← Back
Theo August 7, 2026 44m

Meta's Claude Code clone is INSANELY cheap

Read full transcript 34 segments
  1. I have something I'm a bit ashamed to I have something I'm a bit ashamed to admit. I'm a pretty big fan of Meta. Not admit. I'm a pretty big fan of Meta. Not admit. I'm a pretty big fan of Meta. Not of things like Facebook, like no, gross. of things like Facebook, like no, gross. of things like Facebook, like no, gross. Not my thing at all. But when it comes Not my thing at all. But when it comes Not my thing at all. But when it comes to their actual software contributions to their actual software contributions to their actual software contributions and things they do for the ecosystem, I and things they do for the ecosystem, I and things they do for the ecosystem, I found them to be pretty solid, if not found them to be pretty solid, if not found them to be pretty solid, if not great. Projects like React and React great. Projects like React and React great. Projects like React and React Native were essential to the growth of Native were essential to the growth of Native were essential to the growth of the web and to an extent mobile as well, the web and to an extent mobile as well, the web and to an extent mobile as well, and are genuinely incredible things that and are genuinely incredible things that and are genuinely incredible things that they've put out there for free for they've put out there for free for they've put out there for free for people to use however they want. Their people to use however they want. Their people to use however they want. Their contributions to the AR and VR ecosystem contributions to the AR and VR ecosystem contributions to the AR and VR ecosystem were all things that I enjoyed heavily were all things that I enjoyed heavily were all things that I enjoyed heavily at the time because I'm a big VR nerd. at the time because I'm a big VR nerd. at the time because I'm a big VR nerd. And when they started getting into AI, And when they started getting into AI, And when they started getting into AI, they did that similarly well with their they did that similarly well with their they did that similarly well with their focus on openweight models with the focus on openweight models with the focus on openweight models with the llama line. To this day, llama is still llama line. To this day, llama is still llama line. To this day, llama is still used almost like a generic term for used almost like a generic term for used almost like a generic term for openweight models. And it's crazy how openweight models. And it's crazy how openweight models. And it's crazy how far the llama models went, but it's even far the llama models went, but it's even far the llama models went, but it's even crazier how far the entire industry has crazier how far the entire industry has crazier how far the entire industry has moved ahead of meta. They've been trying moved ahead of meta. They've been trying moved ahead of meta. They've been trying to correct course for a while with this to correct course for a while with this to correct course for a while with this hidden quietly worked on line of models hidden quietly worked on line of models hidden quietly worked on line of models called Muse. And historically, they've called Muse. And historically, they've called Muse. And historically, they've not been accessible beyond their little not been accessible beyond their little not been accessible beyond their little web interface to try it out. But web interface to try it out. But web interface to try it out. But apparently, they've been trying to make apparently, they've been trying to make apparently, they've been trying to make it ready for code. Not just the model, it ready for code. Not just the model, it ready for code. Not just the model, but the tools around it as well. And now but the tools around it as well. And now but the tools around it as well. And now they're finally ready to release.

  2. they're finally ready to release. they're finally ready to release. Zuckerberg just announced on Twitter, by Zuckerberg just announced on Twitter, by Zuckerberg just announced on Twitter, by the way, that Muse Code is out in beta the way, that Muse Code is out in beta the way, that Muse Code is out in beta today. This is their clone of Claude today. This is their clone of Claude today. This is their clone of Claude Code and it's powered by Muse Spark 1.2, Code and it's powered by Muse Spark 1.2, Code and it's powered by Muse Spark 1.2, which is their new coding focused model. which is their new coding focused model. which is their new coding focused model. It's been a while since I've seen a lab It's been a while since I've seen a lab It's been a while since I've seen a lab publish benchmark numbers where they're publish benchmark numbers where they're publish benchmark numbers where they're not in first place in anything, but that not in first place in anything, but that not in first place in anything, but that doesn't mean it's a bad model. As you doesn't mean it's a bad model. As you doesn't mean it's a bad model. As you guys probably saw in my video about Gro guys probably saw in my video about Gro guys probably saw in my video about Gro 4.5, there is still a lot of value in 4.5, there is still a lot of value in 4.5, there is still a lot of value in wellpriced, fast, useful models. And wellpriced, fast, useful models. And wellpriced, fast, useful models. And with Meta's focus on using every model with Meta's focus on using every model with Meta's focus on using every model for everything lately and collecting all for everything lately and collecting all for everything lately and collecting all the data they get as a result, it's the data they get as a result, it's the data they get as a result, it's actually going to be pretty interesting actually going to be pretty interesting actually going to be pretty interesting to see how well this performs. So, I'm to see how well this performs. So, I'm to see how well this performs. So, I'm going to go through this with y'all going to go through this with y'all going to go through this with y'all together. I'm going to try it live. I'm together. I'm going to try it live. I'm together. I'm going to try it live. I'm going to read through what they have to going to read through what they have to going to read through what they have to say. I'm going to see what others are say. I'm going to see what others are say. I'm going to see what others are doing with it, get a feel for the price doing with it, get a feel for the price doing with it, get a feel for the price and the weird quirks as well. And I just and the weird quirks as well. And I just and the weird quirks as well. And I just had to sign into my terminal with had to sign into my terminal with had to sign into my terminal with Facebook. So, I'm feeling a little Facebook. So, I'm feeling a little Facebook. So, I'm feeling a little gross. This is going to be a journey and gross. This is going to be a journey and gross. This is going to be a journey and I hope you enjoy it with me. But first, I hope you enjoy it with me. But first, I hope you enjoy it with me. But first, a quick break for today's sponsor. a quick break for today's sponsor. a quick break for today's sponsor. There's a lot of code review bots around There's a lot of code review bots around There's a lot of code review bots around nowadays and they're pretty good at nowadays and they're pretty good at nowadays and they're pretty good at reading the code, but I've noticed that reading the code, but I've noticed that reading the code, but I've noticed that running cloud code and codecs on my running cloud code and codecs on my running cloud code and codecs on my machine tends to be better simply machine tends to be better simply machine tends to be better simply because the models can actually run because the models can actually run because the models can actually run things and verify the changes. Wouldn't things and verify the changes. Wouldn't things and verify the changes. Wouldn't it be great if those review bots had the it be great if those review bots had the it be great if those review bots had the ability to actually verify changes on a ability to actually verify changes on a ability to actually verify changes on a real computer? Reptile thought so, and real computer? Reptile thought so, and real computer? Reptile thought so, and that's why they introduced T-Rex. It's a that's why they introduced T-Rex. It's a that's why they introduced T-Rex. It's a new sandbox that can actually run your new sandbox that can actually run your new sandbox that can actually run your code and verify the changes rather than code and verify the changes rather than code and verify the changes rather than just reading the syntax and hoping it's just reading the syntax and hoping it's just reading the syntax and hoping it's good. If you've ever added an element to good. If you've ever added an element to good. If you've ever added an element to your UI and had it overlap something your UI and had it overlap something your UI and had it overlap something that was obvious when you ran it that that was obvious when you ran it that that was obvious when you ran it that the code review bots missed, that's what the code review bots missed, that's what the code review bots missed, that's what this is for. Because the models can this is for. Because the models can this is for. Because the models can actually click through and test things, actually click through and test things, actually click through and test things, and it can even respond with screenshots

  3. and it can even respond with screenshots and it can even respond with screenshots of the things it finds along the way. of the things it finds along the way. of the things it finds along the way. Their website lists a bunch of realworld Their website lists a bunch of realworld Their website lists a bunch of realworld open source projects that have had bugs open source projects that have had bugs open source projects that have had bugs prevented through T-Rex. Even companies prevented through T-Rex. Even companies prevented through T-Rex. Even companies like Work OS are finding real bugs. For like Work OS are finding real bugs. For like Work OS are finding real bugs. For example, this code that seems like a example, this code that seems like a example, this code that seems like a totally safe filter, but has a really totally safe filter, but has a really totally safe filter, but has a really rough edge case that you'll only find rough edge case that you'll only find rough edge case that you'll only find when you run it on real data. If you've when you run it on real data. If you've when you run it on real data. If you've debug in the last 3 months, you owe it debug in the last 3 months, you owe it debug in the last 3 months, you owe it to your team to check out to your team to check out to your team to check out soyv.link/gretile. soyv.link/gretile. soyv.link/gretile. So, let's start with what Zuckerberg had So, let's start with what Zuckerberg had So, let's start with what Zuckerberg had to say about their most recent release. to say about their most recent release. to say about their most recent release. Releasing Muse Code in beta today. It's Releasing Muse Code in beta today. It's Releasing Muse Code in beta today. It's a terminal coding agent that takes on a terminal coding agent that takes on a terminal coding agent that takes on complex software engineering tasks complex software engineering tasks complex software engineering tasks across large repos, lining changes, across large repos, lining changes, across large repos, lining changes, writing code, validating the results. writing code, validating the results. writing code, validating the results. It's powered by Muse Spark 1.2, which is It's powered by Muse Spark 1.2, which is It's powered by Muse Spark 1.2, which is a coding focused model update. This is a coding focused model update. This is a coding focused model update. This is going to be an interesting one because a going to be an interesting one because a going to be an interesting one because a lot of the other legacy companies like lot of the other legacy companies like lot of the other legacy companies like uh I don't know Microsoft and Google uh I don't know Microsoft and Google uh I don't know Microsoft and Google don't necessarily have the best stack don't necessarily have the best stack don't necessarily have the best stack internally, whereas Facebook has done internally, whereas Facebook has done internally, whereas Facebook has done some incredible things with their some incredible things with their some incredible things with their internal tech. They didn't use git internal tech. They didn't use git internal tech. They didn't use git because it was too slow for the scale because it was too slow for the scale because it was too slow for the scale they were moving at. So they built their they were moving at. So they built their they were moving at. So they built their own custom everything on top of own custom everything on top of own custom everything on top of Mercurial and it's really powerful to Mercurial and it's really powerful to Mercurial and it's really powerful to the point where people have been trying the point where people have been trying the point where people have been trying to copy their workflows with things like to copy their workflows with things like to copy their workflows with things like stacked PRs and stacked diffs. I know a stacked PRs and stacked diffs. I know a stacked PRs and stacked diffs. I know a lot of people who used to work at Meta lot of people who used to work at Meta lot of people who used to work at Meta and left and just missed all of that and left and just missed all of that and left and just missed all of that tooling that made it so much easier to tooling that made it so much easier to tooling that made it so much easier to work in these gigantic projects that are work in these gigantic projects that are work in these gigantic projects that are met scale. Other companies like Google met scale. Other companies like Google met scale. Other companies like Google also have gigantic monor repos, but they also have gigantic monor repos, but they also have gigantic monor repos, but they are less focused on fixing them. In are less focused on fixing them. In are less focused on fixing them. In fact, they've went as far as laying off fact, they've went as far as laying off fact, they've went as far as laying off the teams that built the tooling to make the teams that built the tooling to make the teams that built the tooling to make them usable. Meta's always done the them usable. Meta's always done the them usable. Meta's always done the opposite where they'll rewrite a opposite where they'll rewrite a opposite where they'll rewrite a language if they have to, like they language if they have to, like they language if they have to, like they rewrote PHP and hack in order to try and rewrote PHP and hack in order to try and rewrote PHP and hack in order to try and make their code bases work better. Since

  4. make their code bases work better. Since make their code bases work better. Since Meta's code bases are so huge, it's Meta's code bases are so huge, it's Meta's code bases are so huge, it's often easier to rewrite the language often easier to rewrite the language often easier to rewrite the language that's underneath them than it is to that's underneath them than it is to that's underneath them than it is to rewrite the codebase itself on a better rewrite the codebase itself on a better rewrite the codebase itself on a better language. And that type of thinking is, language. And that type of thinking is, language. And that type of thinking is, at least it was unique to Meta for a at least it was unique to Meta for a at least it was unique to Meta for a long time. And I think it kind of long time. And I think it kind of long time. And I think it kind of positions them well to dive into this positions them well to dive into this positions them well to dive into this world of using models and tools to work world of using models and tools to work world of using models and tools to work with your code bases at that scale. And with your code bases at that scale. And with your code bases at that scale. And that's why they've been using models that's why they've been using models that's why they've been using models like Fable and Opus so heavily recently. like Fable and Opus so heavily recently. like Fable and Opus so heavily recently. So, it's interesting to see what they're So, it's interesting to see what they're So, it's interesting to see what they're cooking here. They shared some benches cooking here. They shared some benches cooking here. They shared some benches here. And the first notable thing is here. And the first notable thing is here. And the first notable thing is that they're not putting Fable or Soul that they're not putting Fable or Soul that they're not putting Fable or Soul in these lists. They're only putting in these lists. They're only putting in these lists. They're only putting Opus and 56 Terra in, which is a choice. Opus and 56 Terra in, which is a choice. Opus and 56 Terra in, which is a choice. In Terminal Bench 2.1, they slightly In Terminal Bench 2.1, they slightly In Terminal Bench 2.1, they slightly beat out Terra and are slightly behind beat out Terra and are slightly behind beat out Terra and are slightly behind Opus 5. There also a big jump from US Opus 5. There also a big jump from US Opus 5. There also a big jump from US Spark 1.1 to 1.2. 1.1 was barely Spark 1.1 to 1.2. 1.1 was barely Spark 1.1 to 1.2. 1.1 was barely available to the public. 1.2 2 is now available to the public. 1.2 2 is now available to the public. 1.2 2 is now actually out and the pricing is really actually out and the pricing is really actually out and the pricing is really interesting which we'll show in a little interesting which we'll show in a little interesting which we'll show in a little bit. Deep SWE which is one of the bit. Deep SWE which is one of the bit. Deep SWE which is one of the benchmarks I prefer. It did pretty well benchmarks I prefer. It did pretty well benchmarks I prefer. It did pretty well here getting a slightly higher score here getting a slightly higher score here getting a slightly higher score than Grock 4.5 which is nuts because I than Grock 4.5 which is nuts because I than Grock 4.5 which is nuts because I actually found Grock 45 to be a very actually found Grock 45 to be a very actually found Grock 45 to be a very pleasant to use model but it's still pleasant to use model but it's still pleasant to use model but it's still lagging behind Terra and Opus which lagging behind Terra and Opus which lagging behind Terra and Opus which means it's far behind models like Soul means it's far behind models like Soul means it's far behind models like Soul and Fable. And then Meta's internal and Fable. And then Meta's internal and Fable. And then Meta's internal coding bench, it came in second place in coding bench, it came in second place in coding bench, it came in second place in with Muse Spark being much closer in with Muse Spark being much closer in with Muse Spark being much closer in third. And if this model is as far off third. And if this model is as far off third. And if this model is as far off as other benches showed, I think their as other benches showed, I think their as other benches showed, I think their internal bench might not be the best.

  5. internal bench might not be the best. internal bench might not be the best. They also in their internal bench have They also in their internal bench have They also in their internal bench have Terra in Gemini 36 Flash suspiciously Terra in Gemini 36 Flash suspiciously Terra in Gemini 36 Flash suspiciously close. So, we shall see as we go. Oh close. So, we shall see as we go. Oh close. So, we shall see as we go. Oh god, they're already getting community god, they're already getting community god, they're already getting community noted. Community note: Meta is still noted. Community note: Meta is still noted. Community note: Meta is still dead last in AI with this release. Okay, dead last in AI with this release. Okay, dead last in AI with this release. Okay, maybe not. Apple is, but they're behind maybe not. Apple is, but they're behind maybe not. Apple is, but they're behind just as much as Google Gemini is. That's just as much as Google Gemini is. That's just as much as Google Gemini is. That's not a thing you need to put in a not a thing you need to put in a not a thing you need to put in a community note, nerds. Anyways, community note, nerds. Anyways, community note, nerds. Anyways, according to Zuck, Muse Code runs according to Zuck, Muse Code runs according to Zuck, Muse Code runs specialized background agents that stay specialized background agents that stay specialized background agents that stay active your whole session. So, they active your whole session. So, they active your whole session. So, they build up context over time instead of build up context over time instead of build up context over time instead of starting from scratch on every task. starting from scratch on every task. starting from scratch on every task. Interesting. When a job is big enough, Interesting. When a job is big enough, Interesting. When a job is big enough, it fans out the separate sub agents it fans out the separate sub agents it fans out the separate sub agents working in parallel in isolated work working in parallel in isolated work working in parallel in isolated work trees. Your working copies never trees. Your working copies never trees. Your working copies never touched. In testing, we had it build six touched. In testing, we had it build six touched. In testing, we had it build six features for a game simultaneously with features for a game simultaneously with features for a game simultaneously with no collisions. Huh, very interesting. no collisions. Huh, very interesting. no collisions. Huh, very interesting. Again, big codebase stuff. We pointed Again, big codebase stuff. We pointed Again, big codebase stuff. We pointed Muse Spark 1.2 at a kernel optimization Muse Spark 1.2 at a kernel optimization Muse Spark 1.2 at a kernel optimization task and let it run. a thousand tool task and let it run. a thousand tool task and let it run. a thousand tool calls over 24 hours on Nvidia Hopper. It calls over 24 hours on Nvidia Hopper. It calls over 24 hours on Nvidia Hopper. It kept finding substantial improvements kept finding substantial improvements kept finding substantial improvements well beyond the initial exploration well beyond the initial exploration well beyond the initial exploration phase. This is also kind of funny to me phase. This is also kind of funny to me phase. This is also kind of funny to me because Meta rented a bunch of TPUs from because Meta rented a bunch of TPUs from because Meta rented a bunch of TPUs from Google. They're one of the few companies Google. They're one of the few companies Google. They're one of the few companies to actually buy and rack hardware from to actually buy and rack hardware from to actually buy and rack hardware from Google and then Google went and rented a Google and then Google went and rented a Google and then Google went and rented a bunch of GPUs from Nvidia, all from XAI bunch of GPUs from Nvidia, all from XAI bunch of GPUs from Nvidia, all from XAI and SpaceX. So Google is reselling their and SpaceX. So Google is reselling their and SpaceX. So Google is reselling their shitty chips to Meta so they can use the shitty chips to Meta so they can use the shitty chips to Meta so they can use the money to go buy better chips from Nvidia money to go buy better chips from Nvidia money to go buy better chips from Nvidia through XAI. And now apparently through XAI. And now apparently through XAI. And now apparently Zuckerberg is still stuck on Nvidia.

  6. Zuckerberg is still stuck on Nvidia. Zuckerberg is still stuck on Nvidia. Auditable by design, every model call, Auditable by design, every model call, Auditable by design, every model call, tool run, and edit has a local event log tool run, and edit has a local event log tool run, and edit has a local event log before it executes. If it crashes mid before it executes. If it crashes mid before it executes. If it crashes mid task, it picks up exactly where it left task, it picks up exactly where it left task, it picks up exactly where it left off from that log. No lost work and no off from that log. No lost work and no off from that log. No lost work and no reprompting. Pricing. It's easy and low reprompting. Pricing. It's easy and low reprompting. Pricing. It's easy and low cost to get started. Install Muse Code cost to get started. Install Muse Code cost to get started. Install Muse Code with one line and you can start on our with one line and you can start on our with one line and you can start on our contributor tier. Contributor tier is an contributor tier. Contributor tier is an contributor tier. Contributor tier is an interesting piece we'll talk about in a interesting piece we'll talk about in a interesting piece we'll talk about in a second. Musepark 2 is our next step as second. Musepark 2 is our next step as second. Musepark 2 is our next step as we push towards Front Tier with larger, we push towards Front Tier with larger, we push towards Front Tier with larger, more capable models on the way. Install more capable models on the way. Install more capable models on the way. Install it, use it, and tell us what you think. it, use it, and tell us what you think. it, use it, and tell us what you think. Well, the first thing I think is that Well, the first thing I think is that Well, the first thing I think is that when you Google search Muse Spark, you when you Google search Muse Spark, you when you Google search Muse Spark, you get all of these old articles about 1.1. get all of these old articles about 1.1. get all of these old articles about 1.1. This old developer page. Okay, this one This old developer page. Okay, this one This old developer page. Okay, this one has 1.2 on it now, but it was actually has 1.2 on it now, but it was actually has 1.2 on it now, but it was actually kind of hard and annoying to find this kind of hard and annoying to find this kind of hard and annoying to find this initially. Here's the part I was looking initially. Here's the part I was looking initially. Here's the part I was looking for, though. The price. $1.25 per mill for, though. The price. $1.25 per mill for, though. The price. $1.25 per mill in, 15 cents for cashed 1 mil in, and in, 15 cents for cashed 1 mil in, and in, 15 cents for cashed 1 mil in, and 425 per mill out. Not great pricing 425 per mill out. Not great pricing 425 per mill out. Not great pricing until you realize I was hiding something until you realize I was hiding something until you realize I was hiding something from you. The contributor price tier. from you. The contributor price tier. from you. The contributor price tier. This tier is 10 cents per mill in and 20 This tier is 10 cents per mill in and 20 This tier is 10 cents per mill in and 20 cents per mill out with 0.2 cents per cents per mill out with 0.2 cents per cents per mill out with 0.2 cents per cached input token. That is a 10 to 20x cached input token. That is a 10 to 20x cached input token. That is a 10 to 20x price gap. They're effectively giving price gap. They're effectively giving price gap. They're effectively giving out the contributor version for free out the contributor version for free out the contributor version for free because they need training data so because they need training data so because they need training data so goddamn badly. Meta's released Muse goddamn badly. Meta's released Muse goddamn badly. Meta's released Muse Spark 1.2. It's their third release in Spark 1.2. It's their third release in Spark 1.2. It's their third release in four months and it scores a 54 on the four months and it scores a 54 on the four months and it scores a 54 on the intelligence index, significantly intelligence index, significantly intelligence index, significantly improving agentic knowledge war improving agentic knowledge war improving agentic knowledge war capabilities over prior releases, capabilities over prior releases, capabilities over prior releases, putting meta next to SpaceX in a tie for putting meta next to SpaceX in a tie for putting meta next to SpaceX in a tie for third place amongst US labs. Crazy they third place amongst US labs. Crazy they third place amongst US labs. Crazy they have to specify this now because the have to specify this now because the have to specify this now because the Chinese labs have caught up so much, Chinese labs have caught up so much, Chinese labs have caught up so much, especially with K3. Muse Spark 1.2 lands

  7. especially with K3. Muse Spark 1.2 lands especially with K3. Muse Spark 1.2 lands at a 54, up three points from 1.1 and 11 at a 54, up three points from 1.1 and 11 at a 54, up three points from 1.1 and 11 points from Muse Spark 1.0 which came points from Muse Spark 1.0 which came points from Muse Spark 1.0 which came out in April. That's a pretty big jump. out in April. That's a pretty big jump. out in April. That's a pretty big jump. Like if we're just looking at the Like if we're just looking at the Like if we're just looking at the trajectory of what meta right here, trajectory of what meta right here, trajectory of what meta right here, that's a lot of improvement in not a that's a lot of improvement in not a that's a lot of improvement in not a whole lot of time. It's effectively tied whole lot of time. It's effectively tied whole lot of time. It's effectively tied with GBT 55 and Gro 4.5 narrowly behind with GBT 55 and Gro 4.5 narrowly behind with GBT 55 and Gro 4.5 narrowly behind current frontier models like Opus 5, current frontier models like Opus 5, current frontier models like Opus 5, Fable 5, and 56 Soul as well as Kimmy Fable 5, and 56 Soul as well as Kimmy Fable 5, and 56 Soul as well as Kimmy K3. It's among the most costefficient K3. It's among the most costefficient K3. It's among the most costefficient models at its intelligence level. It's models at its intelligence level. It's models at its intelligence level. It's 40 cents per intelligence index task at 40 cents per intelligence index task at 40 cents per intelligence index task at Meta's unchanged $125 and $425 per Meta's unchanged $125 and $425 per Meta's unchanged $125 and $425 per million token pricing. They don't million token pricing. They don't million token pricing. They don't mention I I was right when I was reading mention I I was right when I was reading mention I I was right when I was reading that earlier. They are not measuring the that earlier. They are not measuring the that earlier. They are not measuring the prices based on the contributor tier. prices based on the contributor tier. prices based on the contributor tier. They are measuring them based on the They are measuring them based on the They are measuring them based on the normal pricing that they would charge normal pricing that they would charge normal pricing that they would charge actually. So if we go back to the actually. So if we go back to the actually. So if we go back to the pricing chart here, it's about a 10 to pricing chart here, it's about a 10 to pricing chart here, it's about a 10 to 20x difference in price, which means 20x difference in price, which means 20x difference in price, which means Muspark 1.2 on the contributor tier is Muspark 1.2 on the contributor tier is Muspark 1.2 on the contributor tier is actually the cheapest model currently actually the cheapest model currently actually the cheapest model currently here because it would be 2 to 3 cents here because it would be 2 to 3 cents here because it would be 2 to 3 cents per task, which is comparable to V4 per task, which is comparable to V4 per task, which is comparable to V4 flash as well as 56 Luna. Very flash as well as 56 Luna. Very flash as well as 56 Luna. Very interesting. This model, especially if interesting. This model, especially if interesting. This model, especially if you're down to like let Meta have your you're down to like let Meta have your you're down to like let Meta have your data, could be a really good value for data, could be a really good value for data, could be a really good value for the short term. Amnition's absention the short term. Amnition's absention the short term. Amnition's absention rates have increased. The model score rates have increased. The model score rates have increased. The model score rose from an 18 to a 22 and the rose from an 18 to a 22 and the rose from an 18 to a 22 and the hallucination rate fell 10 points. The hallucination rate fell 10 points. The hallucination rate fell 10 points. The attempt rates dropped from 82 to 67.

  8. attempt rates dropped from 82 to 67. attempt rates dropped from 82 to 67. This is this bench is measuring how This is this bench is measuring how This is this bench is measuring how likely is the model to lie when it likely is the model to lie when it likely is the model to lie when it doesn't know something and it's much doesn't know something and it's much doesn't know something and it's much less likely now even though for my less likely now even though for my less likely now even though for my experience that is not the case. experience that is not the case. experience that is not the case. Scientific reasoning results remain Scientific reasoning results remain Scientific reasoning results remain largely unchanged. does much better at largely unchanged. does much better at largely unchanged. does much better at GDP val, but I still don't love that GDP val, but I still don't love that GDP val, but I still don't love that bench. Good analysis. We can work with bench. Good analysis. We can work with bench. Good analysis. We can work with that. Muse is trained heavily on data that. Muse is trained heavily on data that. Muse is trained heavily on data from people using models like anthropic from people using models like anthropic from people using models like anthropic models internally, and all of the work models internally, and all of the work models internally, and all of the work they do. I've heard numbers as high as they do. I've heard numbers as high as they do. I've heard numbers as high as 50% of staff at Meta have been moved 50% of staff at Meta have been moved 50% of staff at Meta have been moved over to some type of data labeling or over to some type of data labeling or over to some type of data labeling or like management tasks in order to help like management tasks in order to help like management tasks in order to help them train smarter models. And one of them train smarter models. And one of them train smarter models. And one of the results here is this wonderful tweet the results here is this wonderful tweet the results here is this wonderful tweet from Luke. Muse Spark is trained on the from Luke. Muse Spark is trained on the from Luke. Muse Spark is trained on the screen recordings of Meta employees, screen recordings of Meta employees, screen recordings of Meta employees, which makes it best-in-class at applying which makes it best-in-class at applying which makes it best-in-class at applying for jobs at Anthropic. Absolute banger. for jobs at Anthropic. Absolute banger. for jobs at Anthropic. Absolute banger. I want to go through other numbers that I want to go through other numbers that I want to go through other numbers that matter before showing off the code that matter before showing off the code that matter before showing off the code that I can write with Muse. I already have it I can write with Muse. I already have it I can write with Muse. I already have it working in the background and it's working in the background and it's working in the background and it's surprisingly fast. We'll talk about that surprisingly fast. We'll talk about that surprisingly fast. We'll talk about that in a sec, but first I want to look at in a sec, but first I want to look at in a sec, but first I want to look at the numbers from artificial analysis. As the numbers from artificial analysis. As the numbers from artificial analysis. As you can see here, Muse Spark 1.2 falls you can see here, Muse Spark 1.2 falls you can see here, Muse Spark 1.2 falls right between Kimmy K3 and Gro 4.5. I right between Kimmy K3 and Gro 4.5. I right between Kimmy K3 and Gro 4.5. I will say it's a little bit embarrassing will say it's a little bit embarrassing will say it's a little bit embarrassing to release a closedweight model as a lab to release a closedweight model as a lab to release a closedweight model as a lab that has billions of dollars like Meta that has billions of dollars like Meta that has billions of dollars like Meta does and have it be behind an openweight does and have it be behind an openweight does and have it be behind an openweight model that came out much before it. That model that came out much before it. That model that came out much before it. That is a little embarrassing. But also, if is a little embarrassing. But also, if is a little embarrassing. But also, if it's really fast and cheap and reliable, it's really fast and cheap and reliable, it's really fast and cheap and reliable, these numbers don't mean everything cuz these numbers don't mean everything cuz these numbers don't mean everything cuz I can tell you confidently that Opus 5, I can tell you confidently that Opus 5, I can tell you confidently that Opus 5, despite being at the front of this list despite being at the front of this list despite being at the front of this list and tricking me into thinking it was a and tricking me into thinking it was a and tricking me into thinking it was a great model, ended up all being Copus great model, ended up all being Copus great model, ended up all being Copus because I hate using Opus 5. Now, it

  9. because I hate using Opus 5. Now, it because I hate using Opus 5. Now, it once I started merging the code it once I started merging the code it once I started merging the code it wrote, I realized how bad it was cuz I wrote, I realized how bad it was cuz I wrote, I realized how bad it was cuz I had to have Fable and 56 Soul come in had to have Fable and 56 Soul come in had to have Fable and 56 Soul come in and clean up the mess. Not great. But and clean up the mess. Not great. But and clean up the mess. Not great. But again, Kimmy K3 actually felt quite again, Kimmy K3 actually felt quite again, Kimmy K3 actually felt quite usable, as did Grock 4.5. But Opus at usable, as did Grock 4.5. But Opus at usable, as did Grock 4.5. But Opus at the front doesn't feel usable for real the front doesn't feel usable for real the front doesn't feel usable for real code work at all. But what I really want code work at all. But what I really want code work at all. But what I really want to see is output tokens. How did they do to see is output tokens. How did they do to see is output tokens. How did they do here? Interesting. Okay, so Muse Sparks here? Interesting. Okay, so Muse Sparks here? Interesting. Okay, so Muse Sparks runs did about 30k tokens per task runs did about 30k tokens per task runs did about 30k tokens per task compared to about 36k for fable 5 and compared to about 36k for fable 5 and compared to about 36k for fable 5 and compared to about 17k for 56 soul. So compared to about 17k for 56 soul. So compared to about 17k for 56 soul. So it's not quite as inefficient as it's not quite as inefficient as it's not quite as inefficient as anthropic models, but it is far from the anthropic models, but it is far from the anthropic models, but it is far from the efficiency that we see from models like efficiency that we see from models like efficiency that we see from models like the GPT line or Gro 4.5. Very the GPT line or Gro 4.5. Very the GPT line or Gro 4.5. Very interesting. Definitely puts it in a interesting. Definitely puts it in a interesting. Definitely puts it in a weird different spot. Like this model weird different spot. Like this model weird different spot. Like this model isn't just another openweight model isn't just another openweight model isn't just another openweight model tuned or something like this really is a tuned or something like this really is a tuned or something like this really is a new model and it should feel quite new model and it should feel quite new model and it should feel quite different especially when you look at different especially when you look at different especially when you look at the speed that this model has too. I the speed that this model has too. I the speed that this model has too. I don't think they have recent speed don't think they have recent speed don't think they have recent speed numbers yet on artificial analysis but numbers yet on artificial analysis but numbers yet on artificial analysis but they do have them on open router. Open they do have them on open router. Open they do have them on open router. Open router is an easy way to use pretty much router is an easy way to use pretty much router is an easy way to use pretty much every model and they track the every model and they track the every model and they track the throughput and speeds that users are throughput and speeds that users are throughput and speeds that users are actually seeing when they use the models actually seeing when they use the models actually seeing when they use the models through open router and they're seeing through open router and they're seeing through open router and they're seeing an average of 191 tokens per second an average of 191 tokens per second an average of 191 tokens per second which is absolutely nuts. For reference which is absolutely nuts. For reference which is absolutely nuts. For reference 56 soul is only at about 30 tokens per 56 soul is only at about 30 tokens per 56 soul is only at about 30 tokens per second. So that is a gigantic second. So that is a gigantic second. So that is a gigantic improvement like 5x plus absolute best improvement like 5x plus absolute best improvement like 5x plus absolute best case with soul is 134 TPS but case with soul is 134 TPS but case with soul is 134 TPS but realistically speaking the p50s in the

  10. realistically speaking the p50s in the realistically speaking the p50s in the 40s to 50s on openai's official infra 40s to 50s on openai's official infra 40s to 50s on openai's official infra and the 50s for Azure at best. To and the 50s for Azure at best. To and the 50s for Azure at best. To compare Muse Spark gets 316 TPS at its compare Muse Spark gets 316 TPS at its compare Muse Spark gets 316 TPS at its best and 162 at its P50. That is pretty best and 162 at its P50. That is pretty best and 162 at its P50. That is pretty nuts and I feel the result coding with nuts and I feel the result coding with nuts and I feel the result coding with it already. Man, that cerebrous it already. Man, that cerebrous it already. Man, that cerebrous deployment of 56 soul can't come fast deployment of 56 soul can't come fast deployment of 56 soul can't come fast enough. It was supposed to already be enough. It was supposed to already be enough. It was supposed to already be out, so I'm not sure what happened out, so I'm not sure what happened out, so I'm not sure what happened there. No one's given me any info, there. No one's given me any info, there. No one's given me any info, sadly. Then we get to cost and Muse sadly. Then we get to cost and Muse sadly. Then we get to cost and Muse Spark is performing absurdly well here Spark is performing absurdly well here Spark is performing absurdly well here at 40 cents per task on average through at 40 cents per task on average through at 40 cents per task on average through the artificial analysis intelligence the artificial analysis intelligence the artificial analysis intelligence index. That's still more expensive than index. That's still more expensive than index. That's still more expensive than V4 flash is, as well as 56 Luna on max, V4 flash is, as well as 56 Luna on max, V4 flash is, as well as 56 Luna on max, but it's a lot cheaper than other but it's a lot cheaper than other but it's a lot cheaper than other things. Even Mistro medium is more things. Even Mistro medium is more things. Even Mistro medium is more expensive. And Kimmy K3 is more than expensive. And Kimmy K3 is more than expensive. And Kimmy K3 is more than twice the cost for similar work. And I'm twice the cost for similar work. And I'm twice the cost for similar work. And I'm assuming this is the normal cost too, assuming this is the normal cost too, assuming this is the normal cost too, not that contributor tier, which is not that contributor tier, which is not that contributor tier, which is effectively free. Think it's about time effectively free. Think it's about time effectively free. Think it's about time we fire off some prompts. I just opened we fire off some prompts. I just opened we fire off some prompts. I just opened Muse Code inside of the T3 codebase.

  11. Muse Code inside of the T3 codebase. Muse Code inside of the T3 codebase. It's defaulting to one two contributor It's defaulting to one two contributor It's defaulting to one two contributor on high. And since T3 code is fully open on high. And since T3 code is fully open on high. And since T3 code is fully open source, I don't mind if they get my source, I don't mind if they get my source, I don't mind if they get my data. They probably have it anyways. I data. They probably have it anyways. I data. They probably have it anyways. I asked it to give me an overview of the asked it to give me an overview of the asked it to give me an overview of the architecture of the codebase. And it had architecture of the codebase. And it had architecture of the codebase. And it had this under 30 seconds. Not bad. Let's this under 30 seconds. Not bad. Let's this under 30 seconds. Not bad. Let's see where this goes. If I ask it for see where this goes. If I ask it for see where this goes. If I ask it for something a little harder where it has something a little harder where it has something a little harder where it has to investigate more. I asked it if it to investigate more. I asked it if it to investigate more. I asked it if it can find anything suspicious about the can find anything suspicious about the can find anything suspicious about the event sourcing model. Things cause weird event sourcing model. Things cause weird event sourcing model. Things cause weird behaviors on different platforms, behaviors on different platforms, behaviors on different platforms, unnecessary bandwidth, usage and lag, unnecessary bandwidth, usage and lag, unnecessary bandwidth, usage and lag, etc. Let's see what it does for that. etc. Let's see what it does for that. etc. Let's see what it does for that. Another interesting thing I noticed when Another interesting thing I noticed when Another interesting thing I noticed when I set this up is that it pulled all of I set this up is that it pulled all of I set this up is that it pulled all of my skills and personal rules from claude my skills and personal rules from claude my skills and personal rules from claude code. Normally things don't use the code. Normally things don't use the code. Normally things don't use the cloud directory because it's clouds. cloud directory because it's clouds. cloud directory because it's clouds. They use the shared aagents standard and They use the shared aagents standard and They use the shared aagents standard and agentmd stuff like that. They seem to be agentmd stuff like that. They seem to be agentmd stuff like that. They seem to be really focused on quad. And from really focused on quad. And from really focused on quad. And from everything I know about how meta's been everything I know about how meta's been everything I know about how meta's been working lately, they're definitely an working lately, they're definitely an working lately, they're definitely an enthropic house. They use a lot of opus enthropic house. They use a lot of opus enthropic house. They use a lot of opus and a lot of fable there. Yeah. 174 TPS. and a lot of fable there. Yeah. 174 TPS. and a lot of fable there. Yeah. 174 TPS. Okay. Okay. Okay. Okay. Okay. Okay. Now we're talking. For comparison, Grock Now we're talking. For comparison, Grock Now we're talking. For comparison, Grock 4.5, which is still quite fast, is at 50 4.5, which is still quite fast, is at 50 4.5, which is still quite fast, is at 50 to 52 TPS. And a model like 56 Soul is to 52 TPS. And a model like 56 Soul is to 52 TPS. And a model like 56 Soul is 30 to 50. It's crazy that Azure is 30 to 50. It's crazy that Azure is 30 to 50. It's crazy that Azure is performing that much better than Open performing that much better than Open performing that much better than Open Eyes official servers are right now. But Eyes official servers are right now. But Eyes official servers are right now. But yeah, you're welcome. I had to fight yeah, you're welcome. I had to fight yeah, you're welcome. I had to fight hard to get Azure to move. It had some hard to get Azure to move. It had some hard to get Azure to move. It had some interesting findings here. I'm going to interesting findings here. I'm going to interesting findings here. I'm going to tell it to summarize them. See our tell it to summarize them. See our tell it to summarize them. See our findings in an easy to digest HTML file.

  12. findings in an easy to digest HTML file. findings in an easy to digest HTML file. Let's see if it's able to find my HTML Let's see if it's able to find my HTML Let's see if it's able to find my HTML plan skill and actually take all of this plan skill and actually take all of this plan skill and actually take all of this work and synthesize it. Look at that. work and synthesize it. Look at that. work and synthesize it. Look at that. Already found the skill. Now it is Already found the skill. Now it is Already found the skill. Now it is thinking. Hopefully, it will synthesize thinking. Hopefully, it will synthesize thinking. Hopefully, it will synthesize this properly and then put things up. this properly and then put things up. this properly and then put things up. It's interesting to state how much data It's interesting to state how much data It's interesting to state how much data this skill is. There's definitely some this skill is. There's definitely some this skill is. There's definitely some slop in here. things like the way that slop in here. things like the way that slop in here. things like the way that the full path for files is included in the full path for files is included in the full path for files is included in the UI and stuff like this CLI is a bit the UI and stuff like this CLI is a bit the UI and stuff like this CLI is a bit slopp here is the HTML it made for me to slopp here is the HTML it made for me to slopp here is the HTML it made for me to summarize the findings that it had pure summarize the findings that it had pure summarize the findings that it had pure decider plus single serialized worker is decider plus single serialized worker is decider plus single serialized worker is correct cost surfaces at the wire full correct cost surfaces at the wire full correct cost surfaces at the wire full snapshots on subscribe unbounded snapshots on subscribe unbounded snapshots on subscribe unbounded payloads in a global sequence that payloads in a global sequence that payloads in a global sequence that couples every thread so here's what couples every thread so here's what couples every thread so here's what we'll do for a comparison I'm going to we'll do for a comparison I'm going to we'll do for a comparison I'm going to do a cla run with fable we're going to do a cla run with fable we're going to do a cla run with fable we're going to paste this I have a plan here you know paste this I have a plan here you know paste this I have a plan here you know what I I want to try DeepSeek 4 model. what I I want to try DeepSeek 4 model. what I I want to try DeepSeek 4 model. Deep seek v4 flash. Let's copy over the Deep seek v4 flash. Let's copy over the Deep seek v4 flash. Let's copy over the same prompt then.

  13. God, that is fast as [ __ ] too. For God, that is fast as [ __ ] too. For reference, Muse 1.2 got me this feedback reference, Muse 1.2 got me this feedback reference, Muse 1.2 got me this feedback and even had an HTML page generated all and even had an HTML page generated all and even had an HTML page generated all in well under a minute. I've just been in well under a minute. I've just been in well under a minute. I've just been trying to get similar feedback from trying to get similar feedback from trying to get similar feedback from Fable and we're over four minutes in and Fable and we're over four minutes in and Fable and we're over four minutes in and still don't have any useful anything still don't have any useful anything still don't have any useful anything here. I also have Deepseek Flash running here. I also have Deepseek Flash running here. I also have Deepseek Flash running in the background as well just as a in the background as well just as a in the background as well just as a comparison because the new Flash release comparison because the new Flash release comparison because the new Flash release was really solid too just to get a was really solid too just to get a was really solid too just to get a comparison across things here. RG wasn't comparison across things here. RG wasn't comparison across things here. RG wasn't finding them because I used - G quote finding them because I used - G quote finding them because I used - G quote star.ts which doesn't work on Mac OS star.ts which doesn't work on Mac OS star.ts which doesn't work on Mac OS with rip grip. This is the Deepseek one with rip grip. This is the Deepseek one with rip grip. This is the Deepseek one running by the way. God, these are both running by the way. God, these are both running by the way. God, these are both so slow and not getting me useful so slow and not getting me useful so slow and not getting me useful feedback so far in comparison. It's been feedback so far in comparison. It's been feedback so far in comparison. It's been very interesting. My plan is I'm going very interesting. My plan is I'm going very interesting. My plan is I'm going to take the findings that Muse had and to take the findings that Muse had and to take the findings that Muse had and throw them at these other two models throw them at these other two models throw them at these other two models after and say, "How do these compare?" after and say, "How do these compare?" after and say, "How do these compare?" So, my favorite ways to get a feel for So, my favorite ways to get a feel for So, my favorite ways to get a feel for what model strengths and weaknesses are what model strengths and weaknesses are what model strengths and weaknesses are is to have them all do the same task, is to have them all do the same task, is to have them all do the same task, have them write up the results, and then have them write up the results, and then have them write up the results, and then have them compare each other's results. have them compare each other's results. have them compare each other's results. Deepseek Flash just finished its Deepseek Flash just finished its Deepseek Flash just finished its analysis. I'll ask it write up your analysis. I'll ask it write up your analysis. I'll ask it write up your findings in HTML. See if it can figure findings in HTML. See if it can figure findings in HTML. See if it can figure out my HTML skill. Oh, um, use my HTML out my HTML skill. Oh, um, use my HTML out my HTML skill. Oh, um, use my HTML skill. You can read it from my agents skill. You can read it from my agents skill. You can read it from my agents orclaude.

  14. orclaude. orclaude. It doesn't have skills built in by It doesn't have skills built in by It doesn't have skills built in by default if I recall. Oh, it did. It default if I recall. Oh, it did. It default if I recall. Oh, it did. It found it. Nice. After waiting way too found it. Nice. After waiting way too found it. Nice. After waiting way too long for both Fable and Deepseek to do a long for both Fable and Deepseek to do a long for both Fable and Deepseek to do a similar investigation, I asked them both similar investigation, I asked them both similar investigation, I asked them both to review their findings and compare it to review their findings and compare it to review their findings and compare it to what was found with Muse 1.2. And to what was found with Muse 1.2. And to what was found with Muse 1.2. And here we see the way that Fable here we see the way that Fable here we see the way that Fable categorizes the differences. And categorizes the differences. And categorizes the differences. And remember, Fable, when doing a similar remember, Fable, when doing a similar remember, Fable, when doing a similar task with Opus, thought that Opus had a task with Opus, thought that Opus had a task with Opus, thought that Opus had a much better investigation overall. It much better investigation overall. It much better investigation overall. It seems like that is not the case with the seems like that is not the case with the seems like that is not the case with the Muse investigations. There are a few Muse investigations. There are a few Muse investigations. There are a few places where it said that Muse did places where it said that Muse did places where it said that Muse did better. For example, the coverage breath better. For example, the coverage breath better. For example, the coverage breath it thinks Muse went a little further on, it thinks Muse went a little further on, it thinks Muse went a little further on, but for the most part, it thinks it did but for the most part, it thinks it did but for the most part, it thinks it did worse and that Fable's plan was much worse and that Fable's plan was much worse and that Fable's plan was much more accurate and overall properly found more accurate and overall properly found more accurate and overall properly found root causes. Meanwhile, Deepseek seems root causes. Meanwhile, Deepseek seems root causes. Meanwhile, Deepseek seems to think that the Muse plan was to think that the Muse plan was to think that the Muse plan was meaningfully better. Very interesting. meaningfully better. Very interesting. meaningfully better. Very interesting. Yeah, it's a solid model for the price, Yeah, it's a solid model for the price, Yeah, it's a solid model for the price, but I don't know if I would trust it for but I don't know if I would trust it for but I don't know if I would trust it for really heavy endto-end work yet. But really heavy endto-end work yet. But really heavy endto-end work yet. But we'll test it a little more. I had to we'll test it a little more. I had to we'll test it a little more. I had to try one of my favorite tasks where I try one of my favorite tasks where I try one of my favorite tasks where I take my crappy fish game and have it take my crappy fish game and have it take my crappy fish game and have it remake it both in 2D and in 3D. And the remake it both in 2D and in 3D. And the remake it both in 2D and in 3D. And the most impressive thing is how quickly it most impressive thing is how quickly it most impressive thing is how quickly it did it. It made the 2D version in about did it. It made the 2D version in about did it. It made the 2D version in about 2 and 1/2 minutes and the 3D version in 2 and 1/2 minutes and the 3D version in 2 and 1/2 minutes and the 3D version in under five. For reference, Opus 5 took under five. For reference, Opus 5 took under five. For reference, Opus 5 took over an hour for both of these. It over an hour for both of these. It over an hour for both of these. It doesn't move properly. Like I can't turn doesn't move properly. Like I can't turn doesn't move properly. Like I can't turn Kind of nuts that it made this whole thing of nuts that it made this whole thing from scratch in four minutes, though.

  15. from scratch in four minutes, though. from scratch in four minutes, though. Like, that part is impressive. Like, that part is impressive. Like, that part is impressive. It's just that the rest is broken. Mouse It's just that the rest is broken. Mouse It's just that the rest is broken. Mouse look does not work. And it made a 2D look does not work. And it made a 2D look does not work. And it made a 2D version of Fishlop as well in 2 minutes version of Fishlop as well in 2 minutes version of Fishlop as well in 2 minutes and 39 seconds. For reference, when I and 39 seconds. For reference, when I and 39 seconds. For reference, when I put these same prompts to remake this put these same prompts to remake this put these same prompts to remake this game into other models like Opus, it game into other models like Opus, it game into other models like Opus, it took multiple hours to do. And this made took multiple hours to do. And this made took multiple hours to do. And this made it in under five minutes for both the 2D it in under five minutes for both the 2D it in under five minutes for both the 2D and 3D versions. and 3D versions. and 3D versions. Is it perfect? No, far from it. But it Is it perfect? No, far from it. But it Is it perfect? No, far from it. But it does have a couple nicities. Like it it does have a couple nicities. Like it it does have a couple nicities. Like it it made solid animations using the sprites made solid animations using the sprites made solid animations using the sprites that I had access to from the other that I had access to from the other that I had access to from the other project. It did shadows and like project. It did shadows and like project. It did shadows and like contrast well. Like it it has a weird contrast well. Like it it has a weird contrast well. Like it it has a weird type of taste that I haven't seen models type of taste that I haven't seen models type of taste that I haven't seen models have. It's obviously still jank as [ __ ] have. It's obviously still jank as [ __ ] have. It's obviously still jank as [ __ ] in a lot of ways, but like the little in a lot of ways, but like the little in a lot of ways, but like the little shadows are nice. The way things are shadows are nice. The way things are shadows are nice. The way things are moving is nice. It feels fluid. It's moving is nice. It feels fluid. It's moving is nice. It feels fluid. It's this model feels different. It doesn't this model feels different. It doesn't this model feels different. It doesn't feel like it's just distilled on other feel like it's just distilled on other feel like it's just distilled on other popular things. It does have a vibe to popular things. It does have a vibe to popular things. It does have a vibe to the results. And I am definitely curious the results. And I am definitely curious the results. And I am definitely curious what's going to happen when they make what's going to happen when they make what's going to happen when they make bigger and smarter versions. It bigger and smarter versions. It bigger and smarter versions. It allegedly just fixed the 3D mouse broken allegedly just fixed the 3D mouse broken allegedly just fixed the 3D mouse broken stuff in a few seconds.

  16. And it does appear that it has. I can And it does appear that it has. I can now look around as expected. The fish are all swimming backwards The fish are all swimming backwards though. That's hilarious. Fish are though. That's hilarious. Fish are though. That's hilarious. Fish are swimming backwards. I can't look up or swimming backwards. I can't look up or swimming backwards. I can't look up or down. Only left and right. Can you use down. Only left and right. Can you use down. Only left and right. Can you use computer use to test it yourself? Oh, if computer use to test it yourself? Oh, if computer use to test it yourself? Oh, if I might codeex computer use skill. Oh I might codeex computer use skill. Oh I might codeex computer use skill. Oh no, that is not the solution. As if it no, that is not the solution. As if it no, that is not the solution. As if it can do it itself without codecs. Does it can do it itself without codecs. Does it can do it itself without codecs. Does it have vision? Oh [ __ ] it does. That's have vision? Oh [ __ ] it does. That's have vision? Oh [ __ ] it does. That's huge. Okay, now the up and down movement huge. Okay, now the up and down movement huge. Okay, now the up and down movement of the mouse tilts this. That is broken of the mouse tilts this. That is broken of the mouse tilts this. That is broken as [ __ ] And the fish are still moving as [ __ ] And the fish are still moving as [ __ ] And the fish are still moving backwards. It promised me it fixed that. backwards. It promised me it fixed that. backwards. It promised me it fixed that. Oh, wait. No, it didn't. Yes, please fix Oh, wait. No, it didn't. Yes, please fix Oh, wait. No, it didn't. Yes, please fix and rebuild. Okay. Nope. Fish moved the and rebuild. Okay. Nope. Fish moved the and rebuild. Okay. Nope. Fish moved the right way now, but I still can't look up right way now, but I still can't look up right way now, but I still can't look up and down.

  17. spaces both feed and go up. Yeah, spaces both feed and go up. Yeah, there's a lot of little things it got there's a lot of little things it got there's a lot of little things it got wrong here. Like it wrong here. Like it wrong here. Like it it doesn't understand it doesn't understand it doesn't understand that things can collide with each other, that things can collide with each other, that things can collide with each other, which is interesting. So, one of the which is interesting. So, one of the which is interesting. So, one of the things they said the model was good at things they said the model was good at things they said the model was good at was like breaking things up and letting was like breaking things up and letting was like breaking things up and letting lots of stuff work at the same time. lots of stuff work at the same time. lots of stuff work at the same time. But, it's it's getting work done and But, it's it's getting work done and But, it's it's getting work done and it's doing it alarmingly fast. Like, it's doing it alarmingly fast. Like, it's doing it alarmingly fast. Like, that's one of the nicest things here is that's one of the nicest things here is that's one of the nicest things here is using it to like touch something up. using it to like touch something up. using it to like touch something up. seems to be a useful method. Dra just seems to be a useful method. Dra just seems to be a useful method. Dra just got me one of my favorite things to look got me one of my favorite things to look got me one of my favorite things to look at when new models drop. The comparison at when new models drop. The comparison at when new models drop. The comparison of how it handles design on his of how it handles design on his of how it handles design on his witchai.dev site. Very nice for like witchai.dev site. Very nice for like witchai.dev site. Very nice for like looking at the designs that different looking at the designs that different looking at the designs that different models do and comparing them. And I'm models do and comparing them. And I'm models do and comparing them. And I'm already seeing some interesting things already seeing some interesting things already seeing some interesting things here. Like I love how it uses the scroll here. Like I love how it uses the scroll here. Like I love how it uses the scroll areas where this like graphic on the areas where this like graphic on the areas where this like graphic on the side scrolls with you until you hit a side scrolls with you until you hit a side scrolls with you until you hit a certain point. certain point. certain point. The sections are tasteful, too. I I The sections are tasteful, too. I I The sections are tasteful, too. I I don't hate this. I'm tired of the pills don't hate this. I'm tired of the pills don't hate this. I'm tired of the pills on everything. It feels a little on everything. It feels a little on everything. It feels a little templaty, but not bad at all. And this templaty, but not bad at all. And this templaty, but not bad at all. And this little tilted reminder recall thing here little tilted reminder recall thing here little tilted reminder recall thing here is nice, too. This is all without the is nice, too. This is all without the is nice, too. This is all without the design skill as well. If we switch to design skill as well. If we switch to design skill as well. If we switch to the second one, we end up with one of the second one, we end up with one of the second one, we end up with one of the cringy code style ones. It's laid the cringy code style ones. It's laid the cringy code style ones. It's laid out well. Like, its page layouts are out well. Like, its page layouts are out well. Like, its page layouts are solid so far, but I don't love the solid so far, but I don't love the solid so far, but I don't love the design and the graph it made there. It design and the graph it made there. It design and the graph it made there. It really likes doing these tilted things.

  18. really likes doing these tilted things. really likes doing these tilted things. It has one here as well. I actually I It has one here as well. I actually I It has one here as well. I actually I really like this one of them so far. really like this one of them so far. really like this one of them so far. This design looks decent. And I like the This design looks decent. And I like the This design looks decent. And I like the way it's using the like squirle here. I way it's using the like squirle here. I way it's using the like squirle here. I don't like that it switches from squirle don't like that it switches from squirle don't like that it switches from squirle to rounded there. This isn't bad though. to rounded there. This isn't bad though. to rounded there. This isn't bad though. This again is reminiscent of Gemini to This again is reminiscent of Gemini to This again is reminiscent of Gemini to an extent. Way too much text. Way too an extent. Way too much text. Way too an extent. Way too much text. Way too much text, but otherwise not bad. These much text, but otherwise not bad. These much text, but otherwise not bad. These brutalist ones are getting tiring. I say brutalist ones are getting tiring. I say brutalist ones are getting tiring. I say after making lawn video, which is after making lawn video, which is after making lawn video, which is absolutely of this style. But yeah, it's absolutely of this style. But yeah, it's absolutely of this style. But yeah, it's not bad. Like nothing jumps out as like not bad. Like nothing jumps out as like not bad. Like nothing jumps out as like that is horrible. I don't love having that is horrible. I don't love having that is horrible. I don't love having the bar here when it's already a very the bar here when it's already a very the bar here when it's already a very brutalist like lineheavy style and now brutalist like lineheavy style and now brutalist like lineheavy style and now the classic like Tailwind template the classic like Tailwind template the classic like Tailwind template homepage. homepage. homepage. Oh, it does these little floaties. Well, Oh, it does these little floaties. Well, Oh, it does these little floaties. Well, I I like how it uses those other models I I like how it uses those other models I I like how it uses those other models get that stuff really wrong and I fought get that stuff really wrong and I fought get that stuff really wrong and I fought it a ton on things like the T3 code it a ton on things like the T3 code it a ton on things like the T3 code marketing site. That's like what's marketing site. That's like what's marketing site. That's like what's interesting about this model just has a interesting about this model just has a interesting about this model just has a bit of a different flavor. You know bit of a different flavor. You know bit of a different flavor. You know what? Let's give it something harder. what? Let's give it something harder. what? Let's give it something harder. Does it have work tree support first Does it have work tree support first Does it have work tree support first off? It does. Huge. That means I could off? It does. Huge. That means I could off? It does. Huge. That means I could play around with less risk a bit here.

  19. play around with less risk a bit here. play around with less risk a bit here. Can I yolo mode from here? Nice. Oh, Can I yolo mode from here? Nice. Oh, Can I yolo mode from here? Nice. Oh, that's a bit annoying cuz I had that in that's a bit annoying cuz I had that in that's a bit annoying cuz I had that in a work tree. D-workree- yolo. Now we're a work tree. D-workree- yolo. Now we're a work tree. D-workree- yolo. Now we're in a safe work tree where I can ask it in a safe work tree where I can ask it in a safe work tree where I can ask it to do stupid things like let's try their to do stupid things like let's try their to do stupid things like let's try their voice. Actually, I would like to voice. Actually, I would like to voice. Actually, I would like to implement muse as a provider inside of implement muse as a provider inside of implement muse as a provider inside of T3 code. I'm not sure what Muse code T3 code. I'm not sure what Muse code T3 code. I'm not sure what Muse code supports in terms of integration supports in terms of integration supports in terms of integration methods. I haven't played with it enough methods. I haven't played with it enough methods. I haven't played with it enough or looked into the SDK. I don't even or looked into the SDK. I don't even or looked into the SDK. I don't even know if it's open source. What I really know if it's open source. What I really know if it's open source. What I really want is to integrate it through a layer want is to integrate it through a layer want is to integrate it through a layer like ACP similar to how we have like ACP similar to how we have like ACP similar to how we have integrated other providers in the past. integrated other providers in the past. integrated other providers in the past. But if we need something more custom, I But if we need something more custom, I But if we need something more custom, I am down. I would like you to start by am down. I would like you to start by am down. I would like you to start by investigating everything you need here. investigating everything you need here. investigating everything you need here. Both how we can implement within T3 Both how we can implement within T3 Both how we can implement within T3 code, but also what is offered by Muse code, but also what is offered by Muse code, but also what is offered by Muse itself. You'll have to do some digging itself. You'll have to do some digging itself. You'll have to do some digging to find source code where available, to find source code where available, to find source code where available, docs, SDKs, and whatnot. Might be a bit docs, SDKs, and whatnot. Might be a bit docs, SDKs, and whatnot. Might be a bit hard to find because this is all still hard to find because this is all still hard to find because this is all still so new. Use a lot of sub aents to break so new. Use a lot of sub aents to break so new. Use a lot of sub aents to break up the work as you go. I actually liked up the work as you go. I actually liked up the work as you go. I actually liked that voice to text. It showed things as that voice to text. It showed things as that voice to text. It showed things as I spoke, which a lot of other solutions I spoke, which a lot of other solutions I spoke, which a lot of other solutions don't. They just show it when you're don't. They just show it when you're don't. They just show it when you're done. So, let's send that over and see done. So, let's send that over and see done. So, let's send that over and see how it goes. That's the first good how it goes. That's the first good how it goes. That's the first good terminal voice to text I've seen in any terminal voice to text I've seen in any terminal voice to text I've seen in any of these CLIs for being real. Look at of these CLIs for being real. Look at of these CLIs for being real. Look at that. is already spawning sub aents and that. is already spawning sub aents and that. is already spawning sub aents and we can look at them almost the exact we can look at them almost the exact we can look at them almost the exact same way we can in claude. I I know I same way we can in claude. I I know I same way we can in claude. I I know I click baited a bit saying this is a cla click baited a bit saying this is a cla click baited a bit saying this is a cla code clone, but like this is such a code clone, but like this is such a code clone, but like this is such a cloud code clone. It's cool that it spun cloud code clone. It's cool that it spun cloud code clone. It's cool that it spun up all of these intelligently. Like it's up all of these intelligently. Like it's up all of these intelligently. Like it's it's clear they are trying to make this it's clear they are trying to make this it's clear they are trying to make this model good at this type of like broken

  20. model good at this type of like broken model good at this type of like broken up heavy sub aent breakup work. I just up heavy sub aent breakup work. I just up heavy sub aent breakup work. I just got hit with a rate limit. got hit with a rate limit. got hit with a rate limit. Are you kidding? I put in a credit card Are you kidding? I put in a credit card Are you kidding? I put in a credit card and everything. Are these sub agents and everything. Are these sub agents and everything. Are these sub agents getting hit with them, too? They are. getting hit with them, too? They are. getting hit with them, too? They are. Great. Why why would you make a sub aent Great. Why why would you make a sub aent Great. Why why would you make a sub aent flow like this if your APIs can't even flow like this if your APIs can't even flow like this if your APIs can't even handle it meta? Uh how much does it cost handle it meta? Uh how much does it cost handle it meta? Uh how much does it cost so far is a good question. Interesting. so far is a good question. Interesting. so far is a good question. Interesting. I actually checked out the dashboard I actually checked out the dashboard I actually checked out the dashboard before. They have instructions on how to before. They have instructions on how to before. They have instructions on how to set it up with other providers too like set it up with other providers too like set it up with other providers too like open code, cloud code, codeex as well as open code, cloud code, codeex as well as open code, cloud code, codeex as well as curl and python directly. Assuming the curl and python directly. Assuming the curl and python directly. Assuming the Python is showing you how to like set it Python is showing you how to like set it Python is showing you how to like set it up with code. I'm assuming this is like up with code. I'm assuming this is like up with code. I'm assuming this is like an OpenAI compatible API. I want to see an OpenAI compatible API. I want to see an OpenAI compatible API. I want to see how much I have spent. I don't know how how much I have spent. I don't know how how much I have spent. I don't know how accurate or up to date this is, but accurate or up to date this is, but accurate or up to date this is, but everything I've done so far apparently everything I've done so far apparently everything I've done so far apparently is only about 13 cents, including both is only about 13 cents, including both is only about 13 cents, including both of those game rewrites as well as the of those game rewrites as well as the of those game rewrites as well as the codebase audit. So, yeah, absurdly codebase audit. So, yeah, absurdly codebase audit. So, yeah, absurdly cheap. It's making progress here. Why is cheap. It's making progress here. Why is cheap. It's making progress here. Why is it talking about anti-gravity? it talking about anti-gravity? it talking about anti-gravity? I just asked why is so much of this I just asked why is so much of this I just asked why is so much of this about anti-gravity? I didn't mention about anti-gravity? I didn't mention about anti-gravity? I didn't mention anti-gravity at all. It did a deep dive anti-gravity at all. It did a deep dive anti-gravity at all. It did a deep dive and it took under three minutes despite and it took under three minutes despite and it took under three minutes despite the fact that it was hitting rate the fact that it was hitting rate the fact that it was hitting rate limits, but like the results are limits, but like the results are limits, but like the results are useless. Contributor has very low rate useless. Contributor has very low rate useless. Contributor has very low rate limits. Okay, that's the problem. It's limits. Okay, that's the problem. It's limits. Okay, that's the problem. It's cuz I'm on that contributor tier. Okay, cuz I'm on that contributor tier. Okay, cuz I'm on that contributor tier. Okay, I shouldn't be so amused by this, but I shouldn't be so amused by this, but I shouldn't be so amused by this, but this is actually objectively hilarious.

  21. this is actually objectively hilarious. this is actually objectively hilarious. I I I resplap heavily in recent search results and I heavily in recent search results and I ran the whole investigation on that ran the whole investigation on that ran the whole investigation on that assumption without confirming. I don't assumption without confirming. I don't assumption without confirming. I don't have a clear hit for Muse as a distinct have a clear hit for Muse as a distinct have a clear hit for Muse as a distinct product. Web searches mostly return Muse product. Web searches mostly return Muse product. Web searches mostly return Muse Spark, which is Meta's new coding model. Spark, which is Meta's new coding model. Spark, which is Meta's new coding model. July 2026. An unreleased Muse Markdown July 2026. An unreleased Muse Markdown July 2026. An unreleased Muse Markdown OS project. Something definitive for a OS project. Something definitive for a OS project. Something definitive for a Muse CLI or SDK. Can you point to me to Muse CLI or SDK. Can you point to me to Muse CLI or SDK. Can you point to me to what you mean by Muse? A URL, the what you mean by Muse? A URL, the what you mean by Muse? A URL, the company behind it, or what the binary or company behind it, or what the binary or company behind it, or what the binary or SDK is called would be enough. I'll give SDK is called would be enough. I'll give SDK is called would be enough. I'll give you a hint. I'm in it right now. Let's you a hint. I'm in it right now. Let's you a hint. I'm in it right now. Let's see if this can figure out what it is. see if this can figure out what it is. see if this can figure out what it is. I'm way too amused. Yes, Muse is Muse I'm way too amused. Yes, Muse is Muse I'm way too amused. Yes, Muse is Muse Code itself. The CLI you're in right Code itself. The CLI you're in right Code itself. The CLI you're in right now. Powered by Muse Code 0.1.0. Powered now. Powered by Muse Code 0.1.0. Powered now. Powered by Muse Code 0.1.0. Powered by Muse Spark. My first investigation by Muse Spark. My first investigation by Muse Spark. My first investigation mapped it to anti-gravity by mistake. mapped it to anti-gravity by mistake. mapped it to anti-gravity by mistake. Completely wrong target. So of course Completely wrong target. So of course Completely wrong target. So of course the Muse code layer here does not have a the Muse code layer here does not have a the Muse code layer here does not have a lot of the functionality that I would lot of the functionality that I would lot of the functionality that I would want to integrate with it. Things like want to integrate with it. Things like want to integrate with it. Things like it being open source. Things like it it being open source. Things like it it being open source. Things like it having an SDK I can plug into ACP as a having an SDK I can plug into ACP as a having an SDK I can plug into ACP as a way to communicate with it. It's missing way to communicate with it. It's missing way to communicate with it. It's missing a lot of the stuff I need. I do have to a lot of the stuff I need. I do have to a lot of the stuff I need. I do have to write up a plan on how we would write up a plan on how we would write up a plan on how we would integrate Muse code. You know what? I'm integrate Muse code. You know what? I'm integrate Muse code. You know what? I'm going to send the smarter model through going to send the smarter model through going to send the smarter model through this with a smarter tool code. code.

  22. this with a smarter tool code. code. this with a smarter tool code. code. I'll do it on this machine so that I I'll do it on this machine so that I I'll do it on this machine so that I have the results here. I want to have the results here. I want to have the results here. I want to integrate the new Muse code integrate the new Muse code integrate the new Muse code from meta into T3 code as a new from meta into T3 code as a new from meta into T3 code as a new provider. Not much documentation exists, provider. Not much documentation exists, provider. Not much documentation exists, but I do have it installed right now. but I do have it installed right now. but I do have it installed right now. Investigate and help me figure out how Investigate and help me figure out how Investigate and help me figure out how we can integrate. So I have Fable both we can integrate. So I have Fable both we can integrate. So I have Fable both investigating how it would do it and investigating how it would do it and investigating how it would do it and separately in parallel I have it separately in parallel I have it separately in parallel I have it reviewing the plan that was written by reviewing the plan that was written by reviewing the plan that was written by Muse. O the page layout's super broken. Muse. O the page layout's super broken. Muse. O the page layout's super broken. Bit annoying to make those types of Bit annoying to make those types of Bit annoying to make those types of mistakes still. There's a lot of things mistakes still. There's a lot of things mistakes still. There's a lot of things here that feel really last generation, here that feel really last generation, here that feel really last generation, but a lot of things that feel next but a lot of things that feel next but a lot of things that feel next generation, too, where it's like it generation, too, where it's like it generation, too, where it's like it seems to know how to break things up seems to know how to break things up seems to know how to break things up with sub agents really well, but it also with sub agents really well, but it also with sub agents really well, but it also seems to struggle meaningfully with like seems to struggle meaningfully with like seems to struggle meaningfully with like basic page layout stuff or staying on basic page layout stuff or staying on basic page layout stuff or staying on task and not hallucinating its way down task and not hallucinating its way down task and not hallucinating its way down an entirely incorrect path. Do we get uh an entirely incorrect path. Do we get uh an entirely incorrect path. Do we get uh omniscience scores here yet? Um, oh, omniscience scores here yet? Um, oh, omniscience scores here yet? Um, oh, apparently it does well on a apparently it does well on a apparently it does well on a omniscience. That is weird to me cuz we omniscience. That is weird to me cuz we omniscience. That is weird to me cuz we just watched it hallucinate just watched it hallucinate just watched it hallucinate aggressively. Allegedly, it is about as aggressively. Allegedly, it is about as aggressively. Allegedly, it is about as bad of hallucinations as 56 Soul is and bad of hallucinations as 56 Soul is and bad of hallucinations as 56 Soul is and slightly better than Kimmy K3, but slightly better than Kimmy K3, but slightly better than Kimmy K3, but considering how aggressively we just considering how aggressively we just considering how aggressively we just watched it do that, that was rough. I watched it do that, that was rough. I watched it do that, that was rough. I had Fable give feedback on Muse's plan had Fable give feedback on Muse's plan had Fable give feedback on Muse's plan to integrate Muse and it said it's a to integrate Muse and it said it's a to integrate Muse and it said it's a solid plan and gave a little bit of solid plan and gave a little bit of solid plan and gave a little bit of feedback. So, I'm going to do what I

  23. feedback. So, I'm going to do what I feedback. So, I'm going to do what I usually do in my real world day-to-day usually do in my real world day-to-day usually do in my real world day-to-day work when doing things like this and I'm work when doing things like this and I'm work when doing things like this and I'm going to take that response. I'm going going to take that response. I'm going going to take that response. I'm going to paste it into the original chat with to paste it into the original chat with to paste it into the original chat with the first model that made the plan and the first model that made the plan and the first model that made the plan and see what it does. Oh, it updated the see what it does. Oh, it updated the see what it does. Oh, it updated the plan alarmingly quickly as well. It's plan alarmingly quickly as well. It's plan alarmingly quickly as well. It's It's very nice how fast this model is. It's very nice how fast this model is. It's very nice how fast this model is. And since it's fast and token efficient, And since it's fast and token efficient, And since it's fast and token efficient, it just feels great. Plan updated. Take it just feels great. Plan updated. Take it just feels great. Plan updated. Take another look. It is nice having all of another look. It is nice having all of another look. It is nice having all of these fast, cheap models getting smart these fast, cheap models getting smart these fast, cheap models getting smart again. It's been a while cuz everybody's again. It's been a while cuz everybody's again. It's been a while cuz everybody's been fighting so hard on the frontier been fighting so hard on the frontier been fighting so hard on the frontier considering that OpenAI just lowered the considering that OpenAI just lowered the considering that OpenAI just lowered the price of Luna by 80%. Which is a massive price of Luna by 80%. Which is a massive price of Luna by 80%. Which is a massive drop on a model that was already really drop on a model that was already really drop on a model that was already really cheap. $120 per mill out and 20 cents cheap. $120 per mill out and 20 cents cheap. $120 per mill out and 20 cents per mill in is just unbelievable. It's per mill in is just unbelievable. It's per mill in is just unbelievable. It's crazy that you can get a model with a crazy that you can get a model with a crazy that you can get a model with a million context window with the million context window with the million context window with the capabilities that Luna has for this capabilities that Luna has for this capabilities that Luna has for this price. And I've been using it a lot more price. And I've been using it a lot more price. And I've been using it a lot more for like random background tasks, for like random background tasks, for like random background tasks, categorization, titles, and stuff like categorization, titles, and stuff like categorization, titles, and stuff like that. We went from having almost no good that. We went from having almost no good that. We went from having almost no good small models drop for months if not like small models drop for months if not like small models drop for months if not like a year to having Luna come out and then a year to having Luna come out and then a year to having Luna come out and then price drop massively by like 5x.

  24. price drop massively by like 5x. price drop massively by like 5x. Deepseek V4 Flash just got a new Deepseek V4 Flash just got a new Deepseek V4 Flash just got a new snapshot that I was testing earlier that snapshot that I was testing earlier that snapshot that I was testing earlier that is really good, really fast, and really is really good, really fast, and really is really good, really fast, and really cheap. And now we have this new Spark cheap. And now we have this new Spark cheap. And now we have this new Spark 1.2 model that's also seemingly really 1.2 model that's also seemingly really 1.2 model that's also seemingly really good for the money, especially if you're good for the money, especially if you're good for the money, especially if you're on that contributor tier where they are on that contributor tier where they are on that contributor tier where they are reading all of the things that get sent. reading all of the things that get sent. reading all of the things that get sent. After the one passive feedback from After the one passive feedback from After the one passive feedback from Fable, apparently it thinks that the Fable, apparently it thinks that the Fable, apparently it thinks that the plan Muse wrote is ready to go. That's plan Muse wrote is ready to go. That's plan Muse wrote is ready to go. That's kind of nuts, man. I [ __ ] love T3 kind of nuts, man. I [ __ ] love T3 kind of nuts, man. I [ __ ] love T3 code. It's so nice being able to just code. It's so nice being able to just code. It's so nice being able to just like hop between models, harnesses, and like hop between models, harnesses, and like hop between models, harnesses, and computers. I'm now spinning up soul on computers. I'm now spinning up soul on computers. I'm now spinning up soul on one of my Linux boxes to compare the one of my Linux boxes to compare the one of my Linux boxes to compare the plans that Fable and Muse wrote plans that Fable and Muse wrote plans that Fable and Muse wrote separately. See, we're up to 36 cents of separately. See, we're up to 36 cents of separately. See, we're up to 36 cents of spend now that we've done all of the spend now that we've done all of the spend now that we've done all of the planning. Oh, nope. 40 cents. Wow, I'm planning. Oh, nope. 40 cents. Wow, I'm planning. Oh, nope. 40 cents. Wow, I'm going to go broke at this rate. For going to go broke at this rate. For going to go broke at this rate. For reference, just my little tests with reference, just my little tests with reference, just my little tests with Fable and Opus on this computer for Fable and Opus on this computer for Fable and Opus on this computer for reading those plans and investigating reading those plans and investigating reading those plans and investigating them is already $32 almost of spend in them is already $32 almost of spend in them is already $32 almost of spend in Claude. So that is a 100x gap. And I've Claude. So that is a 100x gap. And I've Claude. So that is a 100x gap. And I've done way less work with these models done way less work with these models done way less work with these models than I have with the Muse one. That is than I have with the Muse one. That is than I have with the Muse one. That is on the they have access to all my data on the they have access to all my data on the they have access to all my data tier. But to be fair, Fable 5 also does tier. But to be fair, Fable 5 also does tier. But to be fair, Fable 5 also does have them storing the data. They're not have them storing the data. They're not have them storing the data. They're not training on it allegedly, but they are training on it allegedly, but they are training on it allegedly, but they are storing it for safety reasons. Opus, storing it for safety reasons. Opus, storing it for safety reasons. Opus, they don't do that with. And if you they don't do that with. And if you they don't do that with. And if you spend the much higher 10 to 20 times spend the much higher 10 to 20 times spend the much higher 10 to 20 times more on the standard tier for Spark 1.2, more on the standard tier for Spark 1.2, more on the standard tier for Spark 1.2, then you will end up paying still not then you will end up paying still not then you will end up paying still not anywhere near this much money. But if anywhere near this much money. But if anywhere near this much money. But if it's 10 to 20 times more, which is it's 10 to 20 times more, which is it's 10 to 20 times more, which is roughly what it is, that would be 4 to 8 roughly what it is, that would be 4 to 8 roughly what it is, that would be 4 to 8 at most for a bunch of work, that is

  25. at most for a bunch of work, that is at most for a bunch of work, that is decent. Oh boy, it's spinning up those decent. Oh boy, it's spinning up those decent. Oh boy, it's spinning up those sub agents now. Probably going to hit sub agents now. Probably going to hit sub agents now. Probably going to hit rate limits again. Yeah, I'm hitting rate limits again. Yeah, I'm hitting rate limits again. Yeah, I'm hitting rate limits again. I might have to rate limits again. I might have to rate limits again. I might have to switch. I'll stop all of these. switch. I'll stop all of these. switch. I'll stop all of these. We'll move off the contributor tier. We'll move off the contributor tier. We'll move off the contributor tier. Let's see how it handles. Uh, continue. Let's see how it handles. Uh, continue. Let's see how it handles. Uh, continue. It is nice only having one model and It is nice only having one model and It is nice only having one model and just picking between the two versions just picking between the two versions just picking between the two versions that have different pricing for the that have different pricing for the that have different pricing for the exact same thing. But when one of them exact same thing. But when one of them exact same thing. But when one of them hits rate limited that hard, it's a hits rate limited that hard, it's a hits rate limited that hard, it's a little annoying. I just spend way too little annoying. I just spend way too little annoying. I just spend way too much money in comparison. Even though much money in comparison. Even though much money in comparison. Even though all the data I'm about to get here is all the data I'm about to get here is all the data I'm about to get here is technically available for anyone to technically available for anyone to technically available for anyone to train on because it's in my videos. The train on because it's in my videos. The train on because it's in my videos. The review sub agents are finishing up their review sub agents are finishing up their review sub agents are finishing up their work. I I switched to the higher price work. I I switched to the higher price work. I I switched to the higher price tier and I'm still getting rate limited. tier and I'm still getting rate limited. tier and I'm still getting rate limited. Are you kidding? Like I'm just doing the Are you kidding? Like I'm just doing the Are you kidding? Like I'm just doing the sub agent stuff. They said that they sub agent stuff. They said that they sub agent stuff. They said that they support well and I'm hitting rate limits support well and I'm hitting rate limits support well and I'm hitting rate limits constantly. This is obnoxious. And it constantly. This is obnoxious. And it constantly. This is obnoxious. And it sucks. It's like the harness actually sucks. It's like the harness actually sucks. It's like the harness actually feels pretty good. It's like a minimal feels pretty good. It's like a minimal feels pretty good. It's like a minimal polished up quad code. It is making my polished up quad code. It is making my polished up quad code. It is making my laptop a little warmer than I would laptop a little warmer than I would laptop a little warmer than I would like. But uh yeah, this is like. But uh yeah, this is like. But uh yeah, this is I am annoyed by the rate limit thing I am annoyed by the rate limit thing I am annoyed by the rate limit thing more than anything else here. The rest more than anything else here. The rest more than anything else here. The rest is not the worst. Like rate limits on is not the worst. Like rate limits on is not the worst. Like rate limits on subscription plans make some sense. rate subscription plans make some sense. rate subscription plans make some sense. rate limits on paid per token tiers where I limits on paid per token tiers where I limits on paid per token tiers where I am using it the way it's intended is am using it the way it's intended is am using it the way it's intended is pretty rough. Oh, cool. Soul is now done pretty rough. Oh, cool. Soul is now done pretty rough. Oh, cool. Soul is now done comparing things here. Yeah, this is comparing things here. Yeah, this is comparing things here. Yeah, this is what I expected. It's a slaughterhouse.

  26. what I expected. It's a slaughterhouse. what I expected. It's a slaughterhouse. It also decided to weight the different It also decided to weight the different It also decided to weight the different categories more and less heavily categories more and less heavily categories more and less heavily depending on how it felt about them. So, depending on how it felt about them. So, depending on how it felt about them. So, it said for the current repository and it said for the current repository and it said for the current repository and API fidelity that the plan from Muse is API fidelity that the plan from Muse is API fidelity that the plan from Muse is 4 out of 10. The plan from Fable is 7 4 out of 10. The plan from Fable is 7 4 out of 10. The plan from Fable is 7 out of 10. For the protocol research, out of 10. For the protocol research, out of 10. For the protocol research, they tied roughly. For end-to-end and they tied roughly. For end-to-end and they tied roughly. For end-to-end and multi-urface completeness, Muse was multi-urface completeness, Muse was multi-urface completeness, Muse was nowhere near there. Fable was a lot nowhere near there. Fable was a lot nowhere near there. Fable was a lot further along at 8 out of 10. Life cycle further along at 8 out of 10. Life cycle further along at 8 out of 10. Life cycle permission safety recovery pretty close. permission safety recovery pretty close. permission safety recovery pretty close. Delivery plan closeish with Fable Little Delivery plan closeish with Fable Little Delivery plan closeish with Fable Little had at seven versus five. Simplicity and had at seven versus five. Simplicity and had at seven versus five. Simplicity and maintainability both didn't score great. maintainability both didn't score great. maintainability both didn't score great. But overall, Muse's plan got a 4.8 out But overall, Muse's plan got a 4.8 out But overall, Muse's plan got a 4.8 out of 10 and Fables got a seven. Yeah, of 10 and Fables got a seven. Yeah, of 10 and Fables got a seven. Yeah, considering the gap in cost, that is considering the gap in cost, that is considering the gap in cost, that is reasonable. But considering the reality reasonable. But considering the reality reasonable. But considering the reality that you have to merge the code when that you have to merge the code when that you have to merge the code when it's done, this is much less reasonable. it's done, this is much less reasonable. it's done, this is much less reasonable. Still a very interesting model, just not Still a very interesting model, just not Still a very interesting model, just not necessarily necessarily necessarily one I would trust for making big one I would trust for making big one I would trust for making big sweeping changes to my code. Think now sweeping changes to my code. Think now sweeping changes to my code. Think now is a good time to answer the important is a good time to answer the important is a good time to answer the important question, why would someone use this question, why would someone use this question, why would someone use this model? What makes Muse 1.2 useful enough model? What makes Muse 1.2 useful enough model? What makes Muse 1.2 useful enough that someone should use it? Well, the that someone should use it? Well, the that someone should use it? Well, the obvious reason, like the number one obvious reason, like the number one obvious reason, like the number one thing that would make someone want to thing that would make someone want to thing that would make someone want to use this model is that they work at use this model is that they work at use this model is that they work at Meta, in which case they would probably Meta, in which case they would probably Meta, in which case they would probably use this as their second or third most use this as their second or third most use this as their second or third most used model compared to Opus and Fable used model compared to Opus and Fable used model compared to Opus and Fable because I know they love those models because I know they love those models because I know they love those models there. Main reason to use Muse is that there. Main reason to use Muse is that there. Main reason to use Muse is that you work at Meta. But there are other you work at Meta. But there are other you work at Meta. But there are other reasons that I'm seeing a bit of. The reasons that I'm seeing a bit of. The reasons that I'm seeing a bit of. The biggest, of course, is that you like to biggest, of course, is that you like to biggest, of course, is that you like to try new things, especially when those try new things, especially when those try new things, especially when those things are cheap. Because this model is

  27. things are cheap. Because this model is things are cheap. Because this model is cheap. Even if you're not willing to cheap. Even if you're not willing to cheap. Even if you're not willing to share your data and you're using it on share your data and you're using it on share your data and you're using it on the paid tier that's a little higher up, the paid tier that's a little higher up, the paid tier that's a little higher up, it's still a very, very cheap model. it's still a very, very cheap model. it's still a very, very cheap model. Although, now that I've switched out of Although, now that I've switched out of Although, now that I've switched out of that way cheaper tier, I just went from that way cheaper tier, I just went from that way cheaper tier, I just went from 40 to $5.32 for the work I've done in 40 to $5.32 for the work I've done in 40 to $5.32 for the work I've done in the past 10 minutes because it's going the past 10 minutes because it's going the past 10 minutes because it's going so fast. But it did make its integration so fast. But it did make its integration so fast. But it did make its integration for Muse inside of T3 code. So, I will for Muse inside of T3 code. So, I will for Muse inside of T3 code. So, I will ask it to spin up dev server and share a ask it to spin up dev server and share a ask it to spin up dev server and share a link. If it got this working properly link. If it got this working properly link. If it got this working properly first try after a little bit of plan first try after a little bit of plan first try after a little bit of plan review from Fable, then credit to them. review from Fable, then credit to them. review from Fable, then credit to them. They made a model that's pretty good. If They made a model that's pretty good. If They made a model that's pretty good. If this fails out right, then expected this fails out right, then expected this fails out right, then expected models this cheap can't really do tasks models this cheap can't really do tasks models this cheap can't really do tasks this exploratory and undefined because this exploratory and undefined because this exploratory and undefined because again they didn't put out the SDK I need again they didn't put out the SDK I need again they didn't put out the SDK I need to build this. So the model had to to build this. So the model had to to build this. So the model had to integrate Muse into T3 code with hacks integrate Muse into T3 code with hacks integrate Muse into T3 code with hacks more than anything. Moment of truth. more than anything. Moment of truth. more than anything. Moment of truth. Let's see if we are in here. Muse Let's see if we are in here. Muse Let's see if we are in here. Muse Spark12. What is this project? Not Spark12. What is this project? Not Spark12. What is this project? Not looking good so far. [snorts] Didn't looking good so far. [snorts] Didn't looking good so far. [snorts] Didn't even add it in the provider section in even add it in the provider section in even add it in the provider section in settings. It just added it in the UI settings. It just added it in the UI settings. It just added it in the UI here and it does not appear to work.

  28. here and it does not appear to work. here and it does not appear to work. Ah, was a nice attempt. The Ah, was a nice attempt. The Ah, was a nice attempt. The autogenerated title was from codeex autogenerated title was from codeex autogenerated title was from codeex because that is the default. If you have because that is the default. If you have because that is the default. If you have codecs, I just use 56 Luna on low for codecs, I just use 56 Luna on low for codecs, I just use 56 Luna on low for it. So, the title gen was nothing to do it. So, the title gen was nothing to do it. So, the title gen was nothing to do with Muse. It just appears to be broken. with Muse. It just appears to be broken. with Muse. It just appears to be broken. Okay. So uh sadly what this means is you Okay. So uh sadly what this means is you Okay. So uh sadly what this means is you cannot trust it for longer running cannot trust it for longer running cannot trust it for longer running things for sure. I was hoping it would things for sure. I was hoping it would things for sure. I was hoping it would be a little more capable at that but it be a little more capable at that but it be a little more capable at that but it is not. So what other reasons would you is not. So what other reasons would you is not. So what other reasons would you use this? You like really fast models use this? You like really fast models use this? You like really fast models that are somewhat capable. Like this is that are somewhat capable. Like this is that are somewhat capable. Like this is not anywhere near as good as a model not anywhere near as good as a model not anywhere near as good as a model like 55. Even though the benches suggest like 55. Even though the benches suggest like 55. Even though the benches suggest otherwise, it just doesn't get it as otherwise, it just doesn't get it as otherwise, it just doesn't get it as well. And that's what makes this model well. And that's what makes this model well. And that's what makes this model so strange to me is that it understands so strange to me is that it understands so strange to me is that it understands breaking up work in sub agents and not breaking up work in sub agents and not breaking up work in sub agents and not stepping on each other's toes when it stepping on each other's toes when it stepping on each other's toes when it does that type of thing really well. But does that type of thing really well. But does that type of thing really well. But it's nowhere near as good at actually it's nowhere near as good at actually it's nowhere near as good at actually seeing complex work through. It's almost seeing complex work through. It's almost seeing complex work through. It's almost like it it knows how to act like a like it it knows how to act like a like it it knows how to act like a modern smart model, but it doesn't know modern smart model, but it doesn't know modern smart model, but it doesn't know what the modern smart models know. I what the modern smart models know. I what the modern smart models know. I like the comparisons people are making like the comparisons people are making like the comparisons people are making with gro code fast because it does feel with gro code fast because it does feel with gro code fast because it does feel similar there where it has a lot of the similar there where it has a lot of the similar there where it has a lot of the the layers that make the models work the layers that make the models work the layers that make the models work well, but it's not good enough to really well, but it's not good enough to really well, but it's not good enough to really be trusted. And I'll be honest, it's be trusted. And I'll be honest, it's be trusted. And I'll be honest, it's hard for me to justify using models like hard for me to justify using models like hard for me to justify using models like this for a lot of my work just because I this for a lot of my work just because I this for a lot of my work just because I would rather wait two to three times would rather wait two to three times would rather wait two to three times longer and have something I can almost longer and have something I can almost longer and have something I can almost certainly merge versus trying it five certainly merge versus trying it five certainly merge versus trying it five times with a fast model and still have a times with a fast model and still have a times with a fast model and still have a mess inside of it. Honestly, the most mess inside of it. Honestly, the most mess inside of it. Honestly, the most impressive thing here is the CLI. It's a

  29. impressive thing here is the CLI. It's a impressive thing here is the CLI. It's a lot more stable and less annoying than lot more stable and less annoying than lot more stable and less annoying than Cloud Code. It's still not my favorite. Cloud Code. It's still not my favorite. Cloud Code. It's still not my favorite. Like Pi still smokes it overall, but Like Pi still smokes it overall, but Like Pi still smokes it overall, but it's solid for a thing that like they it's solid for a thing that like they it's solid for a thing that like they threw together for this release. I am threw together for this release. I am threw together for this release. I am tired of the labs making new CLIs and tired of the labs making new CLIs and tired of the labs making new CLIs and things when they're already in last things when they're already in last things when they're already in last place, forcing us to install yet another place, forcing us to install yet another place, forcing us to install yet another thing, but they did provide integrations thing, but they did provide integrations thing, but they did provide integrations for things like open code. So, that was for things like open code. So, that was for things like open code. So, that was nice and right directionish. And it nice and right directionish. And it nice and right directionish. And it still crushes everything Google is doing still crushes everything Google is doing still crushes everything Google is doing by quite a bit. It is funny to see Meta by quite a bit. It is funny to see Meta by quite a bit. It is funny to see Meta quickly leaprog Google but still be far quickly leaprog Google but still be far quickly leaprog Google but still be far behind everyone else. It's almost like behind everyone else. It's almost like behind everyone else. It's almost like these big companies are fighting to be these big companies are fighting to be these big companies are fighting to be like fifth place and all the startups like fifth place and all the startups like fifth place and all the startups that are really embracing the power of that are really embracing the power of that are really embracing the power of these new models are excelling in these new models are excelling in these new models are excelling in jumping far ahead. Obviously, Anthropic jumping far ahead. Obviously, Anthropic jumping far ahead. Obviously, Anthropic and OpenAI are far ahead, but companies and OpenAI are far ahead, but companies and OpenAI are far ahead, but companies like Moonshot with Kimmy, like Zai with like Moonshot with Kimmy, like Zai with like Moonshot with Kimmy, like Zai with the GLM series and more are all far the GLM series and more are all far the GLM series and more are all far ahead of this in my opinion. Hell, I ahead of this in my opinion. Hell, I ahead of this in my opinion. Hell, I would still use Grock 4.5 above this. I would still use Grock 4.5 above this. I would still use Grock 4.5 above this. I do have one last test I want to give it do have one last test I want to give it do have one last test I want to give it though. I want you to audit all of the though. I want you to audit all of the though. I want you to audit all of the poll requests I have open on T3 code.

  30. poll requests I have open on T3 code. poll requests I have open on T3 code. Figure out which ones are mergeable, Figure out which ones are mergeable, Figure out which ones are mergeable, which ones need more work, which ones which ones need more work, which ones which ones need more work, which ones should be closed, which ones have been should be closed, which ones have been should be closed, which ones have been trumped by other things merging, etc. I trumped by other things merging, etc. I trumped by other things merging, etc. I want you to make a priority list for me want you to make a priority list for me want you to make a priority list for me of what I should look at first and how of what I should look at first and how of what I should look at first and how confident you are in me merging it. confident you are in me merging it. confident you are in me merging it. Break up this work into lots of sub Break up this work into lots of sub Break up this work into lots of sub aents in order to get through it faster. aents in order to get through it faster. aents in order to get through it faster. Cool. We'll see how it does with that. Cool. We'll see how it does with that. Cool. We'll see how it does with that. Bad gateway when trying to hit the Bad gateway when trying to hit the Bad gateway when trying to hit the GitHub API through the CLI. That's fun. GitHub API through the CLI. That's fun. GitHub API through the CLI. That's fun. Is GitHub down? No. Is this formatting Is GitHub down? No. Is this formatting Is GitHub down? No. Is this formatting things wrong? Fun. It's doing some things wrong? Fun. It's doing some things wrong? Fun. It's doing some sketchy [ __ ] to get in. Interesting. It sketchy [ __ ] to get in. Interesting. It sketchy [ __ ] to get in. Interesting. It went through all my PRs in 4 minutes. went through all my PRs in 4 minutes. went through all my PRs in 4 minutes. That's like genuinely impressive. That's like genuinely impressive. That's like genuinely impressive. Especially cuz I spent the first two Especially cuz I spent the first two Especially cuz I spent the first two minutes just trying to get into GitHub minutes just trying to get into GitHub minutes just trying to get into GitHub and off it. If this review is of decent and off it. If this review is of decent and off it. If this review is of decent quality, then this might be what I use quality, then this might be what I use quality, then this might be what I use the model for. I might just use this as the model for. I might just use this as the model for. I might just use this as my go-to. like go review all the PRs I my go-to. like go review all the PRs I my go-to. like go review all the PRs I have open. Oh, that's a nice little have open. Oh, that's a nice little have open. Oh, that's a nice little thing. When I have Codeex and Claude thing. When I have Codeex and Claude thing. When I have Codeex and Claude make these PR audit pages, they often make these PR audit pages, they often make these PR audit pages, they often don't link the actual PRs with these don't link the actual PRs with these don't link the actual PRs with these like links here, and this did. That's like links here, and this did. That's like links here, and this did. That's really nice. I have been annoyed at how really nice. I have been annoyed at how really nice. I have been annoyed at how many models don't make these clickable many models don't make these clickable many models don't make these clickable links. Even gave little confidence links. Even gave little confidence links. Even gave little confidence scores. This is one of the better PR scores. This is one of the better PR scores. This is one of the better PR review pages I've gotten. For reference, review pages I've gotten. For reference, review pages I've gotten. For reference, here's one I made with a different here's one I made with a different here's one I made with a different model, and I actually had to tell it, model, and I actually had to tell it, model, and I actually had to tell it, "Please make sure these are links." I "Please make sure these are links." I "Please make sure these are links." I genuinely prefer the layout of the genuinely prefer the layout of the genuinely prefer the layout of the version that Muse did here by quite a version that Muse did here by quite a version that Muse did here by quite a bit. I might even use this as like my bit. I might even use this as like my bit. I might even use this as like my go-to template in the future. This is go-to template in the future. This is go-to template in the future. This is super readable to me. I like the little super readable to me. I like the little super readable to me. I like the little confidence scores. I like that it tells confidence scores. I like that it tells confidence scores. I like that it tells you if it's clean or dirty merge, what

  31. you if it's clean or dirty merge, what you if it's clean or dirty merge, what it thinks the status of things are. This it thinks the status of things are. This it thinks the status of things are. This is good. is good. is good. This might be one of the use cases I end This might be one of the use cases I end This might be one of the use cases I end up using this for a bunch. So to go back up using this for a bunch. So to go back up using this for a bunch. So to go back to use cases quick like fast models is to use cases quick like fast models is to use cases quick like fast models is one. One of the big ones I've now one. One of the big ones I've now one. One of the big ones I've now learned is for random code adjacent learned is for random code adjacent learned is for random code adjacent analysis work. Obviously I wouldn't analysis work. Obviously I wouldn't analysis work. Obviously I wouldn't trust this model to like actually go trust this model to like actually go trust this model to like actually go merge PRs for me. But as a surface level merge PRs for me. But as a surface level merge PRs for me. But as a surface level like go audit what's going on in this like go audit what's going on in this like go audit what's going on in this repo for 20 or 30 cents. That was really repo for 20 or 30 cents. That was really repo for 20 or 30 cents. That was really good especially with like no additional good especially with like no additional good especially with like no additional effort to try and make it better. Like effort to try and make it better. Like effort to try and make it better. Like that was great. I'm going to tell it to that was great. I'm going to tell it to that was great. I'm going to tell it to go go to go further here. How about you go go to go further here. How about you go go to go further here. How about you do a similar audit for all open PRs that do a similar audit for all open PRs that do a similar audit for all open PRs that have had updates in the past five days? have had updates in the past five days? have had updates in the past five days? If this audit costs less than like a If this audit costs less than like a If this audit costs less than like a dollar, it was absolutely worth it. One dollar, it was absolutely worth it. One dollar, it was absolutely worth it. One of the things that'll be hard for this of the things that'll be hard for this of the things that'll be hard for this model with this though is that it's not model with this though is that it's not model with this though is that it's not trained on GitHub because again, Meta trained on GitHub because again, Meta trained on GitHub because again, Meta mostly uses their internal mercurial mostly uses their internal mercurial mostly uses their internal mercurial stuff where all the other labs are heavy stuff where all the other labs are heavy stuff where all the other labs are heavy on GitHub. So, they've trained the model on GitHub. So, they've trained the model on GitHub. So, they've trained the model to be very good at GitHub. I'm about to to be very good at GitHub. I'm about to to be very good at GitHub. I'm about to hit so many rate limits. It doesn't seem hit so many rate limits. It doesn't seem hit so many rate limits. It doesn't seem to have parallel limits like all the to have parallel limits like all the to have parallel limits like all the other harnesses do. It's more than happy other harnesses do. It's more than happy other harnesses do. It's more than happy to run seven sub agents in the to run seven sub agents in the to run seven sub agents in the background at once. And it gave each of background at once. And it gave each of background at once. And it gave each of these a ton of PRs to look at. There's these a ton of PRs to look at. There's these a ton of PRs to look at. There's $841 before I hit send on this. We have $841 before I hit send on this. We have $841 before I hit send on this. We have spent another 5 cents since. It's going spent another 5 cents since. It's going spent another 5 cents since. It's going to hit rate limits. We'll let that run to hit rate limits. We'll let that run to hit rate limits. We'll let that run in the background. I'll come back to it in the background. I'll come back to it in the background. I'll come back to it if it ends up with good results. But if it ends up with good results. But if it ends up with good results. But back to why you would use it. I already back to why you would use it. I already back to why you would use it. I already said you like to try new things, but I said you like to try new things, but I said you like to try new things, but I really want to emphasize this point

  32. really want to emphasize this point really want to emphasize this point because this model has a different because this model has a different because this model has a different flavor. it like when we went through the flavor. it like when we went through the flavor. it like when we went through the different designs it made, it did things different designs it made, it did things different designs it made, it did things meaningfully differently. It still has meaningfully differently. It still has meaningfully differently. It still has the like early Gemini 3 Pro era style the like early Gemini 3 Pro era style the like early Gemini 3 Pro era style overall to it. Like a lot of these look overall to it. Like a lot of these look overall to it. Like a lot of these look like what I saw out of Gemini 3 and 3.1 like what I saw out of Gemini 3 and 3.1 like what I saw out of Gemini 3 and 3.1 Pro, but it also just has little Pro, but it also just has little Pro, but it also just has little nicities to it that give it a vibe nicities to it that give it a vibe nicities to it that give it a vibe that's different. It's it's like a that's different. It's it's like a that's different. It's it's like a little bit of seasoning that they added little bit of seasoning that they added little bit of seasoning that they added that other stuff doesn't have. This that other stuff doesn't have. This that other stuff doesn't have. This really is a model for enthusiasts right really is a model for enthusiasts right really is a model for enthusiasts right now. Like you want to go play with it now. Like you want to go play with it now. Like you want to go play with it because it's fun to play with new because it's fun to play with new because it's fun to play with new things, not because this model is going things, not because this model is going things, not because this model is going to magically save you a bunch of money to magically save you a bunch of money to magically save you a bunch of money or become your go-to. Like no one should or become your go-to. Like no one should or become your go-to. Like no one should use this model as their day-to-day use this model as their day-to-day use this model as their day-to-day coding model. But what it is is coding model. But what it is is coding model. But what it is is interesting. The things that Meta chose interesting. The things that Meta chose interesting. The things that Meta chose to focus on, the things they didn't to focus on, the things they didn't to focus on, the things they didn't focus on, and the capabilities that I'm focus on, and the capabilities that I'm focus on, and the capabilities that I'm seeing here, it's fascinating, seeing here, it's fascinating, seeing here, it's fascinating, genuinely, and I'm definitely going to genuinely, and I'm definitely going to genuinely, and I'm definitely going to play with this model more. Probably play with this model more. Probably play with this model more. Probably through a better harness, though, through a better harness, though, through a better harness, though, because you can't really integrate this because you can't really integrate this because you can't really integrate this one with anything. I don't know why they one with anything. I don't know why they one with anything. I don't know why they closed source the Muse CLI here. Like closed source the Muse CLI here. Like closed source the Muse CLI here. Like Muse code should just be open source.

  33. Muse code should just be open source. Muse code should just be open source. Meta, you guys know better. You're an Meta, you guys know better. You're an Meta, you guys know better. You're an open source company at heart. Just put open source company at heart. Just put open source company at heart. Just put out the source. In under five minutes, out the source. In under five minutes, out the source. In under five minutes, it was able to index and review 222 poll it was able to index and review 222 poll it was able to index and review 222 poll requests. And this cost me a total of 10 requests. And this cost me a total of 10 requests. And this cost me a total of 10 cents on the contributor tier. That is cents on the contributor tier. That is cents on the contributor tier. That is insane. Being able to go through that insane. Being able to go through that insane. Being able to go through that much real work, like auditing 222 pull much real work, like auditing 222 pull much real work, like auditing 222 pull requests in my codebase, organizing them requests in my codebase, organizing them requests in my codebase, organizing them by how mergeable it thinks they are, so by how mergeable it thinks they are, so by how mergeable it thinks they are, so I can quickly blast through this and I can quickly blast through this and I can quickly blast through this and ship real code. That's insane. That's ship real code. That's insane. That's ship real code. That's insane. That's actually valuable. And this is why it's actually valuable. And this is why it's actually valuable. And this is why it's fun to experiment with the models. Like fun to experiment with the models. Like fun to experiment with the models. Like try the different things you do and take try the different things you do and take try the different things you do and take a look at how much it costs and how fast a look at how much it costs and how fast a look at how much it costs and how fast it runs. Being able to hit a button and it runs. Being able to hit a button and it runs. Being able to hit a button and spend 10 cents and in five minutes you spend 10 cents and in five minutes you spend 10 cents and in five minutes you have a page like this for 200 plus poll have a page like this for 200 plus poll have a page like this for 200 plus poll requests on your project. That's good. requests on your project. That's good. requests on your project. That's good. That's useful. I'm impressed. I would That's useful. I'm impressed. I would That's useful. I'm impressed. I would use this regularly and I might even set use this regularly and I might even set use this regularly and I might even set something up to automatically do this something up to automatically do this something up to automatically do this for me every day. That's cool as [ __ ] for me every day. That's cool as [ __ ] for me every day. That's cool as [ __ ] Remember though that was on the Remember though that was on the Remember though that was on the contributor tier. So if you're not contributor tier. So if you're not contributor tier. So if you're not willing to share this data to Meta, willing to share this data to Meta, willing to share this data to Meta, you're going to be spending 20ish times you're going to be spending 20ish times you're going to be spending 20ish times more. But that goes from 10 cents to $2 more. But that goes from 10 cents to $2 more. But that goes from 10 cents to $2 for this type of work and this much for this type of work and this much for this type of work and this much work. That's genuinely really work. That's genuinely really work. That's genuinely really impressive. I I think you should play impressive. I I think you should play impressive. I I think you should play with this model if you're interested in with this model if you're interested in with this model if you're interested in this type of thing. Obviously, you this type of thing. Obviously, you this type of thing. Obviously, you shouldn't trust it for everything. You shouldn't trust it for everything. You shouldn't trust it for everything. You shouldn't just blindly go through and shouldn't just blindly go through and shouldn't just blindly go through and merge stuff, but as a way to like pull merge stuff, but as a way to like pull merge stuff, but as a way to like pull signals out of noise for really really signals out of noise for really really signals out of noise for really really cheap. It's a fun way to experiment and cheap. It's a fun way to experiment and cheap. It's a fun way to experiment and try things. stuff like generating try things. stuff like generating try things. stuff like generating titles, categorizing PRs, generating

  34. titles, categorizing PRs, generating titles, categorizing PRs, generating summaries for reports, digging through summaries for reports, digging through summaries for reports, digging through logs to find useful stuff. Like this is logs to find useful stuff. Like this is logs to find useful stuff. Like this is solid and doing similar work with Fable solid and doing similar work with Fable solid and doing similar work with Fable cost tens if not hundreds of dollars and cost tens if not hundreds of dollars and cost tens if not hundreds of dollars and this was literally 10 cents. So yeah, this was literally 10 cents. So yeah, this was literally 10 cents. So yeah, not a bad model. I think it's going to not a bad model. I think it's going to not a bad model. I think it's going to get a lot of [ __ ] because it has all the get a lot of [ __ ] because it has all the get a lot of [ __ ] because it has all the weird quirks it has, but when you think weird quirks it has, but when you think weird quirks it has, but when you think about it for the price, the speed, and about it for the price, the speed, and about it for the price, the speed, and the capability, as well as it weird the capability, as well as it weird the capability, as well as it weird things it seems to do not better, but things it seems to do not better, but things it seems to do not better, but different from other stuff, it's fun. It different from other stuff, it's fun. It different from other stuff, it's fun. It almost is like playing with one of those almost is like playing with one of those almost is like playing with one of those like interesting toy programming like interesting toy programming like interesting toy programming languages is how it feels to me. It's languages is how it feels to me. It's languages is how it feels to me. It's it's different in a way that isn't it's different in a way that isn't it's different in a way that isn't necessarily ready for real world work necessarily ready for real world work necessarily ready for real world work all the time, but it's cool as [ __ ] and all the time, but it's cool as [ __ ] and all the time, but it's cool as [ __ ] and very fun to play with. I definitely had very fun to play with. I definitely had very fun to play with. I definitely had ups and downs with this exploration, but ups and downs with this exploration, but ups and downs with this exploration, but overall I'm coming out somewhat overall I'm coming out somewhat overall I'm coming out somewhat impressed. I don't think this model is impressed. I don't think this model is impressed. I don't think this model is going to kill Opus or Fable anytime going to kill Opus or Fable anytime going to kill Opus or Fable anytime soon, but it's a model that I could see soon, but it's a model that I could see soon, but it's a model that I could see myself playing with for lots of weird myself playing with for lots of weird myself playing with for lots of weird things, especially when you consider the things, especially when you consider the things, especially when you consider the price. If I didn't have all of these price. If I didn't have all of these price. If I didn't have all of these accounts across cloud code and codecs accounts across cloud code and codecs accounts across cloud code and codecs that I just burn for all sorts of stuff, that I just burn for all sorts of stuff, that I just burn for all sorts of stuff, I would probably be leaning on models I would probably be leaning on models I would probably be leaning on models like this. And even though I have those like this. And even though I have those like this. And even though I have those other things I can burn, I am still other things I can burn, I am still other things I can burn, I am still going to be trying this to try and just going to be trying this to try and just going to be trying this to try and just organize my work and life better because organize my work and life better because organize my work and life better because it is so surprisingly cheap. I'm curious it is so surprisingly cheap. I'm curious it is so surprisingly cheap. I'm curious how y'all feel though. Am I too Frontier how y'all feel though. Am I too Frontier how y'all feel though. Am I too Frontier Pilled or is this model just not that Pilled or is this model just not that Pilled or is this model just not that impressive? Let me know how y'all feel impressive? Let me know how y'all feel impressive? Let me know how y'all feel and if you'll be using it. And until and if you'll be using it. And until and if you'll be using it. And until next time, peace nerds.

Summary

The main theme is Meta's continued contributions to the tech ecosystem, particularly in AI, despite past criticisms. Key subjects include their open-weight models like Llama and the new Muse Code, a coding assistant powered by Muse Spark 1.2. The practical takeaway is that while Meta's offerings may not always lead benchmarks, their focus on accessible, useful, and well-priced AI tools, like Muse Code, makes them significant players to watch.

View original episode ↗