← Back
Theo July 17, 2026 41m

Kimi K3 is the best model ever made (sometimes)

Read full transcript 31 segments
  1. KimmyK3 is here and it is a huge leap KimmyK3 is here and it is a huge leap for open weight models. I'm going to be for open weight models. I'm going to be for open weight models. I'm going to be honest and tell you guys that I just honest and tell you guys that I just honest and tell you guys that I just haven't been that hyped about open haven't been that hyped about open haven't been that hyped about open weight stuff recently because it hasn't weight stuff recently because it hasn't weight stuff recently because it hasn't been close to where we're at with models been close to where we're at with models been close to where we're at with models like Fable and GPT-5.6. The open weight like Fable and GPT-5.6. The open weight like Fable and GPT-5.6. The open weight frontier caught up to where we were frontier caught up to where we were frontier caught up to where we were before with models like Opus 4.8, kind before with models like Opus 4.8, kind before with models like Opus 4.8, kind of, but nothing has come close to of, but nothing has come close to of, but nothing has come close to surpassing it, especially for day-to-day surpassing it, especially for day-to-day surpassing it, especially for day-to-day use on complex coding work, especially use on complex coding work, especially use on complex coding work, especially once you start orchestrating really long once you start orchestrating really long once you start orchestrating really long runs where agents spin up tons of runs where agents spin up tons of runs where agents spin up tons of sub-agents for complex tasks. I think sub-agents for complex tasks. I think sub-agents for complex tasks. I think this might have changed because Kimmy K3 this might have changed because Kimmy K3 this might have changed because Kimmy K3 is genuinely on the line for frontier, is genuinely on the line for frontier, is genuinely on the line for frontier, if not surpassing where we're already at if not surpassing where we're already at if not surpassing where we're already at for various different things. The for various different things. The for various different things. The benchmarks are showing some pretty benchmarks are showing some pretty benchmarks are showing some pretty absurd numbers with KimmyK3 beating out absurd numbers with KimmyK3 beating out absurd numbers with KimmyK3 beating out GPT-5.6 sole in various tasks, as well GPT-5.6 sole in various tasks, as well GPT-5.6 sole in various tasks, as well as Fable and others. And at the very as Fable and others. And at the very as Fable and others. And at the very least, it's neck and neck throughout least, it's neck and neck throughout least, it's neck and neck throughout pretty much every bench I've seen. The pretty much every bench I've seen. The pretty much every bench I've seen. The benchmarks only tell one part of the benchmarks only tell one part of the benchmarks only tell one part of the story though, so I spent the whole day story though, so I spent the whole day story though, so I spent the whole day building as much as I could with KimmyK3 building as much as I could with KimmyK3 building as much as I could with KimmyK3 just to push it to its limits and see just to push it to its limits and see just to push it to its limits and see what it's capable of. And I'm going to what it's capable of. And I'm going to what it's capable of. And I'm going to be real, I'm blown away. There are be real, I'm blown away. There are be real, I'm blown away. There are definitely some rough edges and I'll do definitely some rough edges and I'll do definitely some rough edges and I'll do my best to show you guys how to work my best to show you guys how to work my best to show you guys how to work around them. My terminal froze because I around them. My terminal froze because I around them. My terminal froze because I was doing so much though, so I'm going was doing so much though, so I'm going was doing so much though, so I'm going to have to fix that first. This video is to have to fix that first. This video is to have to fix that first. This video is going to have a lot of fun in it from going to have a lot of fun in it from going to have a lot of fun in it from how to maximize your usage of the model how to maximize your usage of the model how to maximize your usage of the model to addressing the confusion around the to addressing the confusion around the to addressing the confusion around the different ways to use the model because different ways to use the model because different ways to use the model because there are quite a bit. Talking about how there are quite a bit. Talking about how there are quite a bit. Talking about how the world is seeing this and what the the world is seeing this and what the the world is seeing this and what the impact might be both on how we do dev impact might be both on how we do dev impact might be both on how we do dev work, as well as the economy, but also work, as well as the economy, but also work, as well as the economy, but also possibly most importantly, the security possibly most importantly, the security possibly most importantly, the security implications of a release like this implications of a release like this implications of a release like this because this model is going to be open because this model is going to be open because this model is going to be open weight. And when you have a model this weight. And when you have a model this weight. And when you have a model this capable with no restrictions, there are

  2. capable with no restrictions, there are capable with no restrictions, there are some real concerns we're going to have some real concerns we're going to have some real concerns we're going to have to address. I'm going to go fix my to address. I'm going to go fix my to address. I'm going to go fix my terminal and while I'm doing that, I terminal and while I'm doing that, I terminal and while I'm doing that, I hope you don't mind a quick break for hope you don't mind a quick break for hope you don't mind a quick break for today's sponsor. If you use GitHub today's sponsor. If you use GitHub today's sponsor. If you use GitHub Actions or Docker, trust me, you're Actions or Docker, trust me, you're Actions or Docker, trust me, you're going to want to watch this one because going to want to watch this one because going to want to watch this one because Depot is today's sponsor and they made Depot is today's sponsor and they made Depot is today's sponsor and they made both way, way better. Depot is fully both way, way better. Depot is fully both way, way better. Depot is fully compatible with GitHub Actions, but they compatible with GitHub Actions, but they compatible with GitHub Actions, but they also built their own alternative CI also built their own alternative CI also built their own alternative CI engine that is way faster and it can engine that is way faster and it can engine that is way faster and it can also be called from your coding agents also be called from your coding agents also be called from your coding agents using a CLI, which allows your agents to using a CLI, which allows your agents to using a CLI, which allows your agents to get feedback much faster than they would get feedback much faster than they would get feedback much faster than they would if they had to run all that stuff if they had to run all that stuff if they had to run all that stuff locally or wait for your PR to build it locally or wait for your PR to build it locally or wait for your PR to build it for you. If you do use them for your for you. If you do use them for your for you. If you do use them for your normal actions, you'll still see crazy normal actions, you'll still see crazy normal actions, you'll still see crazy speed-ups, so up to 10 times faster. speed-ups, so up to 10 times faster. speed-ups, so up to 10 times faster. Docker's where they shine even more Docker's where they shine even more Docker's where they shine even more though, making your Docker builds 40 though, making your Docker builds 40 though, making your Docker builds 40 times faster for real-world use cases, times faster for real-world use cases, times faster for real-world use cases, and not just in the cloud, on your and not just in the cloud, on your and not just in the cloud, on your machine, too. The Depot CLI is a drop-in machine, too. The Depot CLI is a drop-in machine, too. The Depot CLI is a drop-in replacement for the Docker CLI that replacement for the Docker CLI that replacement for the Docker CLI that caches all of the layers on their CDN, caches all of the layers on their CDN, caches all of the layers on their CDN, which means all of your employees that which means all of your employees that which means all of your employees that are building the same images can all are building the same images can all are building the same images can all have way faster builds fetching from have way faster builds fetching from have way faster builds fetching from that cache instead of having to do the that cache instead of having to do the that cache instead of having to do the whole thing on their machine. And whole thing on their machine. And whole thing on their machine. And somehow this all just got even faster somehow this all just got even faster somehow this all just got even faster with Depot Metal. As a friend of the with Depot Metal. As a friend of the with Depot Metal. As a friend of the Depot team, I am blown away at how far Depot team, I am blown away at how far Depot team, I am blown away at how far they went with Depot Metal. They went as they went with Depot Metal. They went as they went with Depot Metal. They went as deep as they possibly could on AWS, deep as they possibly could on AWS, deep as they possibly could on AWS, specking out machines directly that they specking out machines directly that they specking out machines directly that they own and control, managing the storage own and control, managing the storage own and control, managing the storage themselves as well. And the results show themselves as well. And the results show themselves as well. And the results show why they made these changes. It's why they made these changes. It's why they made these changes. It's already 30% faster for existing Depot CI already 30% faster for existing Depot CI already 30% faster for existing Depot CI or Sandbox runs. But more importantly, or Sandbox runs. But more importantly, or Sandbox runs. But more importantly, they managed to move from their roughly they managed to move from their roughly they managed to move from their roughly 10-second VM spin-up time to sub-second 10-second VM spin-up time to sub-second 10-second VM spin-up time to sub-second spin-ups, which is crazy for CI jobs.

  3. spin-ups, which is crazy for CI jobs. spin-ups, which is crazy for CI jobs. There's a reason companies like Posthog There's a reason companies like Posthog There's a reason companies like Posthog and PlanetScale do all of their builds and PlanetScale do all of their builds and PlanetScale do all of their builds on Depot, and you can figure it out on Depot, and you can figure it out on Depot, and you can figure it out yourself at soydev.link/depot. yourself at soydev.link/depot. yourself at soydev.link/depot. In case you thought I was joking, I'm In case you thought I was joking, I'm In case you thought I was joking, I'm not. I actually have to force quit tmux not. I actually have to force quit tmux not. I actually have to force quit tmux right now in order to do what I was right now in order to do what I was right now in order to do what I was working on. Let's start with what the working on. Let's start with what the working on. Let's start with what the official Moonshot team had to say about official Moonshot team had to say about official Moonshot team had to say about this release. Today, we're introducing this release. Today, we're introducing this release. Today, we're introducing Kimmi K3, our most capable model. Kimmi Kimmi K3, our most capable model. Kimmi Kimmi K3, our most capable model. Kimmi K3 is a 2.8 trillion parameter model K3 is a 2.8 trillion parameter model K3 is a 2.8 trillion parameter model built on our Kimmi Delta attention and built on our Kimmi Delta attention and built on our Kimmi Delta attention and attention residuals, with native vision attention residuals, with native vision attention residuals, with native vision capabilities and a 1 million token capabilities and a 1 million token capabilities and a 1 million token context window. These are two really context window. These are two really context window. These are two really nice, big changes. Things like GLM-52 nice, big changes. Things like GLM-52 nice, big changes. Things like GLM-52 don't have vision at all, which is one don't have vision at all, which is one don't have vision at all, which is one of the most annoying parts of using of the most annoying parts of using of the most annoying parts of using them. And the 1 million token context them. And the 1 million token context them. And the 1 million token context window is also super useful, even if window is also super useful, even if window is also super useful, even if it's not available in the subscriptions it's not available in the subscriptions it's not available in the subscriptions we'll talk about in a bit. It's very we'll talk about in a bit. It's very we'll talk about in a bit. It's very nice to have when you do need it, and nice to have when you do need it, and nice to have when you do need it, and this model is huge, so it'll be this model is huge, so it'll be this model is huge, so it'll be beneficial for a lot of different beneficial for a lot of different beneficial for a lot of different things. It's crazy how just a year ago things. It's crazy how just a year ago things. It's crazy how just a year ago the Kimmi models had no vision, had the Kimmi models had no vision, had the Kimmi models had no vision, had short context, and didn't even have short context, and didn't even have short context, and didn't even have reasoning. And they've somehow caught up reasoning. And they've somehow caught up reasoning. And they've somehow caught up to the frontier in that time. Saying to the frontier in that time. Saying to the frontier in that time. Saying that they just raised around $2 billion that they just raised around $2 billion that they just raised around $2 billion and have raised almost $4 billion total and have raised almost $4 billion total and have raised almost $4 billion total makes sense that they're aiming for the makes sense that they're aiming for the makes sense that they're aiming for the stars. Like this is a moonshot in the stars. Like this is a moonshot in the stars. Like this is a moonshot in the most literal sense, especially at the most literal sense, especially at the most literal sense, especially at the size of 2.8 trillion parameters. If size of 2.8 trillion parameters. If size of 2.8 trillion parameters. If you're curious how big this is, the you're curious how big this is, the you're curious how big this is, the rough math for FP8 is that a trillion rough math for FP8 is that a trillion rough math for FP8 is that a trillion parameters is roughly a terabyte. So 2.8 parameters is roughly a terabyte. So 2.8 parameters is roughly a terabyte. So 2.8 trillion is 2.8 terabytes of data. At trillion is 2.8 terabytes of data. At trillion is 2.8 terabytes of data. At FP8 it gets cut roughly in half, so it's FP8 it gets cut roughly in half, so it's FP8 it gets cut roughly in half, so it's only 1.4 terabytes. Oh man, that's so only 1.4 terabytes. Oh man, that's so only 1.4 terabytes. Oh man, that's so much more reasonable. To be very, very much more reasonable. To be very, very much more reasonable. To be very, very clear, anyone who's telling you that clear, anyone who's telling you that clear, anyone who's telling you that this is the future of local models has

  4. this is the future of local models has this is the future of local models has no idea what they're talking about and no idea what they're talking about and no idea what they're talking about and they should be ignored forever because a they should be ignored forever because a they should be ignored forever because a 1.4 terabyte model is not fitting in 1.4 terabyte model is not fitting in 1.4 terabyte model is not fitting in memory on any computer owned by anyone memory on any computer owned by anyone memory on any computer owned by anyone watching this. And if I'm wrong about watching this. And if I'm wrong about watching this. And if I'm wrong about that, please contact me. I would love to that, please contact me. I would love to that, please contact me. I would love to borrow your machines for some fun work. borrow your machines for some fun work. borrow your machines for some fun work. It seriously though, this is not running It seriously though, this is not running It seriously though, this is not running on anything any of us have in our homes on anything any of us have in our homes on anything any of us have in our homes unless you happen to live in like the unless you happen to live in like the unless you happen to live in like the Colossus data center. Colossus data center. Colossus data center. This model requires supercomputers to be This model requires supercomputers to be This model requires supercomputers to be used. We don't know how big the closed used. We don't know how big the closed used. We don't know how big the closed weight models from Frontier Labs are, weight models from Frontier Labs are, weight models from Frontier Labs are, but we've seen estimates between 1 and 3 but we've seen estimates between 1 and 3 but we've seen estimates between 1 and 3 trillion per ams for a model like Opus, trillion per ams for a model like Opus, trillion per ams for a model like Opus, usually in the 1 to 2 trillion range. So usually in the 1 to 2 trillion range. So usually in the 1 to 2 trillion range. So having a roughly three trill model having a roughly three trill model having a roughly three trill model that's open weight is insane. This is a that's open weight is insane. This is a that's open weight is insane. This is a massive leap in the size of models that massive leap in the size of models that massive leap in the size of models that are available for us to use, and I are available for us to use, and I are available for us to use, and I genuinely feel bad for our friends over genuinely feel bad for our friends over genuinely feel bad for our friends over at Hugging Face having to host this and at Hugging Face having to host this and at Hugging Face having to host this and deal people downloading 2 plus terabytes deal people downloading 2 plus terabytes deal people downloading 2 plus terabytes of data to use it. They won't have to of data to use it. They won't have to of data to use it. They won't have to for a bit though because their planned for a bit though because their planned for a bit though because their planned release date for the weights is July release date for the weights is July release date for the weights is July 27th. So the weights aren't available 27th. So the weights aren't available 27th. So the weights aren't available yet, which sadly means we have to use yet, which sadly means we have to use yet, which sadly means we have to use their APIs, which uh their APIs, which uh their APIs, which uh they're a Chinese company, so take that they're a Chinese company, so take that they're a Chinese company, so take that as you will. They're going to have as you will. They're going to have as you will. They're going to have access to your code. Some people freak access to your code. Some people freak access to your code. Some people freak out about this and I can understand why.

  5. out about this and I can understand why. out about this and I can understand why. Let's talk about the performance a bit. Let's talk about the performance a bit. Let's talk about the performance a bit. They say that it still trails the most They say that it still trails the most They say that it still trails the most powerful proprietary models like Fable 5 powerful proprietary models like Fable 5 powerful proprietary models like Fable 5 and 5.6 Soul, but it also demonstrates and 5.6 Soul, but it also demonstrates and 5.6 Soul, but it also demonstrates frontier level performance across their frontier level performance across their frontier level performance across their evaluation suite consistently evaluation suite consistently evaluation suite consistently outperforming other tested models. I outperforming other tested models. I outperforming other tested models. I mentioned these in the intro, but we'll mentioned these in the intro, but we'll mentioned these in the intro, but we'll go through the benches real quick here. go through the benches real quick here. go through the benches real quick here. Deep SWE, which is my current favorite Deep SWE, which is my current favorite Deep SWE, which is my current favorite software dev bench, shows 5.6 Soul software dev bench, shows 5.6 Soul software dev bench, shows 5.6 Soul slightly ahead of Fable 5, 73 to 70, and slightly ahead of Fable 5, 73 to 70, and slightly ahead of Fable 5, 73 to 70, and then Kimi K3 behind them at roughly the then Kimi K3 behind them at roughly the then Kimi K3 behind them at roughly the same rate at a 67.5, which puts them same rate at a 67.5, which puts them same rate at a 67.5, which puts them ahead of GPT-5.5, Opus-4.8, and GLM-5.2. ahead of GPT-5.5, Opus-4.8, and GLM-5.2. ahead of GPT-5.5, Opus-4.8, and GLM-5.2. Where things start to get more Where things start to get more Where things start to get more interesting is when we go down a bit to interesting is when we go down a bit to interesting is when we go down a bit to frontier SWE. I'm coming around to this frontier SWE. I'm coming around to this frontier SWE. I'm coming around to this bench because it seems more focused on bench because it seems more focused on bench because it seems more focused on how likely the code is to merge, not how likely the code is to merge, not how likely the code is to merge, not just how well it solves the problem, and just how well it solves the problem, and just how well it solves the problem, and Fable 5 definitely writes the most Fable 5 definitely writes the most Fable 5 definitely writes the most mergeable code in my opinion in my mergeable code in my opinion in my mergeable code in my opinion in my real-world projects. And now that I've real-world projects. And now that I've real-world projects. And now that I've read more of the bench, I understand why read more of the bench, I understand why read more of the bench, I understand why it scores like this. I still think it's it scores like this. I still think it's it scores like this. I still think it's a little higher than it should be, and a little higher than it should be, and a little higher than it should be, and the gap's bigger than I would expect in the gap's bigger than I would expect in the gap's bigger than I would expect in real-world usage, but Kimi K3 sliding in real-world usage, but Kimi K3 sliding in real-world usage, but Kimi K3 sliding in between Fable and Soul with a 10-point between Fable and Soul with a 10-point between Fable and Soul with a 10-point lead on Soul is kind of nuts. It lead on Soul is kind of nuts. It lead on Soul is kind of nuts. It suggests this model is more tasteful is suggests this model is more tasteful is suggests this model is more tasteful is the best I can put it. Like it writes the best I can put it. Like it writes the best I can put it. Like it writes code that better It just looks and feels code that better It just looks and feels code that better It just looks and feels better, rather than just finding a better, rather than just finding a better, rather than just finding a solution to the problem. They have their solution to the problem. They have their solution to the problem. They have their own internal bench where it scored just own internal bench where it scored just own internal bench where it scored just ahead of Opus-4.8 and behind K3, but ahead of Opus-4.8 and behind K3, but ahead of Opus-4.8 and behind K3, but it's worth noting their internal bench it's worth noting their internal bench it's worth noting their internal bench puts Soul behind 5.5, so I don't know puts Soul behind 5.5, so I don't know puts Soul behind 5.5, so I don't know how trustworthy it is. And then over how trustworthy it is. And then over how trustworthy it is. And then over here, we get the terminal bench where it here, we get the terminal bench where it here, we get the terminal bench where it beats out Opus and Fable 5 and is just

  6. beats out Opus and Fable 5 and is just beats out Opus and Fable 5 and is just barely behind Soul. Program bench, which barely behind Soul. Program bench, which barely behind Soul. Program bench, which I haven't really seen much of, but it is I haven't really seen much of, but it is I haven't really seen much of, but it is world-class there, just barely beating world-class there, just barely beating world-class there, just barely beating out Soul, which beats out Fable 5 by a out Soul, which beats out Fable 5 by a out Soul, which beats out Fable 5 by a little more. And SWE marathon, a new little more. And SWE marathon, a new little more. And SWE marathon, a new bench that I know a lot are fond of, bench that I know a lot are fond of, bench that I know a lot are fond of, somehow Opus-4.8 was the lead before, somehow Opus-4.8 was the lead before, somehow Opus-4.8 was the lead before, even against Soul and Fable 5, but Kimi even against Soul and Fable 5, but Kimi even against Soul and Fable 5, but Kimi K3 has now come out in front. General K3 has now come out in front. General K3 has now come out in front. General agent evals are also quite interesting, agent evals are also quite interesting, agent evals are also quite interesting, things like GDP val, job bench, things like GDP val, job bench, things like GDP val, job bench, spreadsheet bench, which is now spreadsheet bench, which is now spreadsheet bench, which is now led by Kimi K3. To whoever is really led by Kimi K3. To whoever is really led by Kimi K3. To whoever is really into spreadsheets and open-weight into spreadsheets and open-weight into spreadsheets and open-weight models, it must be a phenomenal day for models, it must be a phenomenal day for models, it must be a phenomenal day for you. Congratulations. you. Congratulations. you. Congratulations. Browse comp scoring so high is of the Browse comp scoring so high is of the Browse comp scoring so high is of the most interesting to me though, because most interesting to me though, because most interesting to me though, because I'm a huge fan of browser use and I'm a huge fan of browser use and I'm a huge fan of browser use and computer use now after not caring about computer use now after not caring about computer use now after not caring about it for like 2 plus years, because the it for like 2 plus years, because the it for like 2 plus years, because the models got way better at it. It's much models got way better at it. It's much models got way better at it. It's much more interesting. Having an open weight more interesting. Having an open weight more interesting. Having an open weight model that can do this is genuinely model that can do this is genuinely model that can do this is genuinely fascinating, because that means fascinating, because that means fascinating, because that means hypothetically speaking, if I could hypothetically speaking, if I could hypothetically speaking, if I could afford the hardware to run it on, I afford the hardware to run it on, I afford the hardware to run it on, I could have a fully offline runner that could have a fully offline runner that could have a fully offline runner that can control my computer and do real work can control my computer and do real work can control my computer and do real work without having to send my data to without having to send my data to without having to send my data to Anthropic or Open AI. Generally, if you Anthropic or Open AI. Generally, if you Anthropic or Open AI. Generally, if you want to do actual computer use work want to do actual computer use work want to do actual computer use work right now, you're just expecting to send right now, you're just expecting to send right now, you're just expecting to send all of those screenshots of your machine all of those screenshots of your machine all of those screenshots of your machine to one of those labs. Not great. You to one of those labs. Not great. You to one of those labs. Not great. You hear solution there. It's expensive, but hear solution there. It's expensive, but hear solution there. It's expensive, but in the future if the costs come down and in the future if the costs come down and in the future if the costs come down and the opportunity use bottles of this the opportunity use bottles of this the opportunity use bottles of this capability comes more available, that is capability comes more available, that is capability comes more available, that is huge. I mentioned before that it has huge. I mentioned before that it has huge. I mentioned before that it has vision capabilities and it seems to be vision capabilities and it seems to be vision capabilities and it seems to be pretty solid there, too.

  7. pretty solid there, too. pretty solid there, too. But, what's really cool about the But, what's really cool about the But, what's really cool about the browser stuff as I was hinting at before browser stuff as I was hinting at before browser stuff as I was hinting at before is not just that it is industry leading is not just that it is industry leading is not just that it is industry leading by its score on these benches, it's also by its score on these benches, it's also by its score on these benches, it's also comically cheaper than the frontier is. comically cheaper than the frontier is. comically cheaper than the frontier is. It's funny cuz I was just glazing this It's funny cuz I was just glazing this It's funny cuz I was just glazing this chart in the GPT-56 video about how much chart in the GPT-56 video about how much chart in the GPT-56 video about how much better soul was compared to everything better soul was compared to everything better soul was compared to everything else on it. But, I don't know if that's else on it. But, I don't know if that's else on it. But, I don't know if that's fair anymore, cuz Kimi K3 max gets even fair anymore, cuz Kimi K3 max gets even fair anymore, cuz Kimi K3 max gets even cheaper than soul max, roughly the same cheaper than soul max, roughly the same cheaper than soul max, roughly the same price as soul high, but a meaningfully price as soul high, but a meaningfully price as soul high, but a meaningfully better score. That said, an open weight better score. That said, an open weight better score. That said, an open weight model that is priced as cheap as model that is priced as cheap as model that is priced as cheap as possible competing neck and neck with 56 possible competing neck and neck with 56 possible competing neck and neck with 56 soul on cost shows just how far Open AI soul on cost shows just how far Open AI soul on cost shows just how far Open AI has gone in reducing costs to the best has gone in reducing costs to the best has gone in reducing costs to the best of their ability. 56 soul is still the of their ability. 56 soul is still the of their ability. 56 soul is still the fewest tokens to solve these types of fewest tokens to solve these types of fewest tokens to solve these types of problems and benchmarks, but since the problems and benchmarks, but since the problems and benchmarks, but since the tokens are so much more expensive, Kimi tokens are so much more expensive, Kimi tokens are so much more expensive, Kimi K3 eeks out a win there. We'll talk K3 eeks out a win there. We'll talk K3 eeks out a win there. We'll talk about cost more in a little bit, but about cost more in a little bit, but about cost more in a little bit, but I'll just give the numbers now so you I'll just give the numbers now so you I'll just give the numbers now so you have them. 30 cents for a cash hit, $3 have them. 30 cents for a cash hit, $3 have them. 30 cents for a cash hit, $3 per million tokens in and $15 per per million tokens in and $15 per per million tokens in and $15 per million tokens out. For reference, the million tokens out. For reference, the million tokens out. For reference, the standard API price for solid models is standard API price for solid models is standard API price for solid models is $3 per million and 15 per mill out, but $3 per million and 15 per mill out, but $3 per million and 15 per mill out, but Anthropic is temporarily offering a Anthropic is temporarily offering a Anthropic is temporarily offering a discount of 2 per mil in and 10 per mil discount of 2 per mil in and 10 per mil discount of 2 per mil in and 10 per mil out right now. So, this is roughly a out right now. So, this is roughly a out right now. So, this is roughly a Sonnet-level priced model, but it Sonnet-level priced model, but it Sonnet-level priced model, but it doesn't have a lot of Sonnet's problems, doesn't have a lot of Sonnet's problems, doesn't have a lot of Sonnet's problems, which we'll definitely talk about in a which we'll definitely talk about in a which we'll definitely talk about in a bit. Kimmy K3 is available today on bit. Kimmy K3 is available today on bit. Kimmy K3 is available today on kimmi.com, Kimmy Work, Kimmy Code, and kimmi.com, Kimmy Work, Kimmy Code, and kimmi.com, Kimmy Work, Kimmy Code, and the Kimmy API. I would ignore the two in the Kimmy API. I would ignore the two in the Kimmy API. I would ignore the two in the middle here. I will talk about the the middle here. I will talk about the the middle here. I will talk about the kimmi.com and Kimmy API stuff in a bit kimmi.com and Kimmy API stuff in a bit kimmi.com and Kimmy API stuff in a bit when I talk about how to use the model.

  8. when I talk about how to use the model. when I talk about how to use the model. At launch, it will use max thinking At launch, it will use max thinking At launch, it will use max thinking effort by default with low and high effort by default with low and high effort by default with low and high effort modes to be introduced in effort modes to be introduced in effort modes to be introduced in subsequent updates. This is one of the subsequent updates. This is one of the subsequent updates. This is one of the most interesting things I saw about this most interesting things I saw about this most interesting things I saw about this release is it doesn't offer reasoning release is it doesn't offer reasoning release is it doesn't offer reasoning controls at all yet. You just use it and controls at all yet. You just use it and controls at all yet. You just use it and it gets used on max. They're currently it gets used on max. They're currently it gets used on max. They're currently working closely with inference partners working closely with inference partners working closely with inference partners and open-source maintainers to align and open-source maintainers to align and open-source maintainers to align technical details and ensure reliable technical details and ensure reliable technical details and ensure reliable rollout across the ecosystem. This is rollout across the ecosystem. This is rollout across the ecosystem. This is also exciting to see. I know the Kimmy also exciting to see. I know the Kimmy also exciting to see. I know the Kimmy guys have been pretty good about making guys have been pretty good about making guys have been pretty good about making it clear which providers are and aren't it clear which providers are and aren't it clear which providers are and aren't hosting their models correctly with hosting their models correctly with hosting their models correctly with benchmarks to verify the likelihood of benchmarks to verify the likelihood of benchmarks to verify the likelihood of any given provider actually hosting the any given provider actually hosting the any given provider actually hosting the model properly, which is great because model properly, which is great because model properly, which is great because this model is not trivial to host cuz this model is not trivial to host cuz this model is not trivial to host cuz it's an interesting implementation. it's an interesting implementation. it's an interesting implementation. We'll talk briefly about that in a bit. We'll talk briefly about that in a bit. We'll talk briefly about that in a bit. They also haven't put out their They also haven't put out their They also haven't put out their technical report yet, which will have a technical report yet, which will have a technical report yet, which will have a lot more of those details. Moonshot has lot more of those details. Moonshot has lot more of those details. Moonshot has historically had the biggest open weight historically had the biggest open weight historically had the biggest open weight models with the Kimmy line. They put out models with the Kimmy line. They put out models with the Kimmy line. They put out Kimmy K2 as a trillion per am model all Kimmy K2 as a trillion per am model all Kimmy K2 as a trillion per am model all the way back in July 11th of 2025, and the way back in July 11th of 2025, and the way back in July 11th of 2025, and they stayed at that size for all of they stayed at that size for all of they stayed at that size for all of their releases from that point forward their releases from that point forward their releases from that point forward until now, where Kimmy K3 is a huge jump until now, where Kimmy K3 is a huge jump until now, where Kimmy K3 is a huge jump of 2.8 trillion per ams, putting them of 2.8 trillion per ams, putting them of 2.8 trillion per ams, putting them ahead of everyone else, even huge models ahead of everyone else, even huge models ahead of everyone else, even huge models like DeepSeek V4, which was 1.6 trillion like DeepSeek V4, which was 1.6 trillion like DeepSeek V4, which was 1.6 trillion on the pro version. I feel bad for on the pro version. I feel bad for on the pro version. I feel bad for Thinking Machines. They were so hyped to Thinking Machines. They were so hyped to Thinking Machines. They were so hyped to put out Inkling, and it's only a trail put out Inkling, and it's only a trail put out Inkling, and it's only a trail per ams and didn't bench great, and now per ams and didn't bench great, and now per ams and didn't bench great, and now with this coming out the day after, oof, with this coming out the day after, oof, with this coming out the day after, oof, I feel bad for those investors even more I feel bad for those investors even more I feel bad for those investors even more so. K3 is built on their Kimmy Delta so. K3 is built on their Kimmy Delta so. K3 is built on their Kimmy Delta attention and attention residuals attention and attention residuals attention and attention residuals models, two architectural updates models, two architectural updates models, two architectural updates designed to improve how information designed to improve how information designed to improve how information flows across sequence length and model flows across sequence length and model flows across sequence length and model depth. They've also scaled up the depth. They've also scaled up the depth. They've also scaled up the mixture of expert sparsity, effectively

  9. mixture of expert sparsity, effectively mixture of expert sparsity, effectively activating 16 of 896 experts while activating 16 of 896 experts while activating 16 of 896 experts while paired with a stable latent MOE paired with a stable latent MOE paired with a stable latent MOE framework. Again, not going to be framework. Again, not going to be framework. Again, not going to be trivial to host this cuz they invented a trivial to host this cuz they invented a trivial to host this cuz they invented a lot of their own solutions to make a lot of their own solutions to make a lot of their own solutions to make a model this big and capable without the model this big and capable without the model this big and capable without the compute cost being absurd. The memory compute cost being absurd. The memory compute cost being absurd. The memory cost is still very high cuz you're going cost is still very high cuz you're going cost is still very high cuz you're going to need a lot of this in memory for it to need a lot of this in memory for it to need a lot of this in memory for it to make any sense at all, but the sparse to make any sense at all, but the sparse to make any sense at all, but the sparse nature of how these experts are nature of how these experts are nature of how these experts are traversed should hopefully keep costs traversed should hopefully keep costs traversed should hopefully keep costs from being too absurd once the model's from being too absurd once the model's from being too absurd once the model's in memory. When you combine all of those in memory. When you combine all of those in memory. When you combine all of those updates with the refined training and updates with the refined training and updates with the refined training and data recipes that they produced, they data recipes that they produced, they data recipes that they produced, they end up with a 2.5x improvement in end up with a 2.5x improvement in end up with a 2.5x improvement in overall scaling efficiency, which is a overall scaling efficiency, which is a overall scaling efficiency, which is a pretty big jump. pretty big jump. pretty big jump. These guys have always been pretty good These guys have always been pretty good These guys have always been pretty good about sharing their advancements and the about sharing their advancements and the about sharing their advancements and the cool things they do, so I'm extra cool things they do, so I'm extra cool things they do, so I'm extra excited for that technical report to excited for that technical report to excited for that technical report to come out. Now, let's talk about what come out. Now, let's talk about what come out. Now, let's talk about what using it looks like on the coding side using it looks like on the coding side using it looks like on the coding side especially. Kimi K3 has strong long especially. Kimi K3 has strong long especially. Kimi K3 has strong long horizon coding performance, meaning it horizon coding performance, meaning it horizon coding performance, meaning it can run long tasks with minimal human can run long tasks with minimal human can run long tasks with minimal human oversight. One of my favorite test tasks oversight. One of my favorite test tasks oversight. One of my favorite test tasks for these much bigger, more capable for these much bigger, more capable for these much bigger, more capable models is to take my old code base for models is to take my old code base for models is to take my old code base for ping.gg, which is a Zoom app for content ping.gg, which is a Zoom app for content ping.gg, which is a Zoom app for content creators doing live collaborations, and creators doing live collaborations, and creators doing live collaborations, and see if the model is capable of porting see if the model is capable of porting see if the model is capable of porting it. It's been running for like 3 plus it. It's been running for like 3 plus it. It's been running for like 3 plus hours, and it was doing great until like hours, and it was doing great until like hours, and it was doing great until like 10 minutes ago where it hit a limit on 10 minutes ago where it hit a limit on 10 minutes ago where it hit a limit on the context window size. This is a the context window size. This is a the context window size. This is a mistake that's partially my fault mistake that's partially my fault mistake that's partially my fault because I didn't realize that the because I didn't realize that the because I didn't realize that the subscriptions that you can use for Kimi subscriptions that you can use for Kimi subscriptions that you can use for Kimi code don't actually give you the full 1 code don't actually give you the full 1 code don't actually give you the full 1 million token context window, and with million token context window, and with million token context window, and with my setup using it with Claude code, it my setup using it with Claude code, it my setup using it with Claude code, it was hitting a much smaller limit. So, was hitting a much smaller limit. So, was hitting a much smaller limit. So, I'm going to really quickly change that

  10. I'm going to really quickly change that I'm going to really quickly change that max size and get that compacted so it max size and get that compacted so it max size and get that compacted so it can keep going because I want to see if can keep going because I want to see if can keep going because I want to see if that run can complete. Sadly, I won't be that run can complete. Sadly, I won't be that run can complete. Sadly, I won't be able to compact using Kimi K3 because able to compact using Kimi K3 because able to compact using Kimi K3 because again, it won't respond to the API again, it won't respond to the API again, it won't respond to the API request. I'll be able to get it request. I'll be able to get it request. I'll be able to get it compacted with Fable and keep the run compacted with Fable and keep the run compacted with Fable and keep the run going in just a moment though. The fact going in just a moment though. The fact going in just a moment though. The fact that it was able to get through 122 that it was able to get through 122 that it was able to get through 122 tasks with a single like paragraph and a tasks with a single like paragraph and a tasks with a single like paragraph and a half prompt and no additional insight or half prompt and no additional insight or half prompt and no additional insight or effort from me is unbelievable though. effort from me is unbelievable though. effort from me is unbelievable though. I've never seen an open weight model I've never seen an open weight model I've never seen an open weight model come close to staying coherent for even come close to staying coherent for even come close to staying coherent for even half of this length, much less like half of this length, much less like half of this length, much less like actual real-world massive migration actual real-world massive migration actual real-world massive migration work. It's impressive. And since it also work. It's impressive. And since it also work. It's impressive. And since it also has visual reasoning, it's way more has visual reasoning, it's way more has visual reasoning, it's way more capable of UI type work, too. It capable of UI type work, too. It capable of UI type work, too. It leverages screenshots and visuals to leverages screenshots and visuals to leverages screenshots and visuals to optimize game dev, front end, and CAD. optimize game dev, front end, and CAD. optimize game dev, front end, and CAD. And those visual capabilities are nuts And those visual capabilities are nuts And those visual capabilities are nuts when it comes to front end code. We'll when it comes to front end code. We'll when it comes to front end code. We'll talk more about this later, I'm sure, talk more about this later, I'm sure, talk more about this later, I'm sure, but just know in advance that Gemini K3 but just know in advance that Gemini K3 but just know in advance that Gemini K3 is really, really good at front end, at is really, really good at front end, at is really, really good at front end, at least according to Arena AI. From my least according to Arena AI. From my least according to Arena AI. From my experience using it, it has its quirks, experience using it, it has its quirks, experience using it, it has its quirks, but it is very impressive, and I cannot but it is very impressive, and I cannot but it is very impressive, and I cannot wait to show you guys just how cool some wait to show you guys just how cool some wait to show you guys just how cool some of the UI it creates is. On the very of the UI it creates is. On the very of the UI it creates is. On the very opposite end, we have the kernel opposite end, we have the kernel opposite end, we have the kernel optimization capabilities, where optimization capabilities, where optimization capabilities, where apparently they use it to optimize GPU apparently they use it to optimize GPU apparently they use it to optimize GPU kernels, and it did a pretty damn good kernels, and it did a pretty damn good kernels, and it did a pretty damn good job. After being active for around 15 job. After being active for around 15 job. After being active for around 15 hours, it saw slightly better hours, it saw slightly better hours, it saw slightly better improvements than even Fable did, and improvements than even Fable did, and improvements than even Fable did, and quite a bit better than 5.5 and 5.6 quite a bit better than 5.5 and 5.6 quite a bit better than 5.5 and 5.6 Soul, somehow, which is nuts. They have Soul, somehow, which is nuts. They have Soul, somehow, which is nuts. They have four types of kernel optimizations that four types of kernel optimizations that four types of kernel optimizations that they benched here, and somehow K3 came they benched here, and somehow K3 came they benched here, and somehow K3 came out near Frontier or above Frontier in out near Frontier or above Frontier in out near Frontier or above Frontier in all of them. It's also really

  11. all of them. It's also really all of them. It's also really interesting to see how big the gap is interesting to see how big the gap is interesting to see how big the gap is between Soul and Fable in a lot of these between Soul and Fable in a lot of these between Soul and Fable in a lot of these as well. This might be a good bench for as well. This might be a good bench for as well. This might be a good bench for them to actually like put out in the them to actually like put out in the them to actually like put out in the future, not just cuz they're leading it, future, not just cuz they're leading it, future, not just cuz they're leading it, but because it's fascinating to see how but because it's fascinating to see how but because it's fascinating to see how big the gaps are. This is also scary for big the gaps are. This is also scary for big the gaps are. This is also scary for companies like Anthropic who have went companies like Anthropic who have went companies like Anthropic who have went out of their way to hide the model's out of their way to hide the model's out of their way to hide the model's capability of helping with ML work like capability of helping with ML work like capability of helping with ML work like kernel optimization for hosting models. kernel optimization for hosting models. kernel optimization for hosting models. They went out of their way to keep the They went out of their way to keep the They went out of their way to keep the model from sharing those things, and model from sharing those things, and model from sharing those things, and it's one of the restrictions they have it's one of the restrictions they have it's one of the restrictions they have on both Fable and Mythos. So, it's on both Fable and Mythos. So, it's on both Fable and Mythos. So, it's fascinating to see Gemini bragging about fascinating to see Gemini bragging about fascinating to see Gemini bragging about how good their model is at this when how good their model is at this when how good their model is at this when they plan to release the weights, they plan to release the weights, they plan to release the weights, because that means this capability is because that means this capability is because that means this capability is now in the hands of everyone to an now in the hands of everyone to an now in the hands of everyone to an extent. Man, Anthropic has to be extent. Man, Anthropic has to be extent. Man, Anthropic has to be terrified of this release more than terrified of this release more than terrified of this release more than anybody. In the late stages of Gemini K3 anybody. In the late stages of Gemini K3 anybody. In the late stages of Gemini K3 development, an early version of K3 development, an early version of K3 development, an early version of K3 handled the majority of the team's handled the majority of the team's handled the majority of the team's kernel optimization work. That's pretty kernel optimization work. That's pretty kernel optimization work. That's pretty nuts. They also tested if it could a nuts. They also tested if it could a nuts. They also tested if it could a complete GPU programming stack and complete GPU programming stack and complete GPU programming stack and compiler from scratch. They developed compiler from scratch. They developed compiler from scratch. They developed many Triton, a compact Triton-like many Triton, a compact Triton-like many Triton, a compact Triton-like compiler with its own tile level IR compiler with its own tile level IR compiler with its own tile level IR layer over MLIR. I'm sure much smarter layer over MLIR. I'm sure much smarter layer over MLIR. I'm sure much smarter people will know what this all means and people will know what this all means and people will know what this all means and be really pumped about it.

  12. be really pumped about it. be really pumped about it. Yeah, it seems like they did a really Yeah, it seems like they did a really Yeah, it seems like they did a really good job training the model to do these good job training the model to do these good job training the model to do these type of stuff. Here's where I'll be much type of stuff. Here's where I'll be much type of stuff. Here's where I'll be much more useful. Game dev and digital more useful. Game dev and digital more useful. Game dev and digital creation. K3 combines strong 3D creation. K3 combines strong 3D creation. K3 combines strong 3D reasoning, coding, and vision reasoning, coding, and vision reasoning, coding, and vision capabilities to turn concepts, images, capabilities to turn concepts, images, capabilities to turn concepts, images, and videos into fully playable and videos into fully playable and videos into fully playable interactive experiences. It achieves a interactive experiences. It achieves a interactive experiences. It achieves a true vision in the loop by seamlessly true vision in the loop by seamlessly true vision in the loop by seamlessly iterating between code and live iterating between code and live iterating between code and live screenshots, instantly seeing and screenshots, instantly seeing and screenshots, instantly seeing and refining outputs. I actually watched it refining outputs. I actually watched it refining outputs. I actually watched it do this live when I was having it do do this live when I was having it do do this live when I was having it do some changes to T3 code, where it would some changes to T3 code, where it would some changes to T3 code, where it would spin it up, get a screenshot, look at spin it up, get a screenshot, look at spin it up, get a screenshot, look at it, think about it, and then change what it, think about it, and then change what it, think about it, and then change what it was doing. Apparently, it created it was doing. Apparently, it created it was doing. Apparently, it created everything here, including the rider and everything here, including the rider and everything here, including the rider and horse models. So, it like horse models. So, it like horse models. So, it like actually understands 3D? I'm going to actually understands 3D? I'm going to actually understands 3D? I'm going to have to play more. I want to see this in have to play more. I want to see this in have to play more. I want to see this in action. I just spun up Pi to go do a 3D action. I just spun up Pi to go do a 3D action. I just spun up Pi to go do a 3D port of Fish Slap. We'll see how it port of Fish Slap. We'll see how it port of Fish Slap. We'll see how it goes. Hopefully, not well, cuz I don't goes. Hopefully, not well, cuz I don't goes. Hopefully, not well, cuz I don't want to have to record more after I'm want to have to record more after I'm want to have to record more after I'm leaving. I'm actually supposed to be at leaving. I'm actually supposed to be at leaving. I'm actually supposed to be at an event right now, but had to film this an event right now, but had to film this an event right now, but had to film this because it's such a cool model. because it's such a cool model. because it's such a cool model. Yeah, the fact that it's able to do this Yeah, the fact that it's able to do this Yeah, the fact that it's able to do this type of 3D work and modeling is crazy.

  13. type of 3D work and modeling is crazy. type of 3D work and modeling is crazy. Every model I've tried so far is just so Every model I've tried so far is just so Every model I've tried so far is just so rough at 3D stuff that I've been rough at 3D stuff that I've been rough at 3D stuff that I've been impressed by like circles being placed impressed by like circles being placed impressed by like circles being placed properly sometimes. properly sometimes. properly sometimes. If this model can actually do 3D, I'm If this model can actually do 3D, I'm If this model can actually do 3D, I'm going to have to going to have to going to have to that's going to change things. Oh, [ __ ] that's going to change things. Oh, [ __ ] that's going to change things. Oh, [ __ ] Oh, [ __ ] Oh, [ __ ] Oh, [ __ ] Theo from the future here. It got pretty Theo from the future here. It got pretty Theo from the future here. It got pretty far in Fish Slap 3D, so I wanted to far in Fish Slap 3D, so I wanted to far in Fish Slap 3D, so I wanted to check this out quick. check this out quick. check this out quick. Holy [ __ ] The 3D game is like actually Holy [ __ ] The 3D game is like actually Holy [ __ ] The 3D game is like actually working. Uh working. Uh working. Uh Sorry about the audio. Can I mute that Sorry about the audio. Can I mute that Sorry about the audio. Can I mute that easily? No, I can't. easily? No, I can't. easily? No, I can't. I have no idea what it sounds like. I'm I have no idea what it sounds like. I'm I have no idea what it sounds like. I'm not putting on my headphones to see I I not putting on my headphones to see I I not putting on my headphones to see I I I'm I'm I'm I'm putting on I'm so curious. >> Holy [ __ ] the fish textures. They're >> Holy [ __ ] the fish textures. They're not perfect at all, but this is the best not perfect at all, but this is the best not perfect at all, but this is the best fish model I've seen any lab create. And fish model I've seen any lab create. And fish model I've seen any lab create. And the submarine is phenomenal, too. the submarine is phenomenal, too. the submarine is phenomenal, too. For a model generated for a 3D web game For a model generated for a 3D web game For a model generated for a 3D web game like this. like this. like this. It's got its issues, for sure, but like It's got its issues, for sure, but like It's got its issues, for sure, but like it's still working on it. It just said it's still working on it. It just said it's still working on it. It just said like the game works, let me do more. But like the game works, let me do more. But like the game works, let me do more. But the fact that it's already this far, the fact that it's already this far, the fact that it's already this far, unbelievable.

  14. Okay. Okay. Holy [ __ ] Yeah. Yeah. It even has sound effects and things. It even has sound effects and things. It even has sound effects and things. This is so much better than I would have This is so much better than I would have This is so much better than I would have expected. It's still working like I'm expected. It's still working like I'm expected. It's still working like I'm checking this before it's done, and it checking this before it's done, and it checking this before it's done, and it keeps pulling in visualizations like it keeps pulling in visualizations like it keeps pulling in visualizations like it ran it, and I saw the little picture in ran it, and I saw the little picture in ran it, and I saw the little picture in the pie history where it actually like the pie history where it actually like the pie history where it actually like pulled it up and was like, "It's pulled it up and was like, "It's pulled it up and was like, "It's working. Let's continue. Let's fix all working. Let's continue. Let's fix all working. Let's continue. Let's fix all these things." these things." these things." It's a good [ __ ] model. Apparently, It's a good [ __ ] model. Apparently, It's a good [ __ ] model. Apparently, it's also good at chip design. This is it's also good at chip design. This is it's also good at chip design. This is interesting. It might have real like interesting. It might have real like interesting. It might have real like beyond what the frontier currently beyond what the frontier currently beyond what the frontier currently allows capabilities that none of the allows capabilities that none of the allows capabilities that none of the other frontier labs are focusing on. other frontier labs are focusing on. other frontier labs are focusing on. This is actually arguably one of the This is actually arguably one of the This is actually arguably one of the problems with this fixation on code that problems with this fixation on code that problems with this fixation on code that both OpenAI and Anthropic have as they both OpenAI and Anthropic have as they both OpenAI and Anthropic have as they try to win an enterprise. They're not try to win an enterprise. They're not try to win an enterprise. They're not finding new capabilities the same way finding new capabilities the same way finding new capabilities the same way that they used to. Here, it seems like that they used to. Here, it seems like that they used to. Here, it seems like this new frontier open weight model is this new frontier open weight model is this new frontier open weight model is actually able to explore things that the actually able to explore things that the actually able to explore things that the labs here have just not explored. labs here have just not explored. labs here have just not explored. Hopefully, as long as the 3D stuff is as Hopefully, as long as the 3D stuff is as Hopefully, as long as the 3D stuff is as cool as it seems, that's going to be cool as it seems, that's going to be cool as it seems, that's going to be huge. It's also incredible at knowledge huge. It's also incredible at knowledge huge. It's also incredible at knowledge work, according to them. Benchmarks like work, according to them. Benchmarks like work, according to them. Benchmarks like online EXP bench, deck bench, and online EXP bench, deck bench, and online EXP bench, deck bench, and finance bench, it is industry leading finance bench, it is industry leading finance bench, it is industry leading in.

  15. in. in. A lot of that's probably because of how A lot of that's probably because of how A lot of that's probably because of how good it is at visualization stuff. They good it is at visualization stuff. They good it is at visualization stuff. They had it do a bunch of research work, and had it do a bunch of research work, and had it do a bunch of research work, and it did a pretty good job even at it did a pretty good job even at it did a pretty good job even at generating the actual reports, which is generating the actual reports, which is generating the actual reports, which is pretty damn cool. Good at infographic pretty damn cool. Good at infographic pretty damn cool. Good at infographic style presentations. It's still got that style presentations. It's still got that style presentations. It's still got that like AI-generated vibe that a lot of like AI-generated vibe that a lot of like AI-generated vibe that a lot of those have. They show some dashboards. It does They show some dashboards. It does really like this um really like this um really like this um Bento box style UI layout, but it makes Bento box style UI layout, but it makes Bento box style UI layout, but it makes them in a very pretty way, and the them in a very pretty way, and the them in a very pretty way, and the animation taste is actually quite good, animation taste is actually quite good, animation taste is actually quite good, too. I've been impressed with it. too. I've been impressed with it. too. I've been impressed with it. Apparently, it's also good at video Apparently, it's also good at video Apparently, it's also good at video editing, which is crazy. I will not have editing, which is crazy. I will not have editing, which is crazy. I will not have any time to try this anytime soon, but any time to try this anytime soon, but any time to try this anytime soon, but uh if my team ends up liking it, I'll be uh if my team ends up liking it, I'll be uh if my team ends up liking it, I'll be sure to share that in the future. I'm sure to share that in the future. I'm sure to share that in the future. I'm not that into AI-based video editing cuz not that into AI-based video editing cuz not that into AI-based video editing cuz video editing is actually quite fun and video editing is actually quite fun and video editing is actually quite fun and not too tedious if you are any good at not too tedious if you are any good at not too tedious if you are any good at it at all. And they had edit a lot of it at all. And they had edit a lot of it at all. And they had edit a lot of their videos for the launch of the their videos for the launch of the their videos for the launch of the model. Having it go through lots of model. Having it go through lots of model. Having it go through lots of clips, handling clip selection, motion clips, handling clip selection, motion clips, handling clip selection, motion matched cuts, frame accurate beat matched cuts, frame accurate beat matched cuts, frame accurate beat synchronization, audio processing, and synchronization, audio processing, and synchronization, audio processing, and multiple rounds of revision. That's multiple rounds of revision. That's multiple rounds of revision. That's pretty cool. They have some cool info pretty cool. They have some cool info pretty cool. They have some cool info here about how they were able to get a here about how they were able to get a here about how they were able to get a model of this size to be trainable in a model of this size to be trainable in a model of this size to be trainable in a reasonable time frame compute-wise in in reasonable time frame compute-wise in in reasonable time frame compute-wise in in a stable fashion cuz the bigger it gets, a stable fashion cuz the bigger it gets, a stable fashion cuz the bigger it gets, the harder it gets. They found some fun the harder it gets. They found some fun the harder it gets. They found some fun tricks like using FP4 weights for the tricks like using FP4 weights for the tricks like using FP4 weights for the actual stored weights and params, but actual stored weights and params, but actual stored weights and params, but then using FP8 for the activations once then using FP8 for the activations once then using FP8 for the activations once the model's actually running so that it the model's actually running so that it the model's actually running so that it has better short-term memory, and it has better short-term memory, and it has better short-term memory, and it doesn't have to waste a ton of RAM in doesn't have to waste a ton of RAM in doesn't have to waste a ton of RAM in order to preserve the actual weights. At order to preserve the actual weights. At order to preserve the actual weights. At least that's my rough understanding.

  16. least that's my rough understanding. least that's my rough understanding. Smarter people, feel free to correct me Smarter people, feel free to correct me Smarter people, feel free to correct me in the comments. Remember that in the comments. Remember that in the comments. Remember that supercomputer thing I said? Since supercomputer thing I said? Since supercomputer thing I said? Since inference efficiency likewise benefits inference efficiency likewise benefits inference efficiency likewise benefits from larger high-bandwidth communication from larger high-bandwidth communication from larger high-bandwidth communication domains, we recommend deploying K3 on domains, we recommend deploying K3 on domains, we recommend deploying K3 on supernode configurations with 64 or more supernode configurations with 64 or more supernode configurations with 64 or more accelerators. 64 H100s is a bit rough. accelerators. 64 H100s is a bit rough. accelerators. 64 H100s is a bit rough. That's 2.6 million to host the model. That's 2.6 million to host the model. That's 2.6 million to host the model. Yeah. Yeah. Yeah. Very local-friendly. So, that's what we Very local-friendly. So, that's what we Very local-friendly. So, that's what we have from them. Let's take a quick look have from them. Let's take a quick look have from them. Let's take a quick look at what others have said. I mentioned at what others have said. I mentioned at what others have said. I mentioned Arena AI gave it an absurd score for Arena AI gave it an absurd score for Arena AI gave it an absurd score for front-end stuff. We'll show some front front-end stuff. We'll show some front front-end stuff. We'll show some front ends in a bit. We're going to start with ends in a bit. We're going to start with ends in a bit. We're going to start with Artificial Analysis because the model, Artificial Analysis because the model, Artificial Analysis because the model, according to them, is the third smartest according to them, is the third smartest according to them, is the third smartest ever. So, pretty much out of nowhere, we ever. So, pretty much out of nowhere, we ever. So, pretty much out of nowhere, we had two drops in a row that took the had two drops in a row that took the had two drops in a row that took the frontier away from OpenAI and Fable just frontier away from OpenAI and Fable just frontier away from OpenAI and Fable just going back-to-back forever. And suddenly going back-to-back forever. And suddenly going back-to-back forever. And suddenly we have Grok 4.5 and Kimmy K3 coming up we have Grok 4.5 and Kimmy K3 coming up we have Grok 4.5 and Kimmy K3 coming up for those third-place spots right behind for those third-place spots right behind for those third-place spots right behind Soul and Fable. Normally Opus would be Soul and Fable. Normally Opus would be Soul and Fable. Normally Opus would be up there, too, pushing these back, and up there, too, pushing these back, and up there, too, pushing these back, and it's not anymore because both xAI and it's not anymore because both xAI and it's not anymore because both xAI and Kimmy have gotten their [ __ ] so together Kimmy have gotten their [ __ ] so together Kimmy have gotten their [ __ ] so together that they are leapfrogging. And they are that they are leapfrogging. And they are that they are leapfrogging. And they are so far ahead of what Google is cooking, so far ahead of what Google is cooking, so far ahead of what Google is cooking, it's hilarious. I honestly think Google it's hilarious. I honestly think Google it's hilarious. I honestly think Google if they had any reasonable way to buy a if they had any reasonable way to buy a if they had any reasonable way to buy a Chinese company like Moonshot, they Chinese company like Moonshot, they Chinese company like Moonshot, they probably have to at this point cuz they probably have to at this point cuz they probably have to at this point cuz they are just so behind in comparison. Okay, are just so behind in comparison. Okay, are just so behind in comparison. Okay, slight correction. Grok 4.5 is behind slight correction. Grok 4.5 is behind slight correction. Grok 4.5 is behind both Tara and Opus, as well as 5.5, so both Tara and Opus, as well as 5.5, so both Tara and Opus, as well as 5.5, so it's not really up at this frontier it's not really up at this frontier it's not really up at this frontier level, but Kimmy K3 is. It is just

  17. level, but Kimmy K3 is. It is just level, but Kimmy K3 is. It is just behind Soul and Fable according to behind Soul and Fable according to behind Soul and Fable according to Artificial Analysis's benches. Artificial Analysis's benches. Artificial Analysis's benches. But things get much more interesting as But things get much more interesting as But things get much more interesting as we dig in more. I'll read what they said we dig in more. I'll read what they said we dig in more. I'll read what they said on Twitter first. Kimmy K3 scores 57 on on Twitter first. Kimmy K3 scores 57 on on Twitter first. Kimmy K3 scores 57 on the Artificial Analysis Intelligence the Artificial Analysis Intelligence the Artificial Analysis Intelligence Index. Its intelligence is comparable to Index. Its intelligence is comparable to Index. Its intelligence is comparable to Opus 4.8 and 5.5, but remains slightly Opus 4.8 and 5.5, but remains slightly Opus 4.8 and 5.5, but remains slightly behind Fable 5 and 5.6 Soul. Moonshot behind Fable 5 and 5.6 Soul. Moonshot behind Fable 5 and 5.6 Soul. Moonshot has expressed plans to release the 2.8 has expressed plans to release the 2.8 has expressed plans to release the 2.8 trillion parameter model's weights, trillion parameter model's weights, trillion parameter model's weights, which would make it the leading open which would make it the leading open which would make it the leading open weight model. weight model. weight model. It's got really strong agentic It's got really strong agentic It's got really strong agentic performance, as I mentioned before, it's performance, as I mentioned before, it's performance, as I mentioned before, it's meaningfully better than models like GLM meaningfully better than models like GLM meaningfully better than models like GLM 5.2, going from 1668 on GDP val from 5.2, going from 1668 on GDP val from 5.2, going from 1668 on GDP val from 1514, which was the open weight frontier 1514, which was the open weight frontier 1514, which was the open weight frontier before, and even ahead of models like before, and even ahead of models like before, and even ahead of models like Opus 4.8. Opus 4.8. Opus 4.8. It's the second highest score I've ever It's the second highest score I've ever It's the second highest score I've ever seen on Artificial Analysis's knowledge seen on Artificial Analysis's knowledge seen on Artificial Analysis's knowledge bench called briefcase. It will bench called briefcase. It will bench called briefcase. It will absolutely be the leading open weight absolutely be the leading open weight absolutely be the leading open weight model, not a surprise there. The cost model, not a surprise there. The cost model, not a surprise there. The cost per task is similar to 5.6 Soul, but per task is similar to 5.6 Soul, but per task is similar to 5.6 Soul, but it's half the price of Opus and higher it's half the price of Opus and higher it's half the price of Opus and higher than all of the open weight peers. Not than all of the open weight peers. Not than all of the open weight peers. Not just cuz the model pricing is more just cuz the model pricing is more just cuz the model pricing is more expensive, but also remember the amount expensive, but also remember the amount expensive, but also remember the amount of tokens it uses is important, too. For of tokens it uses is important, too. For of tokens it uses is important, too. For reference, in output tokens per task, reference, in output tokens per task, reference, in output tokens per task, Kimmy K3 is a decent bit lower than Kimmy K3 is a decent bit lower than Kimmy K3 is a decent bit lower than Fable 5, which is the most token Fable 5, which is the most token Fable 5, which is the most token efficient model in Fable's put out in a efficient model in Fable's put out in a efficient model in Fable's put out in a while, but compared to things like 5 6 while, but compared to things like 5 6 while, but compared to things like 5 6 Terra or even Soul all the way back here Terra or even Soul all the way back here Terra or even Soul all the way back here at 15k tokens per task, 23 is more, but at 15k tokens per task, 23 is more, but at 15k tokens per task, 23 is more, but also not much more when you consider the also not much more when you consider the also not much more when you consider the price difference and it's so much less price difference and it's so much less price difference and it's so much less than other models, especially other than other models, especially other than other models, especially other noisy open weight ones, things like noisy open weight ones, things like noisy open weight ones, things like DeepSeek V4 Pro or even worse V4 Flash, DeepSeek V4 Pro or even worse V4 Flash, DeepSeek V4 Pro or even worse V4 Flash, which was 45k tokens for the same tasks.

  18. which was 45k tokens for the same tasks. which was 45k tokens for the same tasks. It's pretty token efficient and that's It's pretty token efficient and that's It's pretty token efficient and that's awesome to see because historically only awesome to see because historically only awesome to see because historically only OpenAI has really focused on token OpenAI has really focused on token OpenAI has really focused on token efficiency. Now we have both Grok and efficiency. Now we have both Grok and efficiency. Now we have both Grok and Kimi surrounding them in their little Kimi surrounding them in their little Kimi surrounding them in their little cheap token section here, finally cheap token section here, finally cheap token section here, finally driving the industry towards more driving the industry towards more driving the industry towards more efficient reasoning. 21% more efficient efficient reasoning. 21% more efficient efficient reasoning. 21% more efficient than K2 6 was even though the model is than K2 6 was even though the model is than K2 6 was even though the model is bigger and the tokens are more bigger and the tokens are more bigger and the tokens are more expensive. It has native multimodal expensive. It has native multimodal expensive. It has native multimodal capabilities as I mentioned before, capabilities as I mentioned before, capabilities as I mentioned before, very, very good there. The AI Omniscient very, very good there. The AI Omniscient very, very good there. The AI Omniscient score is also very interesting. If score is also very interesting. If score is also very interesting. If you're not familiar with this bench, you're not familiar with this bench, you're not familiar with this bench, it's Artificial Analysis' attempt to it's Artificial Analysis' attempt to it's Artificial Analysis' attempt to measure hallucinations, specifically if measure hallucinations, specifically if measure hallucinations, specifically if the model is going to tell you when it the model is going to tell you when it the model is going to tell you when it doesn't know something versus will it doesn't know something versus will it doesn't know something versus will it make up something instead. And Kimi K3 make up something instead. And Kimi K3 make up something instead. And Kimi K3 is one of the best open weight models at is one of the best open weight models at is one of the best open weight models at not hallucinating. In fact, I think it not hallucinating. In fact, I think it not hallucinating. In fact, I think it is the best by quite a bit. Even 5 2 is is the best by quite a bit. Even 5 2 is is the best by quite a bit. Even 5 2 is a much worse score here. And this test a much worse score here. And this test a much worse score here. And this test is fun because you can get into the is fun because you can get into the is fun because you can get into the positive when you say I don't know the positive when you say I don't know the positive when you say I don't know the answer and you go negative when you lie. answer and you go negative when you lie. answer and you go negative when you lie. So even some very good models like 5 6 So even some very good models like 5 6 So even some very good models like 5 6 Soul end up scoring a lot lower than Soul end up scoring a lot lower than Soul end up scoring a lot lower than they should here because they're a they should here because they're a they should here because they're a little too quick to lie when they don't little too quick to lie when they don't little too quick to lie when they don't know and the extra points they get for know and the extra points they get for know and the extra points they get for the things they do know get canceled out the things they do know get canceled out the things they do know get canceled out a bit by that. And 5 6 Luna gets hit a bit by that. And 5 6 Luna gets hit a bit by that. And 5 6 Luna gets hit real hard with this, getting into the real hard with this, getting into the real hard with this, getting into the negative as a result. One of the most negative as a result. One of the most negative as a result. One of the most honest open weight models we've ever honest open weight models we've ever honest open weight models we've ever seen and that's genuinely exciting cuz seen and that's genuinely exciting cuz seen and that's genuinely exciting cuz open weight models have not been good at open weight models have not been good at open weight models have not been good at this historically. But if your goal is this historically. But if your goal is this historically. But if your goal is to use this model to save a bunch of to use this model to save a bunch of to use this model to save a bunch of money, you probably shouldn't be too money, you probably shouldn't be too money, you probably shouldn't be too excited yet because that $15 out cost is excited yet because that $15 out cost is excited yet because that $15 out cost is not cheap and since it uses twice as not cheap and since it uses twice as not cheap and since it uses twice as many tokens as something like Soul, many tokens as something like Soul, many tokens as something like Soul, cancels out the 50% discount that you're cancels out the 50% discount that you're cancels out the 50% discount that you're getting and you also lose a lot of the

  19. getting and you also lose a lot of the getting and you also lose a lot of the niceties that you get from Modern niceties that you get from Modern niceties that you get from Modern Frontier Labs. Moonshot themselves even Frontier Labs. Moonshot themselves even Frontier Labs. Moonshot themselves even said that the model still has a said that the model still has a said that the model still has a noticeable gap in user experience noticeable gap in user experience noticeable gap in user experience compared with Fable 5 and 5.6 Soul. So, compared with Fable 5 and 5.6 Soul. So, compared with Fable 5 and 5.6 Soul. So, if Soul is as cheap if not cheaper for if Soul is as cheap if not cheaper for if Soul is as cheap if not cheaper for your use case, it might end up being your use case, it might end up being your use case, it might end up being better just cuz it's nicer to work with. better just cuz it's nicer to work with. better just cuz it's nicer to work with. And I have noticed that with Kimi And I have noticed that with Kimi And I have noticed that with Kimi myself. I noticed things like when a myself. I noticed things like when a myself. I noticed things like when a workflow finishes, it responds to the workflow finishes, it responds to the workflow finishes, it responds to the workflow finishing instead of giving me workflow finishing instead of giving me workflow finishing instead of giving me the context I need. Those are things you the context I need. Those are things you the context I need. Those are things you can work around, but you're going to can work around, but you're going to can work around, but you're going to have to do that with this model because have to do that with this model because have to do that with this model because it's not as R L on the expectations we it's not as R L on the expectations we it's not as R L on the expectations we have as users and they just don't have have as users and they just don't have have as users and they just don't have as much data cuz they're not getting as much data cuz they're not getting as much data cuz they're not getting feedback from people the same way cuz feedback from people the same way cuz feedback from people the same way cuz the vast majority of users of the Kimi the vast majority of users of the Kimi the vast majority of users of the Kimi models are using them on other models are using them on other models are using them on other providers. So, they don't get any of the providers. So, they don't get any of the providers. So, they don't get any of the telemetry they would need to improve telemetry they would need to improve telemetry they would need to improve these things. I'm sure it will get these things. I'm sure it will get these things. I'm sure it will get better over time, but not surprisingly better over time, but not surprisingly better over time, but not surprisingly that there's a real usability gap here. that there's a real usability gap here. that there's a real usability gap here. They also call out that it's too They also call out that it's too They also call out that it's too proactive. Reminds you of a certain proactive. Reminds you of a certain proactive. Reminds you of a certain Rottweiler model, as well as the Rottweiler model, as well as the Rottweiler model, as well as the sensitivity to thinking history that it sensitivity to thinking history that it sensitivity to thinking history that it needs its thinking. Thankfully, it's an needs its thinking. Thankfully, it's an needs its thinking. Thankfully, it's an open weight model, so we get the open weight model, so we get the open weight model, so we get the thinking data, too. It's not like thinking data, too. It's not like thinking data, too. It's not like Infropic or Open AI where the thinking Infropic or Open AI where the thinking Infropic or Open AI where the thinking is hidden on some server. All is hidden on some server. All is hidden on some server. All reasonable, really cool call outs. I reasonable, really cool call outs. I reasonable, really cool call outs. I love this level of transparency from a love this level of transparency from a love this level of transparency from a lab that's releasing something this lab that's releasing something this lab that's releasing something this important. Okay, enough of the research important. Okay, enough of the research important. Okay, enough of the research side. Let's talk about the actual side. Let's talk about the actual side. Let's talk about the actual front-end capabilities cuz I know a lot front-end capabilities cuz I know a lot front-end capabilities cuz I know a lot of you guys are excited about this. I of you guys are excited about this. I of you guys are excited about this. I did one of my recent favorite demos, did one of my recent favorite demos, did one of my recent favorite demos, which is to have things redesign the T3 which is to have things redesign the T3 which is to have things redesign the T3 code marketing site. This is the design code marketing site. This is the design code marketing site. This is the design I came to after working with the Claude I came to after working with the Claude I came to after working with the Claude design product, my own brain, and a lot design product, my own brain, and a lot design product, my own brain, and a lot of back and forth. I got it here. There of back and forth. I got it here. There of back and forth. I got it here. There are little things I would change.

  20. are little things I would change. are little things I would change. Obviously, I want to make the cursor Obviously, I want to make the cursor Obviously, I want to make the cursor icon a different color so it fits icon a different color so it fits icon a different color so it fits better, but you get the idea. It's not better, but you get the idea. It's not better, but you get the idea. It's not bad. I love my little carousel here with bad. I love my little carousel here with bad. I love my little carousel here with all the nice things people said about T3 all the nice things people said about T3 all the nice things people said about T3 code. code. code. So, let's look at what it did. I told it So, let's look at what it did. I told it So, let's look at what it did. I told it to do five different designs on to do five different designs on to do five different designs on different URLs to keep them varied and different URLs to keep them varied and different URLs to keep them varied and here's what it gave me so far. This is here's what it gave me so far. This is here's what it gave me so far. This is the first one. It's doing pills as a the first one. It's doing pills as a the first one. It's doing pills as a certain Open AI models really like to certain Open AI models really like to certain Open AI models really like to do, so that was interesting to see. I do, so that was interesting to see. I do, so that was interesting to see. I did this with Open Code as the harness did this with Open Code as the harness did this with Open Code as the harness if you were curious. So, that's number if you were curious. So, that's number if you were curious. So, that's number one. Not bad, but not something I would one. Not bad, but not something I would one. Not bad, but not something I would ship. Here's two. It looks It's nice ship. Here's two. It looks It's nice ship. Here's two. It looks It's nice except for the fact that this type of except for the fact that this type of except for the fact that this type of design has been copied by so many models design has been copied by so many models design has been copied by so many models for so many things that it's a not as for so many things that it's a not as for so many things that it's a not as cool anymore. I also can't help but cool anymore. I also can't help but cool anymore. I also can't help but notice that this top bar doesn't have notice that this top bar doesn't have notice that this top bar doesn't have enough content, so it's cycling, but enough content, so it's cycling, but enough content, so it's cycling, but it's half empty now as a result. Has a it's half empty now as a result. Has a it's half empty now as a result. Has a lot of nice little animations when you lot of nice little animations when you lot of nice little animations when you hover over things, which is cool. hover over things, which is cool. hover over things, which is cool. Still Still Still not perfect, though. But like not perfect, though. But like not perfect, though. But like considering this is an open weight considering this is an open weight considering this is an open weight model, unbelievably cool. Switch over model, unbelievably cool. Switch over model, unbelievably cool. Switch over here, and now we get the cringe terminal here, and now we get the cringe terminal here, and now we get the cringe terminal one.

  21. one. one. I was so shocked that this snuck in that I was so shocked that this snuck in that I was so shocked that this snuck in that I actually asked if it had pulled in my I actually asked if it had pulled in my I actually asked if it had pulled in my UI skill or something, so I thought I UI skill or something, so I thought I UI skill or something, so I thought I had removed it. And it hadn't. It's had removed it. And it hadn't. It's had removed it. And it hadn't. It's almost like the Claude design skill got almost like the Claude design skill got almost like the Claude design skill got baked in, though, which is baked in, though, which is baked in, though, which is For the fourth design, I channeled more For the fourth design, I channeled more For the fourth design, I channeled more of what I already had, but it with a of what I already had, but it with a of what I already had, but it with a surprisingly not too cringe glow in the surprisingly not too cringe glow in the surprisingly not too cringe glow in the corners, also fixing this icon to be the corners, also fixing this icon to be the corners, also fixing this icon to be the right color. right color. right color. Still don't necessarily love how it Still don't necessarily love how it Still don't necessarily love how it structured things here, too bubbly, but structured things here, too bubbly, but structured things here, too bubbly, but you could steer this somewhere good, for you could steer this somewhere good, for you could steer this somewhere good, for sure. sure. sure. I do think the gradients are pretty I do think the gradients are pretty I do think the gradients are pretty cool. cool. cool. And then we have the Bento box version, And then we have the Bento box version, And then we have the Bento box version, cuz I knew it would do one of these. I cuz I knew it would do one of these. I cuz I knew it would do one of these. I don't think it fits for the type of don't think it fits for the type of don't think it fits for the type of product that T3 Code is, but it's not a product that T3 Code is, but it's not a product that T3 Code is, but it's not a bad design at all. bad design at all. bad design at all. So, from just this one pass, I would say So, from just this one pass, I would say So, from just this one pass, I would say that for doing a marketing page, it's that for doing a marketing page, it's that for doing a marketing page, it's slightly better than what I get out of slightly better than what I get out of slightly better than what I get out of open AI models, but slightly behind what open AI models, but slightly behind what open AI models, but slightly behind what I would expect from Claude. But I would expect from Claude. But I would expect from Claude. But marketing pages are far from the best marketing pages are far from the best marketing pages are far from the best way to measure the UI capabilities of a way to measure the UI capabilities of a way to measure the UI capabilities of a model. So, I gave a slightly harder model. So, I gave a slightly harder model. So, I gave a slightly harder task. It was actually something I was task. It was actually something I was task. It was actually something I was already working on. I've been already working on. I've been already working on. I've been overhauling the sidebar in T3 Code. I overhauling the sidebar in T3 Code. I overhauling the sidebar in T3 Code. I want to find a better way to do it, want to find a better way to do it, want to find a better way to do it, something that's a little more uh something that's a little more uh something that's a little more uh flexible based on how it's being used.

  22. flexible based on how it's being used. flexible based on how it's being used. So, I already had this build where I So, I already had this build where I So, I already had this build where I redid the sidebar, and this build has a redid the sidebar, and this build has a redid the sidebar, and this build has a much better experience there, but I've much better experience there, but I've much better experience there, but I've noticed that the color isn't great. This noticed that the color isn't great. This noticed that the color isn't great. This is dark, and this isn't, and if is dark, and this isn't, and if is dark, and this isn't, and if anything, that should probably be anything, that should probably be anything, that should probably be flipped. And I also am hating the gray flipped. And I also am hating the gray flipped. And I also am hating the gray more and more since I had to remove the more and more since I had to remove the more and more since I had to remove the noise cuz of performance things that'll noise cuz of performance things that'll noise cuz of performance things that'll be in an upcoming video. Keep an eye out be in an upcoming video. Keep an eye out be in an upcoming video. Keep an eye out for that. Fable screwed up the for that. Fable screwed up the for that. Fable screwed up the performance of the app and I had to performance of the app and I had to performance of the app and I had to remove a bunch of [ __ ] to fix it. So, I remove a bunch of [ __ ] to fix it. So, I remove a bunch of [ __ ] to fix it. So, I asked it to do a darker redesign with asked it to do a darker redesign with asked it to do a darker redesign with true black instead of grays as often, true black instead of grays as often, true black instead of grays as often, and this is what it made. I will be and this is what it made. I will be and this is what it made. I will be frank. frank. frank. This is better than what we had. It made This is better than what we had. It made This is better than what we had. It made a better design than what we were a better design than what we were a better design than what we were already doing. already doing. already doing. With an open weight model that is a With an open weight model that is a With an open weight model that is a third the price of the frontier models third the price of the frontier models third the price of the frontier models from the other company that makes ones from the other company that makes ones from the other company that makes ones good at UI. I think we finally have a good at UI. I think we finally have a good at UI. I think we finally have a model that's good at solving real-world model that's good at solving real-world model that's good at solving real-world UI tasks without having to pay Anthropic UI tasks without having to pay Anthropic UI tasks without having to pay Anthropic massive amounts of money. massive amounts of money. massive amounts of money. Very exciting. Another cool UI example Very exciting. Another cool UI example Very exciting. Another cool UI example that Mac shared here is a recreation of that Mac shared here is a recreation of that Mac shared here is a recreation of macOS 27's liquid glass styles in a real macOS 27's liquid glass styles in a real macOS 27's liquid glass styles in a real web app that you can load that Kimmy web app that you can load that Kimmy web app that you can load that Kimmy built all of from scratch. And they're built all of from scratch. And they're built all of from scratch. And they're also hosting it on kimmy.page, their also hosting it on kimmy.page, their also hosting it on kimmy.page, their little like web hosting thing. And this little like web hosting thing. And this little like web hosting thing. And this is nuts. To have something that looks is nuts. To have something that looks is nuts. To have something that looks and works this well that like a open and works this well that like a open and works this well that like a open weight model generated for not too weight model generated for not too weight model generated for not too expensive is kind of crazy. There's a expensive is kind of crazy. There's a expensive is kind of crazy. There's a lot of little things that are broken in lot of little things that are broken in lot of little things that are broken in it, but like this is not bad at all for it, but like this is not bad at all for it, but like this is not bad at all for something that a model generated.

  23. something that a model generated. something that a model generated. Seriously, this is dope. What was much Seriously, this is dope. What was much Seriously, this is dope. What was much more exciting to me is how well it did more exciting to me is how well it did more exciting to me is how well it did this work. I did this work through this work. I did this work through this work. I did this work through OpenCode's bindings in T3 Code because, OpenCode's bindings in T3 Code because, OpenCode's bindings in T3 Code because, as I've mentioned many times, T3 Code is as I've mentioned many times, T3 Code is as I've mentioned many times, T3 Code is not a harness. You have to bring not a harness. You have to bring not a harness. You have to bring something like Claude Code, Codex, or something like Claude Code, Codex, or something like Claude Code, Codex, or OpenCode in order to use T3 Code. So, OpenCode in order to use T3 Code. So, OpenCode in order to use T3 Code. So, since Kimmy isn't available in Claude since Kimmy isn't available in Claude since Kimmy isn't available in Claude Code or Codex, wink, I'll show you Code or Codex, wink, I'll show you Code or Codex, wink, I'll show you something soon, something soon, something soon, I used OpenCode. And this was a fun bit I used OpenCode. And this was a fun bit I used OpenCode. And this was a fun bit of work. I told it to rethink the core of work. I told it to rethink the core of work. I told it to rethink the core UI for T3 Code. It's currently too gray. UI for T3 Code. It's currently too gray. UI for T3 Code. It's currently too gray. In dark mode, I want a cooler black and In dark mode, I want a cooler black and In dark mode, I want a cooler black and white layout with hard blacks inspired white layout with hard blacks inspired white layout with hard blacks inspired by ChatGPT. I gave it a screenshot of by ChatGPT. I gave it a screenshot of by ChatGPT. I gave it a screenshot of the ChatGPT UI so it could have that as the ChatGPT UI so it could have that as the ChatGPT UI so it could have that as a reference. And then an hour and 5 a reference. And then an hour and 5 a reference. And then an hour and 5 minutes later with one steer where I minutes later with one steer where I minutes later with one steer where I told it that it was opening things told it that it was opening things told it that it was opening things incorrectly cuz it kept trying to run incorrectly cuz it kept trying to run incorrectly cuz it kept trying to run the dev command to see, but when it did the dev command to see, but when it did the dev command to see, but when it did that it opened my browser and burned the that it opened my browser and burned the that it opened my browser and burned the code one-time token that we use in T3 code one-time token that we use in T3 code one-time token that we use in T3 code to verify who is connecting. So, it code to verify who is connecting. So, it code to verify who is connecting. So, it couldn't access it the way it wanted to. couldn't access it the way it wanted to. couldn't access it the way it wanted to. But, once I told it that and it stopped But, once I told it that and it stopped But, once I told it that and it stopped letting it open the browser, then it was letting it open the browser, then it was letting it open the browser, then it was able to access it, check its work, and able to access it, check its work, and able to access it, check its work, and keep refining until it got somewhere keep refining until it got somewhere keep refining until it got somewhere that it liked. It also apologized for that it liked. It also apologized for that it liked. It also apologized for opening new tabs constantly cuz it kept opening new tabs constantly cuz it kept opening new tabs constantly cuz it kept forgetting the no browser command. Sorry forgetting the no browser command. Sorry forgetting the no browser command. Sorry about that. Everything I started is about that. Everything I started is about that. Everything I started is killed in the temp homes have been killed in the temp homes have been killed in the temp homes have been removed. Also, a parallel sessions dev removed. Also, a parallel sessions dev removed. Also, a parallel sessions dev server, coding and PP ports got caught server, coding and PP ports got caught server, coding and PP ports got caught in a broad cleanup kill. I restarted in a broad cleanup kill. I restarted in a broad cleanup kill. I restarted those with their original flags. Nice.

  24. those with their original flags. Nice. those with their original flags. Nice. So, it made mistakes, it told me about So, it made mistakes, it told me about So, it made mistakes, it told me about them immediately, and then cleaned them them immediately, and then cleaned them them immediately, and then cleaned them up. Other models will make similar up. Other models will make similar up. Other models will make similar mistakes, but they don't necessarily mistakes, but they don't necessarily mistakes, but they don't necessarily tell you when they do. So, that was very tell you when they do. So, that was very tell you when they do. So, that was very nice to see. I asked it to do a more nice to see. I asked it to do a more nice to see. I asked it to do a more machine task, not like building a machine task, not like building a machine task, not like building a project, but fixing a config because I project, but fixing a config because I project, but fixing a config because I wanted to run this in isolated wanted to run this in isolated wanted to run this in isolated environment so it didn't affect my environment so it didn't affect my environment so it didn't affect my existing T3 code install, just so I existing T3 code install, just so I existing T3 code install, just so I could open it in the browser and play could open it in the browser and play could open it in the browser and play with it a bit as I did here. But, in with it a bit as I did here. But, in with it a bit as I did here. But, in order to do that, I needed content. So, order to do that, I needed content. So, order to do that, I needed content. So, I told it to figure out how to get I told it to figure out how to get I told it to figure out how to get things from my official T3 history, things from my official T3 history, things from my official T3 history, which was breaking this build cuz I have which was breaking this build cuz I have which was breaking this build cuz I have all my sidebar changes. And it managed all my sidebar changes. And it managed all my sidebar changes. And it managed to pull over the history, modify it, and to pull over the history, modify it, and to pull over the history, modify it, and get it in a state where I could actually get it in a state where I could actually get it in a state where I could actually work again. work again. work again. Thought that was really cool that it Thought that was really cool that it Thought that was really cool that it could do that type of thing. I then could do that type of thing. I then could do that type of thing. I then asked it in the same thread to file a asked it in the same thread to file a asked it in the same thread to file a PR. Follow Reborn Adventures for PR PR. Follow Reborn Adventures for PR PR. Follow Reborn Adventures for PR naming, minimal description, include naming, minimal description, include naming, minimal description, include before and after photos, use the file before and after photos, use the file before and after photos, use the file upload skill. If you don't have the upload skill. If you don't have the upload skill. If you don't have the skill, look at the global Codex and skill, look at the global Codex and skill, look at the global Codex and Cloud skills, you can find it there. I Cloud skills, you can find it there. I Cloud skills, you can find it there. I added this because I didn't have the added this because I didn't have the added this because I didn't have the skill in open code and wanted to go find skill in open code and wanted to go find skill in open code and wanted to go find it, and it did, and it got it working. it, and it did, and it got it working. it, and it did, and it got it working. It again struggled a bit with getting It again struggled a bit with getting It again struggled a bit with getting the browser going, but once it realized the browser going, but once it realized the browser going, but once it realized it could do the no browser command and it could do the no browser command and it could do the no browser command and also open in Puppeteer, it pulled it also open in Puppeteer, it pulled it also open in Puppeteer, it pulled it together, and it made this PR where it together, and it made this PR where it together, and it made this PR where it shows the before and the after. You can shows the before and the after. You can shows the before and the after. You can very clearly see how much better the very clearly see how much better the very clearly see how much better the after is here. I'm after is here. I'm after is here. I'm I'm impressed. It did a great job with I'm impressed. It did a great job with I'm impressed. It did a great job with this work. I didn't hold back in any this work. I didn't hold back in any this work. I didn't hold back in any way. I didn't give it an easier thing to way. I didn't give it an easier thing to way. I didn't give it an easier thing to see if it could do it. I worked with see if it could do it. I worked with see if it could do it. I worked with this the way I work with other models, this the way I work with other models, this the way I work with other models, and it did all of the work very well in and it did all of the work very well in and it did all of the work very well in one thread without having to do any one thread without having to do any one thread without having to do any custom anything. Since it did so well on

  25. custom anything. Since it did so well on custom anything. Since it did so well on those tasks, I decided to push its those tasks, I decided to push its those tasks, I decided to push its limits with these bigger overhauls. I limits with these bigger overhauls. I limits with these bigger overhauls. I did have the issue with the token window did have the issue with the token window did have the issue with the token window and compaction with the previous set as and compaction with the previous set as and compaction with the previous set as I mentioned. I'm working on fixing that I mentioned. I'm working on fixing that I mentioned. I'm working on fixing that now, but that did not stop me from now, but that did not stop me from now, but that did not stop me from having success with some other really having success with some other really having success with some other really big jobs. This was one of those jobs. I big jobs. This was one of those jobs. I big jobs. This was one of those jobs. I asked it to do a deep audit for security asked it to do a deep audit for security asked it to do a deep audit for security issues on my cloud product that I'm issues on my cloud product that I'm issues on my cloud product that I'm building called Lakebed. This is one of building called Lakebed. This is one of building called Lakebed. This is one of those tasks that I get a lot of refusals those tasks that I get a lot of refusals those tasks that I get a lot of refusals for when I use Fable and Soul for it. for when I use Fable and Soul for it. for when I use Fable and Soul for it. So, I was curious to see if it would So, I was curious to see if it would So, I was curious to see if it would refuse as well as what it would find cuz refuse as well as what it would find cuz refuse as well as what it would find cuz I have a concern with a really powerful I have a concern with a really powerful I have a concern with a really powerful open weight model like this coming out. open weight model like this coming out. open weight model like this coming out. And it did it. It spun up a bunch of And it did it. It spun up a bunch of And it did it. It spun up a bunch of agents to go do discovery on specific agents to go do discovery on specific agents to go do discovery on specific things. It then followed up with a things. It then followed up with a things. It then followed up with a verification pass for all of those verification pass for all of those verification pass for all of those things that it found doing 25 things that it found doing 25 things that it found doing 25 verification agents. And then it verification agents. And then it verification agents. And then it synthesized it all at the end, ran out synthesized it all at the end, ran out synthesized it all at the end, ran out of token space, I compacted, continued, of token space, I compacted, continued, of token space, I compacted, continued, and it succeeded. It gave me useful and it succeeded. It gave me useful and it succeeded. It gave me useful feedback on things I can do to secure feedback on things I can do to secure feedback on things I can do to secure further. further. further. This is good on my end cuz I can use This is good on my end cuz I can use This is good on my end cuz I can use this to secure my systems, but it's this to secure my systems, but it's this to secure my systems, but it's going to be really bad for us in the going to be really bad for us in the going to be really bad for us in the future when attackers start using it for future when attackers start using it for future when attackers start using it for similar things. All of that effort that similar things. All of that effort that similar things. All of that effort that Anthropic and OpenAI have been putting Anthropic and OpenAI have been putting Anthropic and OpenAI have been putting in to use their models for defense work in to use their models for defense work in to use their models for defense work and blocking them from doing offensive and blocking them from doing offensive and blocking them from doing offensive work is very helpful because that is no work is very helpful because that is no work is very helpful because that is no longer going to be enough to protect us longer going to be enough to protect us longer going to be enough to protect us now that we have an open weight model now that we have an open weight model now that we have an open weight model this capable. And if you're curious how this capable. And if you're curious how this capable. And if you're curious how Moonshot is thinking about security and Moonshot is thinking about security and Moonshot is thinking about security and safety, the word security appears inside safety, the word security appears inside safety, the word security appears inside of some of the demos on their page, but of some of the demos on their page, but of some of the demos on their page, but not in the actual contents of it. And not in the actual contents of it. And not in the actual contents of it. And the word safety does not appear on the the word safety does not appear on the the word safety does not appear on the page at all. As I mentioned before, they page at all. As I mentioned before, they page at all. As I mentioned before, they have not put out a system card yet. They

  26. have not put out a system card yet. They have not put out a system card yet. They plan to put out the technical reporting plan to put out the technical reporting plan to put out the technical reporting in the future, in the future, in the future, but we have no info about safety and but we have no info about safety and but we have no info about safety and security and what the impacts on those security and what the impacts on those security and what the impacts on those will be from this model yet. We don't will be from this model yet. We don't will be from this model yet. We don't even know how they thought about safety even know how they thought about safety even know how they thought about safety and security during training. So, that and security during training. So, that and security during training. So, that is a thing that is worth being concerned is a thing that is worth being concerned is a thing that is worth being concerned about. And if I talk about just how about. And if I talk about just how about. And if I talk about just how concerned I am, this video will be much concerned I am, this video will be much concerned I am, this video will be much longer. So, go check out my other videos longer. So, go check out my other videos longer. So, go check out my other videos about security and my concerns in that about security and my concerns in that about security and my concerns in that general space. Since the reasoning is general space. Since the reasoning is general space. Since the reasoning is visible in this model, I did take the visible in this model, I did take the visible in this model, I did take the opportunity to read said reasoning a opportunity to read said reasoning a opportunity to read said reasoning a little bit and uh little bit and uh little bit and uh I don't know if the other labs do this I don't know if the other labs do this I don't know if the other labs do this too cuz most of them don't share the too cuz most of them don't share the too cuz most of them don't share the reasoning, but it was interesting to see reasoning, but it was interesting to see reasoning, but it was interesting to see the dumb things the model would get the dumb things the model would get the dumb things the model would get stuck on. For example, in my global stuck on. For example, in my global stuck on. For example, in my global agents MD, I'm very strict with the agents MD, I'm very strict with the agents MD, I'm very strict with the models that I don't want them to write models that I don't want them to write models that I don't want them to write any all over the place. I want them to any all over the place. I want them to any all over the place. I want them to actually use TypeScript for its types. actually use TypeScript for its types. actually use TypeScript for its types. So, when it was working on a plan in a So, when it was working on a plan in a So, when it was working on a plan in a markdown, it read the detail that we markdown, it read the detail that we markdown, it read the detail that we need to ensure no any per user need to ensure no any per user need to ensure no any per user preference. The plan is markdown, not preference. The plan is markdown, not preference. The plan is markdown, not code, so okay. Like it took a second to code, so okay. Like it took a second to code, so okay. Like it took a second to think about that. Time was spent think about that. Time was spent think about that. Time was spent reasoning about the fact that it reasoning about the fact that it reasoning about the fact that it shouldn't use any, the TypeScript shouldn't use any, the TypeScript shouldn't use any, the TypeScript syntax, inside of a markdown file. When syntax, inside of a markdown file. When syntax, inside of a markdown file. When you combine that type of bad reasoning you combine that type of bad reasoning you combine that type of bad reasoning with the model's 20 TPS out, which is with the model's 20 TPS out, which is with the model's 20 TPS out, which is less than half of what we see from other less than half of what we see from other less than half of what we see from other labs, the result is that it feels slow.

  27. labs, the result is that it feels slow. labs, the result is that it feels slow. It's spending time doing things it It's spending time doing things it It's spending time doing things it shouldn't. Others might as well, but we shouldn't. Others might as well, but we shouldn't. Others might as well, but we don't see that data. At the very least, don't see that data. At the very least, don't see that data. At the very least, I would guess OpenAI doesn't cuz their I would guess OpenAI doesn't cuz their I would guess OpenAI doesn't cuz their models are so efficient with reasoning. models are so efficient with reasoning. models are so efficient with reasoning. But that is time I spent and money I But that is time I spent and money I But that is time I spent and money I spent on a thing that doesn't actually spent on a thing that doesn't actually spent on a thing that doesn't actually affect the quality of the outputs. affect the quality of the outputs. affect the quality of the outputs. Here's another weird experience I had. I Here's another weird experience I had. I Here's another weird experience I had. I mentioned this one earlier, where I had mentioned this one earlier, where I had mentioned this one earlier, where I had it doing a big workflow to write a plan it doing a big workflow to write a plan it doing a big workflow to write a plan up for me. And at the end, it didn't up for me. And at the end, it didn't up for me. And at the end, it didn't tell me what to do. It ended with no new tell me what to do. It ended with no new tell me what to do. It ended with no new report here, just an idle status ping. report here, just an idle status ping. report here, just an idle status ping. The plan in plan/overhaul.md remains The plan in plan/overhaul.md remains The plan in plan/overhaul.md remains current with all three audits current with all three audits current with all three audits incorporated. It was done. It had incorporated. It was done. It had incorporated. It was done. It had completed the work, but it was completed the work, but it was completed the work, but it was responding to the sub-agents and the responding to the sub-agents and the responding to the sub-agents and the ping, not to me, the user who was ping, not to me, the user who was ping, not to me, the user who was setting this off asking it to do the setting this off asking it to do the setting this off asking it to do the work. Just a weird thing that I had to work. Just a weird thing that I had to work. Just a weird thing that I had to notice and be like, oh, it's not going notice and be like, oh, it's not going notice and be like, oh, it's not going anymore. I need to do a follow-up. The anymore. I need to do a follow-up. The anymore. I need to do a follow-up. The reason I tend to use these models in reason I tend to use these models in reason I tend to use these models in Claude Code is because I really like Claude Code is because I really like Claude Code is because I really like workflows. I've talked a lot about this workflows. I've talked a lot about this workflows. I've talked a lot about this in previous videos. Go watch any of my in previous videos. Go watch any of my in previous videos. Go watch any of my videos about Claude Code versus Codex or videos about Claude Code versus Codex or videos about Claude Code versus Codex or why I use 5/6 in Claude code probably why I use 5/6 in Claude code probably why I use 5/6 in Claude code probably has more detail there as well as how I has more detail there as well as how I has more detail there as well as how I did this. That method worked perfectly did this. That method worked perfectly did this. That method worked perfectly for Kimmy by the way. It was very easy for Kimmy by the way. It was very easy for Kimmy by the way. It was very easy to just do slash model Kimmy K3 and it to just do slash model Kimmy K3 and it to just do slash model Kimmy K3 and it just worked. But what really blew me just worked. But what really blew me just worked. But what really blew me away was that Kimmy K3 can use away was that Kimmy K3 can use away was that Kimmy K3 can use sub-agents and workflows and all of sub-agents and workflows and all of sub-agents and workflows and all of these things correctly, but it did it these things correctly, but it did it these things correctly, but it did it differently. Normally when I tell a differently. Normally when I tell a differently. Normally when I tell a model to use a workflow, it will use one model to use a workflow, it will use one model to use a workflow, it will use one workflow to break down a big task. Kimmy workflow to break down a big task. Kimmy workflow to break down a big task. Kimmy made a bunch of workflows for various made a bunch of workflows for various made a bunch of workflows for various phases of the work it was going to do phases of the work it was going to do phases of the work it was going to do with the port of Paint to a modern with the port of Paint to a modern with the port of Paint to a modern stack. It broke up all these different stack. It broke up all these different stack. It broke up all these different phases, created workflows for them, and phases, created workflows for them, and phases, created workflows for them, and started running them. But the most

  28. started running them. But the most started running them. But the most interesting thing is that it allowed the interesting thing is that it allowed the interesting thing is that it allowed the sub-agents to complete tasks on the sub-agents to complete tasks on the sub-agents to complete tasks on the to-do list a level higher. So it had a to-do list a level higher. So it had a to-do list a level higher. So it had a to-do list, it turned the to-do list to-do list, it turned the to-do list to-do list, it turned the to-do list into workflows, started running some of into workflows, started running some of into workflows, started running some of them, and I noticed the to-dos just like them, and I noticed the to-dos just like them, and I noticed the to-dos just like randomly getting crossed off throughout. randomly getting crossed off throughout. randomly getting crossed off throughout. And that was because it gave those And that was because it gave those And that was because it gave those sub-agents the ability to check them sub-agents the ability to check them sub-agents the ability to check them off. Never seen that before. Even Fable off. Never seen that before. Even Fable off. Never seen that before. Even Fable doesn't do that inside of Claude code doesn't do that inside of Claude code doesn't do that inside of Claude code for my experience. So that was really for my experience. So that was really for my experience. So that was really cool and fascinating to me. Probably cool and fascinating to me. Probably cool and fascinating to me. Probably time to talk about how you can use the time to talk about how you can use the time to talk about how you can use the model today. Since the weights aren't model today. Since the weights aren't model today. Since the weights aren't available yet, the only way to use it is available yet, the only way to use it is available yet, the only way to use it is through the Kimmy platform. To be fair, through the Kimmy platform. To be fair, through the Kimmy platform. To be fair, this has been the case with Anthropic this has been the case with Anthropic this has been the case with Anthropic and OpenAI historically. They do usually and OpenAI historically. They do usually and OpenAI historically. They do usually set up their models on Bedrock with AWS set up their models on Bedrock with AWS set up their models on Bedrock with AWS as well as GCP and Azure, depending on as well as GCP and Azure, depending on as well as GCP and Azure, depending on which lab and which models and which which lab and which models and which which lab and which models and which provider. So there are some other provider. So there are some other provider. So there are some other options, but at the very least you options, but at the very least you options, but at the very least you always have a lot of options that are always have a lot of options that are always have a lot of options that are based out of the US. That is not the based out of the US. That is not the based out of the US. That is not the case for Moonshot. Moonshot is a Chinese case for Moonshot. Moonshot is a Chinese case for Moonshot. Moonshot is a Chinese company, which means either of these company, which means either of these company, which means either of these solutions are going to go through solutions are going to go through solutions are going to go through Chinese servers some amount. I would not Chinese servers some amount. I would not Chinese servers some amount. I would not use either of these methods for real use either of these methods for real use either of these methods for real important sensitive data at all yet. important sensitive data at all yet. important sensitive data at all yet. Wait till the weights are out, you'll be Wait till the weights are out, you'll be Wait till the weights are out, you'll be able to do that then. But for now, these able to do that then. But for now, these able to do that then. But for now, these are your two options. You have the are your two options. You have the are your two options. You have the subscription on kimmy.com or the API, subscription on kimmy.com or the API, subscription on kimmy.com or the API, which is platform.kimmy.ai.

  29. which is platform.kimmy.ai. which is platform.kimmy.ai. These are both hard to find, so I ended These are both hard to find, so I ended These are both hard to find, so I ended up accidentally spending $100 on the API up accidentally spending $100 on the API up accidentally spending $100 on the API because I forgot I had a $40 because I forgot I had a $40 because I forgot I had a $40 subscription on Kimmy. They have an subscription on Kimmy. They have an subscription on Kimmy. They have an interesting breakdown of subscriptions. interesting breakdown of subscriptions. interesting breakdown of subscriptions. They have a $20, a $40, a $100, and a They have a $20, a $40, a $100, and a They have a $20, a $40, a $100, and a $200. And I ended up hitting the $40 $200. And I ended up hitting the $40 $200. And I ended up hitting the $40 pretty fast with K3, at least the 5-hour pretty fast with K3, at least the 5-hour pretty fast with K3, at least the 5-hour window. I hit it in about 30 minutes of window. I hit it in about 30 minutes of window. I hit it in about 30 minutes of my early testing. So, I bumped up the my early testing. So, I bumped up the my early testing. So, I bumped up the $200 tier. And despite my borderline $200 tier. And despite my borderline $200 tier. And despite my borderline abusive usage, I have not even gotten to abusive usage, I have not even gotten to abusive usage, I have not even gotten to 50% of my 5-hour or 10% of my weekly 50% of my 5-hour or 10% of my weekly 50% of my 5-hour or 10% of my weekly since doing that upgrade. It seems to be since doing that upgrade. It seems to be since doing that upgrade. It seems to be quite generous, which makes sense cuz quite generous, which makes sense cuz quite generous, which makes sense cuz the model only costs $15 per mill token the model only costs $15 per mill token the model only costs $15 per mill token out. So, I did a bunch of calculations out. So, I did a bunch of calculations out. So, I did a bunch of calculations on how much my usage has cost. For the on how much my usage has cost. For the on how much my usage has cost. For the $40 sub that I burned it upgraded, the $40 sub that I burned it upgraded, the $40 sub that I burned it upgraded, the $10 or so of API usage I did before $10 or so of API usage I did before $10 or so of API usage I did before realizing this, and all of my usage in realizing this, and all of my usage in realizing this, and all of my usage in that sub so far, my estimates from using that sub so far, my estimates from using that sub so far, my estimates from using CC usage and having my whole fleet CC usage and having my whole fleet CC usage and having my whole fleet analyzed through Soul, is around $63 of analyzed through Soul, is around $63 of analyzed through Soul, is around $63 of usage. That is not bad at all usage. That is not bad at all usage. That is not bad at all considering how much work I've been considering how much work I've been considering how much work I've been getting done with this model. I'm getting done with this model. I'm getting done with this model. I'm impressed. For context, I'm doing like a impressed. For context, I'm doing like a impressed. For context, I'm doing like a $1,000 a day with OpenAI models right $1,000 a day with OpenAI models right $1,000 a day with OpenAI models right now just on weird side project I was now just on weird side project I was now just on weird side project I was burning and looping on, and I do at burning and looping on, and I do at burning and looping on, and I do at least 3 to $400 a day with Fable for my least 3 to $400 a day with Fable for my least 3 to $400 a day with Fable for my day-to-day work. I've done a similar day-to-day work. I've done a similar day-to-day work. I've done a similar amount of work with this model and it amount of work with this model and it amount of work with this model and it ended up being way cheaper. So, that's ended up being way cheaper. So, that's ended up being way cheaper. So, that's good to see. Hopefully, not similar to good to see. Hopefully, not similar to good to see. Hopefully, not similar to the amount of work I do with Soul, to be the amount of work I do with Soul, to be the amount of work I do with Soul, to be clear. I do way too much with that. But, clear. I do way too much with that. But, clear. I do way too much with that. But, the cost to actual work done compared to the cost to actual work done compared to the cost to actual work done compared to Fable, it's a really good ratio. And it Fable, it's a really good ratio. And it Fable, it's a really good ratio. And it does also instinctively so far seem

  30. does also instinctively so far seem does also instinctively so far seem slightly cheaper than 5/6 on lower slightly cheaper than 5/6 on lower slightly cheaper than 5/6 on lower reasoning levels for my day-to-day work. reasoning levels for my day-to-day work. reasoning levels for my day-to-day work. What this means for you is that the API What this means for you is that the API What this means for you is that the API is actually a pretty reasonable option is actually a pretty reasonable option is actually a pretty reasonable option because the costs aren't too bad. But, I because the costs aren't too bad. But, I because the costs aren't too bad. But, I would start with a subscription, at would start with a subscription, at would start with a subscription, at least like the $40 tier to play with it least like the $40 tier to play with it least like the $40 tier to play with it if you're curious enough to do said if you're curious enough to do said if you're curious enough to do said play. And if you do sign up for either play. And if you do sign up for either play. And if you do sign up for either of these, you can bring the API key from of these, you can bring the API key from of these, you can bring the API key from them into Open Code, therefore making it them into Open Code, therefore making it them into Open Code, therefore making it work in something like, I don't know, T3 work in something like, I don't know, T3 work in something like, I don't know, T3 Code. It's also worth noting that the Code. It's also worth noting that the Code. It's also worth noting that the subscriptions work inside of the CLI subscriptions work inside of the CLI subscriptions work inside of the CLI proxy that I use for routing my Claude proxy that I use for routing my Claude proxy that I use for routing my Claude code to both Claude code and Codex, and code to both Claude code and Codex, and code to both Claude code and Codex, and now also to Kimmy. That's some pretty now also to Kimmy. That's some pretty now also to Kimmy. That's some pretty nice, too. If you'd prefer to wait until nice, too. If you'd prefer to wait until nice, too. If you'd prefer to wait until the weights are released so you can use the weights are released so you can use the weights are released so you can use this on American servers, I totally this on American servers, I totally this on American servers, I totally understand. But if your reason for understand. But if your reason for understand. But if your reason for waiting is cost, don't. I ran a bunch of waiting is cost, don't. I ran a bunch of waiting is cost, don't. I ran a bunch of numbers in previous Kimi models, when numbers in previous Kimi models, when numbers in previous Kimi models, when you compare the price they released them you compare the price they released them you compare the price they released them at to the price that was available on at to the price that was available on at to the price that was available on other hosts and providers, they were other hosts and providers, they were other hosts and providers, they were cheaper, but it was like 11 to 15% at cheaper, but it was like 11 to 15% at cheaper, but it was like 11 to 15% at best. These models are just massive, so best. These models are just massive, so best. These models are just massive, so they're not cheap to host, and I would they're not cheap to host, and I would they're not cheap to host, and I would not expect this model to suddenly become not expect this model to suddenly become not expect this model to suddenly become $5 on other providers unless they $5 on other providers unless they $5 on other providers unless they quantize the hell out of it and make it quantize the hell out of it and make it quantize the hell out of it and make it way dumber. That doesn't mean the open way dumber. That doesn't mean the open way dumber. That doesn't mean the open weight nature of the model isn't great weight nature of the model isn't great weight nature of the model isn't great for the market as a whole. It's going to for the market as a whole. It's going to for the market as a whole. It's going to make pricing way more competitive, and I make pricing way more competitive, and I make pricing way more competitive, and I would be genuinely surprised if this would be genuinely surprised if this would be genuinely surprised if this doesn't force Anthropic to rethink some doesn't force Anthropic to rethink some doesn't force Anthropic to rethink some of the Opus release that they have of the Opus release that they have of the Opus release that they have coming in the not-too-distant future. If coming in the not-too-distant future. If coming in the not-too-distant future. If they still price it the same they did they still price it the same they did they still price it the same they did before and the numbers end up worse than before and the numbers end up worse than before and the numbers end up worse than Kimi's, Kimi's, Kimi's, Anthropic's going to be in a rough Anthropic's going to be in a rough Anthropic's going to be in a rough place. Think I've said everything I have place. Think I've said everything I have place. Think I've said everything I have to about this model. It's unbelievable to about this model. It's unbelievable to about this model. It's unbelievable at 3D. It's really damn good at visual

  31. at 3D. It's really damn good at visual at 3D. It's really damn good at visual stuff, especially coding and UI, and stuff, especially coding and UI, and stuff, especially coding and UI, and it's just overall a really good model. it's just overall a really good model. it's just overall a really good model. It no longer feels like it's good for It no longer feels like it's good for It no longer feels like it's good for open weight. It now feels frontier-class open weight. It now feels frontier-class open weight. It now feels frontier-class in many different ways. And while in many different ways. And while in many different ways. And while talking to it and working with it isn't talking to it and working with it isn't talking to it and working with it isn't quite as pleasant as the best models quite as pleasant as the best models quite as pleasant as the best models available, it's close enough that I am available, it's close enough that I am available, it's close enough that I am blown away, and I'm so excited for the blown away, and I'm so excited for the blown away, and I'm so excited for the future of models like this and where future of models like this and where future of models like this and where they're going to go. Let me know how they're going to go. Let me know how they're going to go. Let me know how y'all feel about the Kimi K3 launch, and y'all feel about the Kimi K3 launch, and y'all feel about the Kimi K3 launch, and until next time, until next time, until next time, peace, nerds.

Summary

The main theme discusses the significant advancements in open-weight AI models, specifically highlighting KimmyK3 as potentially surpassing closed models like Fable and GPT-5.6. The takeaway is that while benchmarks show impressive performance, the truly impactful aspect is the genuine capabilities and potential security implications of this highly capable, unrestricted open-weight model. The discussion will cover maximizing its usage, addressing confusion, and exploring its broader impact on development, the economy, and security.

View original episode ↗