I need to rant about local models
Read full transcript 23 segments
-
I'm sorry for this one in advance, but I I'm sorry for this one in advance, but I need to crash out a little bit. I'm not need to crash out a little bit. I'm not need to crash out a little bit. I'm not proud of the things I'm about to say, proud of the things I'm about to say, proud of the things I'm about to say, but if I don't, I'm afraid nobody will. but if I don't, I'm afraid nobody will. but if I don't, I'm afraid nobody will. And I can already tell that this comment And I can already tell that this comment And I can already tell that this comment section will be a disaster because the section will be a disaster because the section will be a disaster because the last time I talked about this in other last time I talked about this in other last time I talked about this in other places, I got like 50 death threats in a places, I got like 50 death threats in a places, I got like 50 death threats in a day. The people who I'm about to talk day. The people who I'm about to talk day. The people who I'm about to talk about are genuinely insane. But before about are genuinely insane. But before about are genuinely insane. But before we do that, I want to be clear about we do that, I want to be clear about we do that, I want to be clear about what we are talking about here and what what we are talking about here and what what we are talking about here and what the scope of this video is. Open weight the scope of this video is. Open weight the scope of this video is. Open weight models are awesome. I love them. They models are awesome. I love them. They models are awesome. I love them. They are so cool and they are essential to an are so cool and they are essential to an are so cool and they are essential to an ecosystem to flourish at all. Open ecosystem to flourish at all. Open ecosystem to flourish at all. Open weight and open source are meaningfully weight and open source are meaningfully weight and open source are meaningfully similar enough that people who think similar enough that people who think similar enough that people who think open weight isn't open source, not open weight isn't open source, not open weight isn't open source, not interested to me. I don't care. Open interested to me. I don't care. Open interested to me. I don't care. Open weight models are essential for us to weight models are essential for us to weight models are essential for us to keep evolving the AI ecosystem and keep evolving the AI ecosystem and keep evolving the AI ecosystem and landscape, especially now that the landscape, especially now that the landscape, especially now that the government is starting to take away government is starting to take away government is starting to take away certain models. We need good open weight certain models. We need good open weight certain models. We need good open weight models. So, what is this video about if models. So, what is this video about if models. So, what is this video about if I'm pro open weight? I think GLM-52 is I'm pro open weight? I think GLM-52 is I'm pro open weight? I think GLM-52 is awesome, by the way. Like unbelievably awesome, by the way. Like unbelievably awesome, by the way. Like unbelievably good for what it is. It's actually close good for what it is. It's actually close good for what it is. It's actually close to catching up to like 5-4 and Opus to catching up to like 5-4 and Opus to catching up to like 5-4 and Opus 4-6-4-7 in an open weight thing you can 4-6-4-7 in an open weight thing you can 4-6-4-7 in an open weight thing you can download yourself. That's crazy. But download yourself. That's crazy. But download yourself. That's crazy. But there's a catch. I said download there's a catch. I said download there's a catch. I said download yourself. I didn't say run yourself.
-
yourself. I didn't say run yourself. yourself. I didn't say run yourself. Models like GLM-52, despite being Models like GLM-52, despite being Models like GLM-52, despite being incredible, are nearly impossible to run incredible, are nearly impossible to run incredible, are nearly impossible to run the full proper versions of. 5-2 is the full proper versions of. 5-2 is the full proper versions of. 5-2 is conservatively 400 GB. conservatively 400 GB. conservatively 400 GB. So, if you don't have that much VRAM, So, if you don't have that much VRAM, So, if you don't have that much VRAM, you are not using the full model. This you are not using the full model. This you are not using the full model. This is a model that runs great on crazy data is a model that runs great on crazy data is a model that runs great on crazy data center setups and it's capable of center setups and it's capable of center setups and it's capable of unbelievable things, but it is not a unbelievable things, but it is not a unbelievable things, but it is not a local model. And if you're curious about local model. And if you're curious about local model. And if you're curious about the full precision GLM-52 BF16, 1.5 TB. the full precision GLM-52 BF16, 1.5 TB. the full precision GLM-52 BF16, 1.5 TB. That's not running on anything anyone That's not running on anything anyone That's not running on anything anyone watching this video owns in their house. watching this video owns in their house. watching this video owns in their house. And if you can't run this on something And if you can't run this on something And if you can't run this on something in your house, I'm very curious why in your house, I'm very curious why in your house, I'm very curious why you're living inside of Colossus 1 or 2. you're living inside of Colossus 1 or 2. you're living inside of Colossus 1 or 2. Even the quantized and pruned versions Even the quantized and pruned versions Even the quantized and pruned versions of models like this are still in the 200 of models like this are still in the 200 of models like this are still in the 200 GB range, which isn't running on almost GB range, which isn't running on almost GB range, which isn't running on almost any consumer hardware. And if it is, any consumer hardware. And if it is, any consumer hardware. And if it is, it's barely going to run. And this is it's barely going to run. And this is it's barely going to run. And this is just the start of why I think the local just the start of why I think the local just the start of why I think the local model thing is really overrated and model thing is really overrated and model thing is really overrated and dumb. Local models are not GLM 52. It's dumb. Local models are not GLM 52. It's dumb. Local models are not GLM 52. It's things like a quantized Gemma 4 that things like a quantized Gemma 4 that things like a quantized Gemma 4 that barely functions. And we're just at the barely functions. And we're just at the barely functions. And we're just at the tip of the iceberg here for all of the tip of the iceberg here for all of the tip of the iceberg here for all of the problems that exist in the local model problems that exist in the local model problems that exist in the local model world. If you're hearing people talking world. If you're hearing people talking world. If you're hearing people talking about how they can replace Codex and about how they can replace Codex and about how they can replace Codex and Claude code with things running on their Claude code with things running on their Claude code with things running on their own GPUs, you're being lied to, and I'm own GPUs, you're being lied to, and I'm own GPUs, you're being lied to, and I'm going to explain why I'll for a quick going to explain why I'll for a quick going to explain why I'll for a quick break for today's sponsor. If I told you break for today's sponsor. If I told you break for today's sponsor. If I told you that today's sponsor could 4x your that today's sponsor could 4x your that today's sponsor could 4x your potential customer base, you would potential customer base, you would potential customer base, you would probably call me insane and a liar. And probably call me insane and a liar. And probably call me insane and a liar. And that makes sense, because I am kind of that makes sense, because I am kind of that makes sense, because I am kind of insane and I am kind of lying. It's insane and I am kind of lying. It's insane and I am kind of lying. It's actually 5x, because General Translation actually 5x, because General Translation actually 5x, because General Translation it makes it trivial to launch your stuff
-
it makes it trivial to launch your stuff it makes it trivial to launch your stuff in every single language. If you're only in every single language. If you're only in every single language. If you're only shipping in English, you're leaving shipping in English, you're leaving shipping in English, you're leaving behind 81% of the world. Only 5% speaks behind 81% of the world. Only 5% speaks behind 81% of the world. Only 5% speaks English as their native tongue, and only English as their native tongue, and only English as their native tongue, and only 19% speaks English overall. And a lot of 19% speaks English overall. And a lot of 19% speaks English overall. And a lot of that 19% prefers other languages, too. that 19% prefers other languages, too. that 19% prefers other languages, too. So, supporting all of them is huge, but So, supporting all of them is huge, but So, supporting all of them is huge, but it's always been way too much work. Take it's always been way too much work. Take it's always been way too much work. Take it from me, I helped build the system it from me, I helped build the system it from me, I helped build the system for this at Twitch, and it was for this at Twitch, and it was for this at Twitch, and it was miserable. General Translation makes it miserable. General Translation makes it miserable. General Translation makes it too easy. It almost feels criminal. You too easy. It almost feels criminal. You too easy. It almost feels criminal. You can run their NPX GT at latest command, can run their NPX GT at latest command, can run their NPX GT at latest command, and it will get you going fast. Or if and it will get you going fast. Or if and it will get you going fast. Or if you're old school like me, you can go you're old school like me, you can go you're old school like me, you can go look at the code yourself. You install look at the code yourself. You install look at the code yourself. You install the package in the project for whatever the package in the project for whatever the package in the project for whatever framework or tool that you're using, you framework or tool that you're using, you framework or tool that you're using, you import their config, you choose the import their config, you choose the import their config, you choose the languages you want to support, and then languages you want to support, and then languages you want to support, and then you wrap your app in their provider. And you wrap your app in their provider. And you wrap your app in their provider. And from this point on, it is trivial to add from this point on, it is trivial to add from this point on, it is trivial to add translation because you use their helper translation because you use their helper translation because you use their helper component, T. You don't have to wrap component, T. You don't have to wrap component, T. You don't have to wrap every string, you just wrap the sections every string, you just wrap the sections every string, you just wrap the sections where there are strings that should be where there are strings that should be where there are strings that should be translated. And now their systems can translated. And now their systems can translated. And now their systems can yank out all of the copy that matters, yank out all of the copy that matters, yank out all of the copy that matters, generate translations for it, and then generate translations for it, and then generate translations for it, and then give them to your app when you're ready give them to your app when you're ready give them to your app when you're ready to ship. It all just works as part of to ship. It all just works as part of to ship. It all just works as part of your CI, and it couldn't be easier to your CI, and it couldn't be easier to your CI, and it couldn't be easier to do. They provide all the annoying do. They provide all the annoying do. They provide all the annoying components you might need for something components you might need for something components you might need for something like this, like a local selector, or like this, like a local selector, or like this, like a local selector, or more importantly, the ability to to more importantly, the ability to to more importantly, the ability to to variable content where things shouldn't variable content where things shouldn't variable content where things shouldn't be formatted or a number where the be formatted or a number where the be formatted or a number where the formatting varies a lot depending on formatting varies a lot depending on formatting varies a lot depending on where you are or even things like date where you are or even things like date where you are or even things like date time and currency which are so annoying time and currency which are so annoying time and currency which are so annoying to get right by hand. There's a reason to get right by hand. There's a reason to get right by hand. There's a reason that companies like Cursor, Ramp, that companies like Cursor, Ramp, that companies like Cursor, Ramp, ClickHouse, Partiful and more are ClickHouse, Partiful and more are ClickHouse, Partiful and more are leaning on general translation to manage leaning on general translation to manage leaning on general translation to manage their translations. It's because they their translations. It's because they their translations. It's because they want users that aren't just English
-
want users that aren't just English want users that aren't just English speakers. Serve more users in their speakers. Serve more users in their speakers. Serve more users in their language at soydev.link/gt. language at soydev.link/gt. language at soydev.link/gt. So let's talk about why I think the So let's talk about why I think the So let's talk about why I think the local model hype is dumb. There's a lot local model hype is dumb. There's a lot local model hype is dumb. There's a lot of layers to this but I want to close of layers to this but I want to close of layers to this but I want to close out the one I was hinting at at the out the one I was hinting at at the out the one I was hinting at at the start which is the gap between runnable start which is the gap between runnable start which is the gap between runnable and good. The models that you can run on and good. The models that you can run on and good. The models that you can run on your laptop are incredibly impressive your laptop are incredibly impressive your laptop are incredibly impressive that they could exist at all but they're that they could exist at all but they're that they could exist at all but they're not doing real work. There are some not doing real work. There are some not doing real work. There are some exceptions here like DS4 by antirez. exceptions here like DS4 by antirez. exceptions here like DS4 by antirez. Antirez is the creator of Redis and he Antirez is the creator of Redis and he Antirez is the creator of Redis and he has been trying to build a from scratch has been trying to build a from scratch has been trying to build a from scratch C based runtime in order to be able to C based runtime in order to be able to C based runtime in order to be able to use the Deep Seek V4 models on his use the Deep Seek V4 models on his use the Deep Seek V4 models on his MacBook. It is a native inference engine MacBook. It is a native inference engine MacBook. It is a native inference engine optimized first for Deep Seek V4 flash optimized first for Deep Seek V4 flash optimized first for Deep Seek V4 flash with some support for V4 Pro on very with some support for V4 Pro on very with some support for V4 Pro on very high memory machines. In order to even high memory machines. In order to even high memory machines. In order to even use flash which remember is the small use flash which remember is the small use flash which remember is the small version of Deep Seek V4, you need at version of Deep Seek V4, you need at version of Deep Seek V4, you need at least 96 gigs of RAM on a Mac or an least 96 gigs of RAM on a Mac or an least 96 gigs of RAM on a Mac or an Nvidia CUDA or Strix Halo machine. But Nvidia CUDA or Strix Halo machine. But Nvidia CUDA or Strix Halo machine. But Theo, I have 128 gigs of RAM in my Theo, I have 128 gigs of RAM in my Theo, I have 128 gigs of RAM in my gaming PC with a 5080 in it. Do you know gaming PC with a 5080 in it. Do you know gaming PC with a 5080 in it. Do you know how much RAM the 5080 has? Cuz that's how much RAM the 5080 has? Cuz that's how much RAM the 5080 has? Cuz that's what actually matters. 16 gigs. You're what actually matters. 16 gigs. You're what actually matters. 16 gigs. You're not fitting [ __ ] in this.
-
not fitting [ __ ] in this. not fitting [ __ ] in this. Oh, well, I can upgrade to the 5090. Do Oh, well, I can upgrade to the 5090. Do Oh, well, I can upgrade to the 5090. Do you know how much VRAM's in that? 32 you know how much VRAM's in that? 32 you know how much VRAM's in that? 32 gigs. gigs. gigs. You still can't fit [ __ ] on it. It does You still can't fit [ __ ] on it. It does You still can't fit [ __ ] on it. It does not matter how much RAM you put into not matter how much RAM you put into not matter how much RAM you put into your fancy gaming PC. It is not your fancy gaming PC. It is not your fancy gaming PC. It is not benefiting inference beyond when the benefiting inference beyond when the benefiting inference beyond when the VRAM is saturated. Well, why is your VRAM is saturated. Well, why is your VRAM is saturated. Well, why is your MacBook with 128 gigs of RAM better? MacBook with 128 gigs of RAM better? MacBook with 128 gigs of RAM better? Because it's also VRAM. It's unified Because it's also VRAM. It's unified Because it's also VRAM. It's unified memory. There's basically three real memory. There's basically three real memory. There's basically three real options you can consider as a consumer options you can consider as a consumer options you can consider as a consumer for getting lots of unified RAM which is for getting lots of unified RAM which is for getting lots of unified RAM which is useful for this type of thing. Option useful for this type of thing. Option useful for this type of thing. Option one, as I said here, is a MacBook with one, as I said here, is a MacBook with one, as I said here, is a MacBook with enough RAM. Would really be a shame if enough RAM. Would really be a shame if enough RAM. Would really be a shame if the prices happened to have just gone up the prices happened to have just gone up the prices happened to have just gone up 30 to 50% and more than 2x for the more 30 to 50% and more than 2x for the more 30 to 50% and more than 2x for the more RAM. The 128 gig option is now 3,000 RAM. The 128 gig option is now 3,000 RAM. The 128 gig option is now 3,000 additional dollars on most Macs. So, additional dollars on most Macs. So, additional dollars on most Macs. So, that's ruled out. Maybe I'll grab a DGX that's ruled out. Maybe I'll grab a DGX that's ruled out. Maybe I'll grab a DGX Spark. If you want to run these things Spark. If you want to run these things Spark. If you want to run these things at like 12 TPS at best, sure. I have a at like 12 TPS at best, sure. I have a at like 12 TPS at best, sure. I have a DGX Spark. It kind of just sits there. I DGX Spark. It kind of just sits there. I DGX Spark. It kind of just sits there. I should probably sell it. It's not should probably sell it. It's not should probably sell it. It's not useful. It's cool if you want to debug useful. It's cool if you want to debug useful. It's cool if you want to debug CUDA [ __ ] with a lot of VRAM in quotes, CUDA [ __ ] with a lot of VRAM in quotes, CUDA [ __ ] with a lot of VRAM in quotes, and you don't care about the and you don't care about the and you don't care about the performance. Other than that, not performance. Other than that, not performance. Other than that, not useful. So, what we're left with is the useful. So, what we're left with is the useful. So, what we're left with is the Strix Halo, which you can get inside of Strix Halo, which you can get inside of Strix Halo, which you can get inside of a handful of laptops, as well as within a handful of laptops, as well as within a handful of laptops, as well as within the Framework desktop, which is probably the Framework desktop, which is probably the Framework desktop, which is probably the best option you have if you don't the best option you have if you don't the best option you have if you don't want to go Mac. Those are your actual want to go Mac. Those are your actual want to go Mac. Those are your actual only consumer options to get enough RAM only consumer options to get enough RAM only consumer options to get enough RAM to even run these things. And since all to even run these things. And since all to even run these things. And since all of those aren't real GPUs, as in not of those aren't real GPUs, as in not of those aren't real GPUs, as in not real high-end Nvidia hardware, they're real high-end Nvidia hardware, they're real high-end Nvidia hardware, they're not going to run great. There's a very not going to run great. There's a very not going to run great. There's a very interesting problem that I see a lot on interesting problem that I see a lot on interesting problem that I see a lot on my Mac here, because again, I have these
-
my Mac here, because again, I have these my Mac here, because again, I have these things working on my MacBook with 128 things working on my MacBook with 128 things working on my MacBook with 128 gigs of RAM, because it has enough RAM. gigs of RAM, because it has enough RAM. gigs of RAM, because it has enough RAM. If you have a model that fits on this If you have a model that fits on this If you have a model that fits on this and fits on my 5090, it'll run three and fits on my 5090, it'll run three and fits on my 5090, it'll run three times plus faster on the 5090. If I take times plus faster on the 5090. If I take times plus faster on the 5090. If I take like Gemma 4, for example, it'll run like Gemma 4, for example, it'll run like Gemma 4, for example, it'll run decently on here. It'll run way better decently on here. It'll run way better decently on here. It'll run way better on my 5090, because it fits in the VRAM on my 5090, because it fits in the VRAM on my 5090, because it fits in the VRAM on both. But, if you switch to a big on both. But, if you switch to a big on both. But, if you switch to a big model like Deep See V4 Flash, not even model like Deep See V4 Flash, not even model like Deep See V4 Flash, not even Pro, it'll run okay on my MacBook, and Pro, it'll run okay on my MacBook, and Pro, it'll run okay on my MacBook, and it will run like absolute [ __ ] on my it will run like absolute [ __ ] on my it will run like absolute [ __ ] on my desktop, because it doesn't have enough desktop, because it doesn't have enough desktop, because it doesn't have enough VRAM, because it's bottlenecked on the VRAM, because it's bottlenecked on the VRAM, because it's bottlenecked on the GPU. So, your options are get enough RAM GPU. So, your options are get enough RAM GPU. So, your options are get enough RAM on one of these unified boxes and have on one of these unified boxes and have on one of these unified boxes and have way less compute, or you can get more way less compute, or you can get more way less compute, or you can get more than enough compute with a high-end chip than enough compute with a high-end chip than enough compute with a high-end chip like a 5090, which is as good as an RTX like a 5090, which is as good as an RTX like a 5090, which is as good as an RTX 6000. It is comparable to an A100, which 6000. It is comparable to an A100, which 6000. It is comparable to an A100, which is a much more expensive enterprise is a much more expensive enterprise is a much more expensive enterprise chip, except for the VRAM limitations. chip, except for the VRAM limitations. chip, except for the VRAM limitations. And Nvidia does that intentionally. The And Nvidia does that intentionally. The And Nvidia does that intentionally. The 5090 is just as powerful as an A100 if 5090 is just as powerful as an A100 if 5090 is just as powerful as an A100 if not better in everything but the VRAM. not better in everything but the VRAM. not better in everything but the VRAM. So, your options are high performance So, your options are high performance So, your options are high performance [ __ ] models or [ __ ] performance slightly [ __ ] models or [ __ ] performance slightly [ __ ] models or [ __ ] performance slightly better models. But as soon as you're better models. But as soon as you're better models. But as soon as you're talking about legitimately good models talking about legitimately good models talking about legitimately good models for our real dev work, we are far for our real dev work, we are far for our real dev work, we are far outside of what you can do on side of outside of what you can do on side of outside of what you can do on side of consumer hardware in any format. That's consumer hardware in any format. That's consumer hardware in any format. That's when you go from this 96 gig limit up to when you go from this 96 gig limit up to when you go from this 96 gig limit up to like 200 plus. And just for reference, like 200 plus. And just for reference, like 200 plus. And just for reference, the 5090 currently goes resale for the 5090 currently goes resale for the 5090 currently goes resale for around $4300 around $4300 around $4300 for just the GPU. That price is a little for just the GPU. That price is a little for just the GPU. That price is a little absurd. Okay, it looks like they are absurd. Okay, it looks like they are absurd. Okay, it looks like they are actually in the four grand range. That's
-
actually in the four grand range. That's actually in the four grand range. That's insane. insane. insane. That's absurd. That's absurd. That's absurd. I never paid near that for mine. I never paid near that for mine. I never paid near that for mine. But if you do need more VRAM, you have But if you do need more VRAM, you have But if you do need more VRAM, you have an option, the RTX an option, the RTX an option, the RTX 6000 Pro. The Pro 6000 Blackwell has 96 6000 Pro. The Pro 6000 Blackwell has 96 6000 Pro. The Pro 6000 Blackwell has 96 gigs of VRAM. So, you can now fit huge gigs of VRAM. So, you can now fit huge gigs of VRAM. So, you can now fit huge models like Deep Seek V4 Flash. And models like Deep Seek V4 Flash. And models like Deep Seek V4 Flash. And ready for the price on these? ready for the price on these? ready for the price on these? $13,000. $13,000. $13,000. And it's the same performance. It's just And it's the same performance. It's just And it's the same performance. It's just more VRAM. The chip on this is nearly more VRAM. The chip on this is nearly more VRAM. The chip on this is nearly identical to the 5090. And it's 3x more identical to the 5090. And it's 3x more identical to the 5090. And it's 3x more expensive than the 5090, which is expensive than the 5090, which is expensive than the 5090, which is already too expensive. Not good. Don't already too expensive. Not good. Don't already too expensive. Not good. Don't worry though, you can just buy a tiny worry though, you can just buy a tiny worry though, you can just buy a tiny box. And for only $75,000, box. And for only $75,000, box. And for only $75,000, you can have four RTX Pro 6000s. you can have four RTX Pro 6000s. you can have four RTX Pro 6000s. This could almost run GLM 52. It'll have This could almost run GLM 52. It'll have This could almost run GLM 52. It'll have to do some crazy things because the VRAM to do some crazy things because the VRAM to do some crazy things because the VRAM doesn't get pulled as you might hope doesn't get pulled as you might hope doesn't get pulled as you might hope when you have a lot of GPUs. You need to when you have a lot of GPUs. You need to when you have a lot of GPUs. You need to find some way to replicate some things find some way to replicate some things find some way to replicate some things and split up others. Yeah, good luck. and split up others. Yeah, good luck. and split up others. Yeah, good luck. Have fun. Have fun. Have fun. Don't worry though, the red is way Don't worry though, the red is way Don't worry though, the red is way cheaper. It's only $12,000.
-
cheaper. It's only $12,000. cheaper. It's only $12,000. And this one works by using a bunch of And this one works by using a bunch of And this one works by using a bunch of AMD GPUs instead of the Nvidia ones. And AMD GPUs instead of the Nvidia ones. And AMD GPUs instead of the Nvidia ones. And it's got a whopping 128 gigs of VRAM. it's got a whopping 128 gigs of VRAM. it's got a whopping 128 gigs of VRAM. So, again, you can use Flash. And almost So, again, you can use Flash. And almost So, again, you can use Flash. And almost nothing else. So, hopefully I have nothing else. So, hopefully I have nothing else. So, hopefully I have helped establish both that the models helped establish both that the models helped establish both that the models you're seeing these crazy open weight you're seeing these crazy open weight you're seeing these crazy open weight drops on, things like GLM 5.2, drops on, things like GLM 5.2, drops on, things like GLM 5.2, as incredible as they are, they're not as incredible as they are, they're not as incredible as they are, they're not running on consumer hardware. If you running on consumer hardware. If you running on consumer hardware. If you want to spend $75,000 on a tiny box, you want to spend $75,000 on a tiny box, you want to spend $75,000 on a tiny box, you can run one instance of it at a can run one instance of it at a can run one instance of it at a decent-ish speed. Well, when I hear decent-ish speed. Well, when I hear decent-ish speed. Well, when I hear local model, I think about things I can local model, I think about things I can local model, I think about things I can run on consumer hardware of some form. run on consumer hardware of some form. run on consumer hardware of some form. Even my $10,000 MacBook can barely run a Even my $10,000 MacBook can barely run a Even my $10,000 MacBook can barely run a lot of those, and this is better lot of those, and this is better lot of those, and this is better equipped than most machines. So, the equipped than most machines. So, the equipped than most machines. So, the first problem I want to make sure we first problem I want to make sure we first problem I want to make sure we fully understand this. Just because a fully understand this. Just because a fully understand this. Just because a model is open weight does not mean that model is open weight does not mean that model is open weight does not mean that you can run it at home. And the models you can run it at home. And the models you can run it at home. And the models that are open weighted unbelievable are that are open weighted unbelievable are that are open weighted unbelievable are very different from the models that you very different from the models that you very different from the models that you can run at home. The other thing, which can run at home. The other thing, which can run at home. The other thing, which I just established, I should have put on I just established, I should have put on I just established, I should have put on the list before, hardware costs are the list before, hardware costs are the list before, hardware costs are insane. If you're buying these types of insane. If you're buying these types of insane. If you're buying these types of GPUs, I think it's a better purchase as GPUs, I think it's a better purchase as GPUs, I think it's a better purchase as a thing you plan to resell than a thing you plan to resell than a thing you plan to resell than purchasing it to actually use. Oh god, purchasing it to actually use. Oh god, purchasing it to actually use. Oh god, do we have cope in the comments?
-
do we have cope in the comments? do we have cope in the comments? What the [ __ ] is Ornith? What the [ __ ] is Ornith? What the [ __ ] is Ornith? Everyday I get sent some new family of Everyday I get sent some new family of Everyday I get sent some new family of locally runnable-ish models that I've locally runnable-ish models that I've locally runnable-ish models that I've never heard of that I don't know anyone never heard of that I don't know anyone never heard of that I don't know anyone using. I know a lot more people plugging using. I know a lot more people plugging using. I know a lot more people plugging these than using them, and apparently these than using them, and apparently these than using them, and apparently the new one is Ornith, which apparently the new one is Ornith, which apparently the new one is Ornith, which apparently can get comparable numbers to Opus 47 can get comparable numbers to Opus 47 can get comparable numbers to Opus 47 and 48 in STV bench verified, which is a and 48 in STV bench verified, which is a and 48 in STV bench verified, which is a saturated, abused, entirely like saturated, abused, entirely like saturated, abused, entirely like compromised benchmark. Or Claw Eval, I compromised benchmark. Or Claw Eval, I compromised benchmark. Or Claw Eval, I don't even want to know what that is. don't even want to know what that is. don't even want to know what that is. When we look at some of these like NL2 When we look at some of these like NL2 When we look at some of these like NL2 repo, we see that it's like half-ish the repo, we see that it's like half-ish the repo, we see that it's like half-ish the score. Or if we look at something like score. Or if we look at something like score. Or if we look at something like Deep SW Deep SW Deep SW which is one of the few actually good which is one of the few actually good which is one of the few actually good benchmarks for these things, and we look benchmarks for these things, and we look benchmarks for these things, and we look at the open models they have here like at the open models they have here like at the open models they have here like Kimmy, Kimmy, Kimmy, it scored a 30% where 54 scored a 52. it scored a 30% where 54 scored a 52. it scored a 30% where 54 scored a 52. Opus on low scored a 41 and was cheaper Opus on low scored a 41 and was cheaper Opus on low scored a 41 and was cheaper to run because these open weight models to run because these open weight models to run because these open weight models just burn [ __ ] tokens. So, that's the just burn [ __ ] tokens. So, that's the just burn [ __ ] tokens. So, that's the other problem is even if you have them other problem is even if you have them other problem is even if you have them running fast, they're going to be slow running fast, they're going to be slow running fast, they're going to be slow as [ __ ] because of the inefficiencies as [ __ ] because of the inefficiencies as [ __ ] because of the inefficiencies that exist in all of those models. Cool.
-
that exist in all of those models. Cool. that exist in all of those models. Cool. Kimmy K 27's like on par with Sonnet 46. Kimmy K 27's like on par with Sonnet 46. Kimmy K 27's like on par with Sonnet 46. Awesome. That's so cool. But, you know Awesome. That's so cool. But, you know Awesome. That's so cool. But, you know what? I'll entertain it. what? I'll entertain it. what? I'll entertain it. Let's enter a hypothetical world where Let's enter a hypothetical world where Let's enter a hypothetical world where the local models do catch up to be the local models do catch up to be the local models do catch up to be roughly the same performance as what we roughly the same performance as what we roughly the same performance as what we get from models like Opus and GPT 5.5 get from models like Opus and GPT 5.5 get from models like Opus and GPT 5.5 today. Hypothetical, but let's just today. Hypothetical, but let's just today. Hypothetical, but let's just assume it does happen. That we get to assume it does happen. That we get to assume it does happen. That we get to the point where a 30 bill per am model the point where a 30 bill per am model the point where a 30 bill per am model that you can run on most higher-end that you can run on most higher-end that you can run on most higher-end MacBooks can be equivalent MacBooks can be equivalent MacBooks can be equivalent capabilities-wise to those big, beefy capabilities-wise to those big, beefy capabilities-wise to those big, beefy Frontier Labs. Let's just assume it can. Frontier Labs. Let's just assume it can. Frontier Labs. Let's just assume it can. I don't think it can, but let's assume I don't think it can, but let's assume I don't think it can, but let's assume it can. That's when you run into the it can. That's when you run into the it can. That's when you run into the next problem, parallelism. next problem, parallelism. next problem, parallelism. When I'm working, I'm not going from When I'm working, I'm not going from When I'm working, I'm not going from zero agent to one agent, I am randomly zero agent to one agent, I am randomly zero agent to one agent, I am randomly fluctuating between one and 40. Each of fluctuating between one and 40. Each of fluctuating between one and 40. Each of those are doing their own inference. those are doing their own inference. those are doing their own inference. Maybe I have two threads running in T3 Maybe I have two threads running in T3 Maybe I have two threads running in T3 code or Codex. Maybe I have an agent code or Codex. Maybe I have an agent code or Codex. Maybe I have an agent break up a bunch of sub agents to go break up a bunch of sub agents to go break up a bunch of sub agents to go explore different things in the project.
-
explore different things in the project. explore different things in the project. Maybe I just want to use Claude and Maybe I just want to use Claude and Maybe I just want to use Claude and Codex at the same time. There's a lot of Codex at the same time. There's a lot of Codex at the same time. There's a lot of reasons I'm running more than one agent reasons I'm running more than one agent reasons I'm running more than one agent at a time, and if you set up your system at a time, and if you set up your system at a time, and if you set up your system in such a crazy, impressive, capable way in such a crazy, impressive, capable way in such a crazy, impressive, capable way where it can actually run GLM 5.2, can where it can actually run GLM 5.2, can where it can actually run GLM 5.2, can it run it twice? Can it run it 10 times? it run it twice? Can it run it 10 times? it run it twice? Can it run it 10 times? Can it run it as many times as I can run Can it run it as many times as I can run Can it run it as many times as I can run sub agents in my current agentic sub agents in my current agentic sub agents in my current agentic workflows? I'll give you a hint, the workflows? I'll give you a hint, the workflows? I'll give you a hint, the answer is [ __ ] no. It absolutely can't. answer is [ __ ] no. It absolutely can't. answer is [ __ ] no. It absolutely can't. This is separate from the other problem, This is separate from the other problem, This is separate from the other problem, which is that you're just wasting the which is that you're just wasting the which is that you're just wasting the money when these sit idle. If you buy money when these sit idle. If you buy money when these sit idle. If you buy enough GPUs to run five agents in enough GPUs to run five agents in enough GPUs to run five agents in parallel, any moment you're not running parallel, any moment you're not running parallel, any moment you're not running five agents in parallel, you're just not five agents in parallel, you're just not five agents in parallel, you're just not benefiting. benefiting. benefiting. Those boxes sitting idle are just money Those boxes sitting idle are just money Those boxes sitting idle are just money wasted. But, if you do need to scale up wasted. But, if you do need to scale up wasted. But, if you do need to scale up to six, and you only bought enough to six, and you only bought enough to six, and you only bought enough hardware for five, you're blocked now. hardware for five, you're blocked now. hardware for five, you're blocked now. You got to wait until one of those You got to wait until one of those You got to wait until one of those workloads is done to have that lane workloads is done to have that lane workloads is done to have that lane available for the next. And if you want available for the next. And if you want available for the next. And if you want to run two different things at the same to run two different things at the same to run two different things at the same time, like you want to run a bigger time, like you want to run a bigger time, like you want to run a bigger model for some tasks and a smaller one model for some tasks and a smaller one model for some tasks and a smaller one for others, maybe you have a big model for others, maybe you have a big model for others, maybe you have a big model orchestrating and then smaller ones orchestrating and then smaller ones orchestrating and then smaller ones reading through the code and doing reading through the code and doing reading through the code and doing subtasks, you need 10 times more VRAM to subtasks, you need 10 times more VRAM to subtasks, you need 10 times more VRAM to even possibly enable that workflow and a even possibly enable that workflow and a even possibly enable that workflow and a lot more compute in order to power it.
-
lot more compute in order to power it. lot more compute in order to power it. The number of models I am running at any The number of models I am running at any The number of models I am running at any given time given time given time varies a lot. varies a lot. varies a lot. And I'm not able to do that on local And I'm not able to do that on local And I'm not able to do that on local hardware at all. And I see a lot of cope hardware at all. And I see a lot of cope hardware at all. And I see a lot of cope and chat about caching and VLLM and and chat about caching and VLLM and and chat about caching and VLLM and batching and all of these types of batching and all of these types of batching and all of these types of things. No, none of that actually solves things. No, none of that actually solves things. No, none of that actually solves this problem meaningfully. And if you this problem meaningfully. And if you this problem meaningfully. And if you think it does, you don't understand the think it does, you don't understand the think it does, you don't understand the problem. And again, we are presuming the problem. And again, we are presuming the problem. And again, we are presuming the models catch up to current state of the models catch up to current state of the models catch up to current state of the art when we talk about all of this art when we talk about all of this art when we talk about all of this because state of the art isn't just it because state of the art isn't just it because state of the art isn't just it can solve this benchmark well, it's it can solve this benchmark well, it's it can solve this benchmark well, it's it can queue up a crazy workflow that has can queue up a crazy workflow that has can queue up a crazy workflow that has lots of different stages and manage it lots of different stages and manage it lots of different stages and manage it as other agents complete those stages as other agents complete those stages as other agents complete those stages and do the whole end-to-end loop. Part and do the whole end-to-end loop. Part and do the whole end-to-end loop. Part of that end-to-end loop is often of that end-to-end loop is often of that end-to-end loop is often computer use. And GLM-52 doesn't even computer use. And GLM-52 doesn't even computer use. And GLM-52 doesn't even have vision. It's an unbelievable coding have vision. It's an unbelievable coding have vision. It's an unbelievable coding model, but it can't look at things. I model, but it can't look at things. I model, but it can't look at things. I can't give it a screenshot. And that's can't give it a screenshot. And that's can't give it a screenshot. And that's state of the art for open weight. state of the art for open weight. state of the art for open weight. Insane. It's still really good. GLM-52 Insane. It's still really good. GLM-52 Insane. It's still really good. GLM-52 is awesome. I'm so thankful it exists. is awesome. I'm so thankful it exists. is awesome. I'm so thankful it exists. But you're not running that on local But you're not running that on local But you're not running that on local hardware, you're not running that with hardware, you're not running that with hardware, you're not running that with the workflows that I do every day, and the workflows that I do every day, and the workflows that I do every day, and you're not going to write code with it you're not going to write code with it you're not going to write code with it for UI stuff even though it is much for UI stuff even though it is much for UI stuff even though it is much better at UI because it can't take the better at UI because it can't take the better at UI because it can't take the feedback. It can't even see what it feedback. It can't even see what it feedback. It can't even see what it made. It just knows the code. There's made. It just knows the code. There's made. It just knows the code. There's also the reality that the best models also the reality that the best models also the reality that the best models are getting bigger. Something like Fable are getting bigger. Something like Fable are getting bigger. Something like Fable is not a one trillion parameter model.
-
is not a one trillion parameter model. is not a one trillion parameter model. It's between two and 10. We have no It's between two and 10. We have no It's between two and 10. We have no idea, but it is massive, massive, idea, but it is massive, massive, idea, but it is massive, massive, massive. You're not running that on massive. You're not running that on massive. You're not running that on anything consumer, ever. anything consumer, ever. anything consumer, ever. And even if you can, we have another And even if you can, we have another And even if you can, we have another problem. problem. problem. Electricity costs. Let's ask a good open Electricity costs. Let's ask a good open Electricity costs. Let's ask a good open weight model about this. GLM 52. How weight model about this. GLM 52. How weight model about this. GLM 52. How much would running an RTX 5090 24/7 much would running an RTX 5090 24/7 much would running an RTX 5090 24/7 cost electricity-wise cost electricity-wise cost electricity-wise in San Francisco? in San Francisco? in San Francisco? Give it search. Give it search. Give it search. Well, at GLM 52, do some math here. Well, at GLM 52, do some math here. Well, at GLM 52, do some math here. The cost for running my 5090 24/7 in San The cost for running my 5090 24/7 in San The cost for running my 5090 24/7 in San Francisco is around $5 a day just for Francisco is around $5 a day just for Francisco is around $5 a day just for electricity. electricity. electricity. That's two grand a year in electricity That's two grand a year in electricity That's two grand a year in electricity costs. Do you understand? It's not costs. Do you understand? It's not costs. Do you understand? It's not cheap. Just cuz you own the hardware cheap. Just cuz you own the hardware cheap. Just cuz you own the hardware doesn't mean you're getting away with doesn't mean you're getting away with doesn't mean you're getting away with this for free. That's just one GPU. If this for free. That's just one GPU. If this for free. That's just one GPU. If you want to run these bigger models, you you want to run these bigger models, you you want to run these bigger models, you need a lot more than one. The costs are need a lot more than one. The costs are need a lot more than one. The costs are crazy.
-
crazy. crazy. I do see one other type of cope I need I do see one other type of cope I need I do see one other type of cope I need to address in chat. Not everyone needs a to address in chat. Not everyone needs a to address in chat. Not everyone needs a frontier model for everything they do. I frontier model for everything they do. I frontier model for everything they do. I don't disagree, but I think the things don't disagree, but I think the things don't disagree, but I think the things you're talking about here that frontier you're talking about here that frontier you're talking about here that frontier doesn't benefit are not the growing use doesn't benefit are not the growing use doesn't benefit are not the growing use case. You need to be realistic here. Do case. You need to be realistic here. Do case. You need to be realistic here. Do you think the growth in token usage you think the growth in token usage you think the growth in token usage month-over-month and year-over-year is month-over-month and year-over-year is month-over-month and year-over-year is mostly going to these small cheap models mostly going to these small cheap models mostly going to these small cheap models for those types of tasks? Tasks that for those types of tasks? Tasks that for those types of tasks? Tasks that could be solved by small cheap models could be solved by small cheap models could be solved by small cheap models got solved by small cheap models a while got solved by small cheap models a while got solved by small cheap models a while ago. The cool thing about the ago. The cool thing about the ago. The cool thing about the development of frontier models is they development of frontier models is they development of frontier models is they increase the types of work that can be increase the types of work that can be increase the types of work that can be done with models. When new better models done with models. When new better models done with models. When new better models come out, my personal token spend goes come out, my personal token spend goes come out, my personal token spend goes up massively because all of a sudden up massively because all of a sudden up massively because all of a sudden more work is able to be done with these more work is able to be done with these more work is able to be done with these models. More ends of the work can be models. More ends of the work can be models. More ends of the work can be reached with them where I can start at reached with them where I can start at reached with them where I can start at an earlier place before the LM comes in an earlier place before the LM comes in an earlier place before the LM comes in and then grab it again at a much later and then grab it again at a much later and then grab it again at a much later space after the LM's done more of the space after the LM's done more of the space after the LM's done more of the work. And you do need better models for work. And you do need better models for work. And you do need better models for things like vision, for things like things like vision, for things like things like vision, for things like accessing APIs to get the feedback from accessing APIs to get the feedback from accessing APIs to get the feedback from the code reviews and then addressing the code reviews and then addressing the code reviews and then addressing them. For owning this long process that them. For owning this long process that them. For owning this long process that is doing millions of tokens instead of is doing millions of tokens instead of is doing millions of tokens instead of dozens, you do need things like the dozens, you do need things like the dozens, you do need things like the current frontier. GLM-52 is the first current frontier. GLM-52 is the first current frontier. GLM-52 is the first open weight model that can actually run open weight model that can actually run open weight model that can actually run for a while, and it's really impressive for a while, and it's really impressive for a while, and it's really impressive that it can do that. It changes the that it can do that. It changes the that it can do that. It changes the economics of using models on a economics of using models on a economics of using models on a fundamental level in a way that's fundamental level in a way that's fundamental level in a way that's genuinely really cool and exciting.
-
genuinely really cool and exciting. genuinely really cool and exciting. That's why I love open weight models. That's why I love open weight models. That's why I love open weight models. But if you think consumers using ChatGPT But if you think consumers using ChatGPT But if you think consumers using ChatGPT are where most of the inference that are where most of the inference that are where most of the inference that OpenAI is doing is, and not the small OpenAI is doing is, and not the small OpenAI is doing is, and not the small number of devs and companies that are number of devs and companies that are number of devs and companies that are cranking the highest end models doing cranking the highest end models doing cranking the highest end models doing crazy long jobs, I can send a single crazy long jobs, I can send a single crazy long jobs, I can send a single prompt that does 10 million tokens. No prompt that does 10 million tokens. No prompt that does 10 million tokens. No consumer can. So, is it cool that 1% of consumer can. So, is it cool that 1% of consumer can. So, is it cool that 1% of token use could theoretically be on a token use could theoretically be on a token use could theoretically be on a local model that runs on your laptop or local model that runs on your laptop or local model that runs on your laptop or phone? Yeah, until you consider the fact phone? Yeah, until you consider the fact phone? Yeah, until you consider the fact it's going to kill their battery and it's going to kill their battery and it's going to kill their battery and cause the thing to overheat. Like Mark cause the thing to overheat. Like Mark cause the thing to overheat. Like Mark says here, things like Gemini 3 and says here, things like Gemini 3 and says here, things like Gemini 3 and fitting on tiny tensor chips on your fitting on tiny tensor chips on your fitting on tiny tensor chips on your phone is really cool because it can do phone is really cool because it can do phone is really cool because it can do message summaries and weather summaries message summaries and weather summaries message summaries and weather summaries without having to go to the cloud, which without having to go to the cloud, which without having to go to the cloud, which is a real valuable use case. Not needing is a real valuable use case. Not needing is a real valuable use case. Not needing to have your data, which is fully to have your data, which is fully to have your data, which is fully unencrypted once it reaches the servers, unencrypted once it reaches the servers, unencrypted once it reaches the servers, by the way, go to the server to serve by the way, go to the server to serve by the way, go to the server to serve the inference. That is awesome. And I do the inference. That is awesome. And I do the inference. That is awesome. And I do want ways to run models without having want ways to run models without having want ways to run models without having to send the data to another company or to send the data to another company or to send the data to another company or cloud, which is why I am excited about cloud, which is why I am excited about cloud, which is why I am excited about some amount of this getting better. And some amount of this getting better. And some amount of this getting better. And like the ways Apple does things on your like the ways Apple does things on your like the ways Apple does things on your phone using a little bit of local phone using a little bit of local phone using a little bit of local intelligence to not have to use a intelligence to not have to use a intelligence to not have to use a server, as well as the secure compute server, as well as the secure compute server, as well as the secure compute stuff, which is even cooler. There is a stuff, which is even cooler. There is a stuff, which is even cooler. There is a problem, though. Your phone only has so problem, though. Your phone only has so problem, though. Your phone only has so much power. This both limits the much power. This both limits the much power. This both limits the capability of the model that you're capability of the model that you're capability of the model that you're running on your phone, but it also means running on your phone, but it also means running on your phone, but it also means that you're going to destroy the battery that you're going to destroy the battery that you're going to destroy the battery on your phone in really egregious ways, on your phone in really egregious ways, on your phone in really egregious ways, which would be much nicer if you could which would be much nicer if you could which would be much nicer if you could just hit an API. But Theo, computers are just hit an API. But Theo, computers are just hit an API. But Theo, computers are getting so much more powerful. Phones getting so much more powerful. Phones getting so much more powerful. Phones get better every year. Aren't they going get better every year. Aren't they going get better every year. Aren't they going to get good enough that you can run to get good enough that you can run to get good enough that you can run basically anything on them?
-
basically anything on them? basically anything on them? I'm going to ask you a question. I'm going to ask you a question. I'm going to ask you a question. How much faster do you think the highest How much faster do you think the highest How much faster do you think the highest end Android phone is today end Android phone is today end Android phone is today than it was 3 years ago? Better than it was 3 years ago? Better than it was 3 years ago? Better question. How much faster do you think a question. How much faster do you think a question. How much faster do you think a $200 Android phone is today than it was $200 Android phone is today than it was $200 Android phone is today than it was 3 years ago? Just get like a rough 3 years ago? Just get like a rough 3 years ago? Just get like a rough number in your head. Here are the number in your head. Here are the number in your head. Here are the numbers. Sure, iPhones continue growing numbers. Sure, iPhones continue growing numbers. Sure, iPhones continue growing and it looks like Xiaomi has come in and and it looks like Xiaomi has come in and and it looks like Xiaomi has come in and actually made some meaningful actually made some meaningful actually made some meaningful improvement on the Android side, which improvement on the Android side, which improvement on the Android side, which is cool to see. That was not the case is cool to see. That was not the case is cool to see. That was not the case before. When you look at the cheap before. When you look at the cheap before. When you look at the cheap devices here, there has been literally devices here, there has been literally devices here, there has been literally no improvement in performance. If no improvement in performance. If no improvement in performance. If anything, there has been slight anything, there has been slight anything, there has been slight regressions in the cheaper phone tier. regressions in the cheaper phone tier. regressions in the cheaper phone tier. And that's before the price hikes And that's before the price hikes And that's before the price hikes because of RAM and demand and all of the because of RAM and demand and all of the because of RAM and demand and all of the manufacturing stuff getting [ __ ] This manufacturing stuff getting [ __ ] This manufacturing stuff getting [ __ ] This is going to be way worse in 2026. If is going to be way worse in 2026. If is going to be way worse in 2026. If you're just counting on our devices to you're just counting on our devices to you're just counting on our devices to get powerful enough to do this, here's a get powerful enough to do this, here's a get powerful enough to do this, here's a chart that tells you you're wrong. It's chart that tells you you're wrong. It's chart that tells you you're wrong. It's not happening. Budget device CPUs have not happening. Budget device CPUs have not happening. Budget device CPUs have not meaningfully improved since 2022. not meaningfully improved since 2022. not meaningfully improved since 2022. We're not anti-open weight models, we're We're not anti-open weight models, we're We're not anti-open weight models, we're just annoyed at local models and the just annoyed at local models and the just annoyed at local models and the people who think they are the future.
-
people who think they are the future. people who think they are the future. This message from Fry nails this. I This message from Fry nails this. I This message from Fry nails this. I don't think the major benefit of the don't think the major benefit of the don't think the major benefit of the open-source models is to run them open-source models is to run them open-source models is to run them locally. They enable competition on the locally. They enable competition on the locally. They enable competition on the provider and hardware side. Hmm. No provider and hardware side. Hmm. No provider and hardware side. Hmm. No provider can offer Opus at a cheaper provider can offer Opus at a cheaper provider can offer Opus at a cheaper price, even if they have spare margins, price, even if they have spare margins, price, even if they have spare margins, because of licensing with Infrobic. Hmm. because of licensing with Infrobic. Hmm. because of licensing with Infrobic. Hmm. There are providers running on places There are providers running on places There are providers running on places with cheaper electricity that offer with cheaper electricity that offer with cheaper electricity that offer cheaper inference of 5.2. Competition is cheaper inference of 5.2. Competition is cheaper inference of 5.2. Competition is a win. a win. a win. Ding ding [ __ ] ding. Ding ding [ __ ] ding. Ding ding [ __ ] ding. You can see the real benefits of this You can see the real benefits of this You can see the real benefits of this when you go somewhere like OpenRouter. when you go somewhere like OpenRouter. when you go somewhere like OpenRouter. Let's look at GPT 5.5 to start. We have Let's look at GPT 5.5 to start. We have Let's look at GPT 5.5 to start. We have options here. We can use Azure, we can options here. We can use Azure, we can options here. We can use Azure, we can use Azure in Europe, or we can use Open use Azure in Europe, or we can use Open use Azure in Europe, or we can use Open AI's models, or or we can use Open AI's AI's models, or or we can use Open AI's AI's models, or or we can use Open AI's hosting. There are your options. There's hosting. There are your options. There's hosting. There are your options. There's all of your options. And they're all all of your options. And they're all all of your options. And they're all priced pretty much identically and the priced pretty much identically and the priced pretty much identically and the performance is relatively similar. Funny performance is relatively similar. Funny performance is relatively similar. Funny enough, it wasn't similar until I enough, it wasn't similar until I enough, it wasn't similar until I crashed out at Azure and now it's a lot crashed out at Azure and now it's a lot crashed out at Azure and now it's a lot better. Apparently, the hosting of Open better. Apparently, the hosting of Open better. Apparently, the hosting of Open AI models on Bedrock entirely broken AI models on Bedrock entirely broken AI models on Bedrock entirely broken right now, so right now, so right now, so I was right in my prediction that uh it I was right in my prediction that uh it I was right in my prediction that uh it was not going to work on Tranium. I was not going to work on Tranium. I was not going to work on Tranium. I called it, everybody said I was wrong. I called it, everybody said I was wrong. I called it, everybody said I was wrong. I was [ __ ] right. I'm going to take was [ __ ] right. I'm going to take was [ __ ] right. I'm going to take that win. I'm going to be proud of it.
-
that win. I'm going to be proud of it. that win. I'm going to be proud of it. But we aren't here to talk about closed But we aren't here to talk about closed But we aren't here to talk about closed models. We're here to talk about open models. We're here to talk about open models. We're here to talk about open ones. And just cuz they're called open ones. And just cuz they're called open ones. And just cuz they're called open AI doesn't mean their models are open, AI doesn't mean their models are open, AI doesn't mean their models are open, sadly. Let's look at 5.2. sadly. Let's look at 5.2. sadly. Let's look at 5.2. Oh. Oh. Oh. That's a lot of options. And they have That's a lot of options. And they have That's a lot of options. And they have different pricing. different pricing. different pricing. Some of them are way faster and cost Some of them are way faster and cost Some of them are way faster and cost more money. Like Wafer Fast at 115 TPS more money. Like Wafer Fast at 115 TPS more money. Like Wafer Fast at 115 TPS and 1025. Or Fireworks Fast, which has a and 1025. Or Fireworks Fast, which has a and 1025. Or Fireworks Fast, which has a little bit of reliability problems, but little bit of reliability problems, but little bit of reliability problems, but still is only 660 out and 130 TPS. still is only 660 out and 130 TPS. still is only 660 out and 130 TPS. That's awesome. Or Friendly, who That's awesome. Or Friendly, who That's awesome. Or Friendly, who apparently is doing 117 TPS for the same apparently is doing 117 TPS for the same apparently is doing 117 TPS for the same 440 that a lot of the others are doing. 440 that a lot of the others are doing. 440 that a lot of the others are doing. But then there's companies like Deep But then there's companies like Deep But then there's companies like Deep Infra that'll go even cheaper. They're a Infra that'll go even cheaper. They're a Infra that'll go even cheaper. They're a little slower at 30 TPS, which is a little slower at 30 TPS, which is a little slower at 30 TPS, which is a compromise that makes sense for some compromise that makes sense for some compromise that makes sense for some people, but you can choose to be cheaper people, but you can choose to be cheaper people, but you can choose to be cheaper there. You can choose to spend a bit there. You can choose to spend a bit there. You can choose to spend a bit more to get more reliability or more more to get more reliability or more more to get more reliability or more speed. You have options, and you have a speed. You have options, and you have a speed. You have options, and you have a lot of them. This is where open weight lot of them. This is where open weight lot of them. This is where open weight models really shine. They allow for models really shine. They allow for models really shine. They allow for competition in the hosting space, which competition in the hosting space, which competition in the hosting space, which allows for the most of the problems I allows for the most of the problems I allows for the most of the problems I was talking about before to be solved. was talking about before to be solved. was talking about before to be solved. Let's like go back to this list again.
-
Let's like go back to this list again. Let's like go back to this list again. Okay. Gap between runnable and good. Okay. Gap between runnable and good. Okay. Gap between runnable and good. Doesn't matter when a cloud can host it Doesn't matter when a cloud can host it Doesn't matter when a cloud can host it for you. Hardware costs are insane. for you. Hardware costs are insane. for you. Hardware costs are insane. Doesn't matter when you're effectively Doesn't matter when you're effectively Doesn't matter when you're effectively renting. Parallelism. Doesn't matter at renting. Parallelism. Doesn't matter at renting. Parallelism. Doesn't matter at all when you're renting. Electricity all when you're renting. Electricity all when you're renting. Electricity costs factored into the token costs costs factored into the token costs costs factored into the token costs you're paying, and they can bring these you're paying, and they can bring these you're paying, and they can bring these GPUs and these data centers to places GPUs and these data centers to places GPUs and these data centers to places where power is cheaper. Almost all of where power is cheaper. Almost all of where power is cheaper. Almost all of the problems I have with this local the problems I have with this local the problems I have with this local model [ __ ] are solved by just throwing model [ __ ] are solved by just throwing model [ __ ] are solved by just throwing them on a cloud host. You do lose one of them on a cloud host. You do lose one of them on a cloud host. You do lose one of the coolest benefits, which is that no the coolest benefits, which is that no the coolest benefits, which is that no one else can see what you're doing. It's one else can see what you're doing. It's one else can see what you're doing. It's just yours. just yours. just yours. And I'm really excited for Secure And I'm really excited for Secure And I'm really excited for Secure Compute to make that a little more Compute to make that a little more Compute to make that a little more viable in the future. But for now, the viable in the future. But for now, the viable in the future. But for now, the best part of these open weight models is best part of these open weight models is best part of these open weight models is that you can get them for good prices that you can get them for good prices that you can get them for good prices with incredible performance on various with incredible performance on various with incredible performance on various different hosts. And believe me, I use different hosts. And believe me, I use different hosts. And believe me, I use those a lot. I think it's awesome that those a lot. I think it's awesome that those a lot. I think it's awesome that you can take these incredible you can take these incredible you can take these incredible developments in the model space that are developments in the model space that are developments in the model space that are all open weight, all available to all open weight, all available to all open weight, all available to download and use and retrain and do download and use and retrain and do download and use and retrain and do whatever you want with it. And then just whatever you want with it. And then just whatever you want with it. And then just run them in the cloud. It's great. It's run them in the cloud. It's great. It's run them in the cloud. It's great. It's awesome. If your work can be done by awesome. If your work can be done by awesome. If your work can be done by these types of models, like if you're these types of models, like if you're these types of models, like if you're doing work that works in Opus or Sonnet, doing work that works in Opus or Sonnet, doing work that works in Opus or Sonnet, but also works fine in GLM 5.2, and the but also works fine in GLM 5.2, and the but also works fine in GLM 5.2, and the failure rate isn't a meaningful failure rate isn't a meaningful failure rate isn't a meaningful difference, you should probably move difference, you should probably move difference, you should probably move that workload to 5.2. It's insane how that workload to 5.2. It's insane how that workload to 5.2. It's insane how much faster and cheaper it can be. There much faster and cheaper it can be. There much faster and cheaper it can be. There is one last catch though, because those is one last catch though, because those is one last catch though, because those prices aren't the whole story. Deep prices aren't the whole story. Deep prices aren't the whole story. Deep Infra has GLM 5.2 at $3 per million out Infra has GLM 5.2 at $3 per million out Infra has GLM 5.2 at $3 per million out versus Opus 4.8, for example, which is versus Opus 4.8, for example, which is versus Opus 4.8, for example, which is 25 per mill out. So, obviously, that's
-
25 per mill out. So, obviously, that's 25 per mill out. So, obviously, that's way cheaper, right? 3 to 25, that's a way cheaper, right? 3 to 25, that's a way cheaper, right? 3 to 25, that's a huge discount. It's almost 10x cheaper. huge discount. It's almost 10x cheaper. huge discount. It's almost 10x cheaper. What if I told you it was nowhere near What if I told you it was nowhere near What if I told you it was nowhere near that much cheaper? In reality, Opus on X that much cheaper? In reality, Opus on X that much cheaper? In reality, Opus on X High is about $8 per run with Deep SWE, High is about $8 per run with Deep SWE, High is about $8 per run with Deep SWE, and 5.2 on Max about $4. So, it's only a and 5.2 on Max about $4. So, it's only a and 5.2 on Max about $4. So, it's only a 2x gap. How the hell does that make any 2x gap. How the hell does that make any 2x gap. How the hell does that make any sense though? Isn't it literally 10 sense though? Isn't it literally 10 sense though? Isn't it literally 10 times cheaper per token? times cheaper per token? times cheaper per token? Here's the problem. Open weight models Here's the problem. Open weight models Here's the problem. Open weight models tend to be a good bit heavier on the tend to be a good bit heavier on the tend to be a good bit heavier on the amount of tokens that they burn during amount of tokens that they burn during amount of tokens that they burn during their runs. Opus 4.8 High more efficient their runs. Opus 4.8 High more efficient their runs. Opus 4.8 High more efficient than 5.2. Every single GPT 5.5 run, even than 5.2. Every single GPT 5.5 run, even than 5.2. Every single GPT 5.5 run, even X High, used less tokens than GLM 5.2 on X High, used less tokens than GLM 5.2 on X High, used less tokens than GLM 5.2 on High. This is the problem. These open High. This is the problem. These open High. This is the problem. These open weight models are not as efficient as weight models are not as efficient as weight models are not as efficient as the frontier models. So, even if the the frontier models. So, even if the the frontier models. So, even if the price per token is way cheaper, the price per token is way cheaper, the price per token is way cheaper, the massive increase in number of tokens is massive increase in number of tokens is massive increase in number of tokens is rough. And on top of that, the speed is rough. And on top of that, the speed is rough. And on top of that, the speed is worse because if it has to generate worse because if it has to generate worse because if it has to generate three times more tokens to get you an three times more tokens to get you an three times more tokens to get you an answer, doesn't matter if it's 20% answer, doesn't matter if it's 20% answer, doesn't matter if it's 20% faster if it has to generate 3x more faster if it has to generate 3x more faster if it has to generate 3x more tokens. This is the problem that kills tokens. This is the problem that kills tokens. This is the problem that kills 3.5 Flash, by the way. It is absurd how 3.5 Flash, by the way. It is absurd how 3.5 Flash, by the way. It is absurd how many tokens it uses. Those asking, "Why many tokens it uses. Those asking, "Why many tokens it uses. Those asking, "Why is the chart backwards?" It's cuz top is the chart backwards?" It's cuz top is the chart backwards?" It's cuz top right good. You want to be up and to the right good. You want to be up and to the right good. You want to be up and to the right. That's how people read charts.
-
right. That's how people read charts. right. That's how people read charts. Most of these types of charts tend to go Most of these types of charts tend to go Most of these types of charts tend to go that direction. One of the few sources that direction. One of the few sources that direction. One of the few sources that does this chart the other direction that does this chart the other direction that does this chart the other direction is artificial analysis. So, the top left is artificial analysis. So, the top left is artificial analysis. So, the top left is the good area. Do you notice is the good area. Do you notice is the good area. Do you notice something? Nothing's in the good area. something? Nothing's in the good area. something? Nothing's in the good area. 31 Pro preview just barely touches the 31 Pro preview just barely touches the 31 Pro preview just barely touches the corner there, but there's very little corner there, but there's very little corner there, but there's very little that is efficient cost-wise and also that is efficient cost-wise and also that is efficient cost-wise and also smart intelligence-wise. Once I clean up smart intelligence-wise. Once I clean up smart intelligence-wise. Once I clean up the chart a bit here, things get much, the chart a bit here, things get much, the chart a bit here, things get much, much clearer. GLM 52 Max is actually much clearer. GLM 52 Max is actually much clearer. GLM 52 Max is actually very interesting where it is slightly very interesting where it is slightly very interesting where it is slightly cheaper than Gemini 35 and also slightly cheaper than Gemini 35 and also slightly cheaper than Gemini 35 and also slightly smarter. That is good. That is awesome. smarter. That is good. That is awesome. smarter. That is good. That is awesome. That is why we like open weight models. That is why we like open weight models. That is why we like open weight models. This is a real competitive dot in this This is a real competitive dot in this This is a real competitive dot in this chart. But, watch what happens when I chart. But, watch what happens when I chart. But, watch what happens when I turn on GPT 55 on, I don't know, medium turn on GPT 55 on, I don't know, medium turn on GPT 55 on, I don't know, medium or low. or low. or low. Huh. Huh. Huh. Looks like 55 medium is also neck and Looks like 55 medium is also neck and Looks like 55 medium is also neck and neck with GLM 52 and is cheaper despite neck with GLM 52 and is cheaper despite neck with GLM 52 and is cheaper despite being way more expensive per token. being way more expensive per token. being way more expensive per token. And I'm not saying this to [ __ ] on GLM And I'm not saying this to [ __ ] on GLM And I'm not saying this to [ __ ] on GLM 52. It is an incredible model and I'm so 52. It is an incredible model and I'm so 52. It is an incredible model and I'm so thankful it exists. It's even better at thankful it exists. It's even better at thankful it exists. It's even better at UI than GPT 55 is somehow, which is UI than GPT 55 is somehow, which is UI than GPT 55 is somehow, which is crazy. They killed it with this model.
-
crazy. They killed it with this model. crazy. They killed it with this model. It's really goddamn good. But, it's not It's really goddamn good. But, it's not It's really goddamn good. But, it's not this magical solution people seem to this magical solution people seem to this magical solution people seem to think when they stare at the prices here think when they stare at the prices here think when they stare at the prices here and think that's all that matters or and think that's all that matters or and think that's all that matters or they hear they can run it locally and they hear they can run it locally and they hear they can run it locally and they get excited about that and then they get excited about that and then they get excited about that and then they try to and realize they literally they try to and realize they literally they try to and realize they literally cannot do that. Think I've said all I cannot do that. Think I've said all I cannot do that. Think I've said all I have to here. This is a long overdue have to here. This is a long overdue have to here. This is a long overdue crash out. I have been annoyed about crash out. I have been annoyed about crash out. I have been annoyed about this stuff for a long time. Open weight this stuff for a long time. Open weight this stuff for a long time. Open weight is awesome. The competitive marketplace is awesome. The competitive marketplace is awesome. The competitive marketplace is necessary. If we don't have good open is necessary. If we don't have good open is necessary. If we don't have good open weight developments, we're not going to weight developments, we're not going to weight developments, we're not going to have an ecosystem that has any incentive have an ecosystem that has any incentive have an ecosystem that has any incentive to keep getting better. Drops like Deep to keep getting better. Drops like Deep to keep getting better. Drops like Deep Seek R1 shifted the whole industry Seek R1 shifted the whole industry Seek R1 shifted the whole industry forward in a really powerful way. And I forward in a really powerful way. And I forward in a really powerful way. And I do not want a world where we are not do not want a world where we are not do not want a world where we are not trying to make the best possible open trying to make the best possible open trying to make the best possible open weight models as an industry. I just weight models as an industry. I just weight models as an industry. I just think it's insane to believe you're think it's insane to believe you're think it's insane to believe you're going to run these things on your local going to run these things on your local going to run these things on your local network on your own GPUs and get network on your own GPUs and get network on your own GPUs and get performance even close to reasonable, performance even close to reasonable, performance even close to reasonable, much less comparable to what we get from much less comparable to what we get from much less comparable to what we get from these frontier labs. And if you take these frontier labs. And if you take these frontier labs. And if you take this video and frame it as Theo hates this video and frame it as Theo hates this video and frame it as Theo hates open weight or open source, I really open weight or open source, I really open weight or open source, I really hope someone clips this part and shows hope someone clips this part and shows hope someone clips this part and shows it to you because I love open weight it to you because I love open weight it to you because I love open weight models. I built T3 chat because of how models. I built T3 chat because of how models. I built T3 chat because of how much I love the deep seek line of open much I love the deep seek line of open much I love the deep seek line of open weight models. I am a huge fan of open weight models. I am a huge fan of open weight models. I am a huge fan of open weight. I just think it's delusional to weight. I just think it's delusional to weight. I just think it's delusional to pretend anyone's going to run good pretend anyone's going to run good pretend anyone's going to run good frontier open weight models on their own frontier open weight models on their own frontier open weight models on their own hardware. And I wish you guys would stop hardware. And I wish you guys would stop hardware. And I wish you guys would stop pretending this is a realistic path pretending this is a realistic path pretending this is a realistic path forward when it obviously is not. So forward when it obviously is not. So forward when it obviously is not. So please, local model people, stop hurting please, local model people, stop hurting please, local model people, stop hurting the open source and open weight movement the open source and open weight movement the open source and open weight movement by pretending these things can run on by pretending these things can run on by pretending these things can run on your computers. We need to be able to your computers. We need to be able to your computers. We need to be able to talk about how good open weight models
-
talk about how good open weight models talk about how good open weight models are without tricking people into are without tricking people into are without tricking people into thinking they can run them on their own thinking they can run them on their own thinking they can run them on their own systems. They can't. It's stupid. It's systems. They can't. It's stupid. It's systems. They can't. It's stupid. It's unrealistic. It is really fun. If you're unrealistic. It is really fun. If you're unrealistic. It is really fun. If you're just doing this because you like the just doing this because you like the just doing this because you like the idea of having a bunch of GPUs in a idea of having a bunch of GPUs in a idea of having a bunch of GPUs in a cluster and playing with these things, cluster and playing with these things, cluster and playing with these things, awesome. That's so cool and fun. And you awesome. That's so cool and fun. And you awesome. That's so cool and fun. And you should talk all about it. But if you're should talk all about it. But if you're should talk all about it. But if you're pretending this is a thing everyone pretending this is a thing everyone pretending this is a thing everyone needs to do in order to get out of the needs to do in order to get out of the needs to do in order to get out of the pockets of Anthropic and OpenAI, pockets of Anthropic and OpenAI, pockets of Anthropic and OpenAI, you're insane. And I really hope you you're insane. And I really hope you you're insane. And I really hope you stop. stop. stop. >> [sighs and gasps] >> [sighs and gasps] >> [sighs and gasps] >> I've said all I have to here. This was a >> I've said all I have to here. This was a >> I've said all I have to here. This was a very, very overdue crash out and I am very, very overdue crash out and I am very, very overdue crash out and I am sure the comment section isn't a sure the comment section isn't a sure the comment section isn't a disaster at all. Let me know how you disaster at all. Let me know how you disaster at all. Let me know how you guys feel and until next time, guys feel and until next time, guys feel and until next time, peace, nerds.
Summary
The discussion focuses on the viability and practical limitations of running powerful open-weight AI models locally, referencing specific models like GLM-52 and comparing them to smaller, functional ones. The takeaway is that while open-weight models are crucial for AI advancement, the expectation of running the most advanced versions on consumer hardware is currently unrealistic due to immense size and resource requirements.