← Back
AI Engineer July 21, 2026 18m

The Desktop Frontier — Ahmad Osman, Osmantic

Read full transcript 14 segments
  1. Hey everyone, Hey everyone, we are about to start this presentation. we are about to start this presentation. we are about to start this presentation. Uh it's called the desktop frontier Uh it's called the desktop frontier Uh it's called the desktop frontier [snorts] and um [clears throat] [snorts] and um [clears throat] [snorts] and um [clears throat] it's basically about it's basically about it's basically about where we started and how far we've come where we started and how far we've come where we started and how far we've come with local and open source models. Um with local and open source models. Um with local and open source models. Um how many like just a quick question how how many like just a quick question how how many like just a quick question how many of you here follow me on X? many of you here follow me on X? many of you here follow me on X? >> I'm I'm amazing. Love you all. Love you >> I'm I'm amazing. Love you all. Love you >> I'm I'm amazing. Love you all. Love you all. Uh so you know I sometimes every all. Uh so you know I sometimes every all. Uh so you know I sometimes every now and then I would say a prediction uh now and then I would say a prediction uh now and then I would say a prediction uh here is a new one here is a new one here is a new one within roughly 18 months we are going to within roughly 18 months we are going to within roughly 18 months we are going to have the equivalent of GLM 5.2 class have the equivalent of GLM 5.2 class have the equivalent of GLM 5.2 class intelligence running on a single RTX intelligence running on a single RTX intelligence running on a single RTX 5090 with 32 GB of VRAM. Um that's 5090 with 32 GB of VRAM. Um that's 5090 with 32 GB of VRAM. Um that's basically late 2027. Um this this is basically late 2027. Um this this is basically late 2027. Um this this is conservative. We might actually get conservative. We might actually get conservative. We might actually get there faster.

  2. So uh you know for a long time uh the So uh you know for a long time uh the story has been bigger models bigger story has been bigger models bigger story has been bigger models bigger models bigger models how can we get to models bigger models how can we get to models bigger models how can we get to the next 5 trillion how can we get to the next 5 trillion how can we get to the next 5 trillion how can we get to the 20 trillion and I'm not saying that the 20 trillion and I'm not saying that the 20 trillion and I'm not saying that there won't ever be like a gap between there won't ever be like a gap between there won't ever be like a gap between frontier intelligence frontier intelligence frontier intelligence um and u you know open source models um and u you know open source models um and u you know open source models there will always be a gap but that gap there will always be a gap but that gap there will always be a gap but that gap um will shrink and the efficiency of the um will shrink and the efficiency of the um will shrink and the efficiency of the models models models will get exponentially better. So the the term that I like to think So the the term that I like to think about is impact per parameter. Um you about is impact per parameter. Um you about is impact per parameter. Um you know what capability are we talking know what capability are we talking know what capability are we talking about? What could a model do? Um what about? What could a model do? Um what about? What could a model do? Um what footprint like hardware footprint did it footprint like hardware footprint did it footprint like hardware footprint did it have last year in comparison to now? and have last year in comparison to now? and have last year in comparison to now? and uh what hardware does that use and what uh what hardware does that use and what uh what hardware does that use and what hardware does it need to use a year ago hardware does it need to use a year ago hardware does it need to use a year ago and uh you know are we moving down for and uh you know are we moving down for and uh you know are we moving down for the same kind of quality on the on that the same kind of quality on the on that the same kind of quality on the on that hardware again as I was saying earlier I hardware again as I was saying earlier I hardware again as I was saying earlier I used to run lama 2 on an RTX 3090 it's used to run lama 2 on an RTX 3090 it's used to run lama 2 on an RTX 3090 it's now running qu 3.5 3.6 6 27 billion now running qu 3.5 3.6 6 27 billion now running qu 3.5 3.6 6 27 billion parameter. That's better than Lamas 405.

  3. parameter. That's better than Lamas 405. parameter. That's better than Lamas 405. That's a 400 billion 400 billion plus That's a 400 billion 400 billion plus That's a 400 billion 400 billion plus parameters model that you beat with a 27 parameters model that you beat with a 27 parameters model that you beat with a 27 billion parameter model a year and a billion parameter model a year and a billion parameter model a year and a half after. So um yeah, as I was saying, similar So um yeah, as I was saying, similar capabilities are moving into smaller capabilities are moving into smaller capabilities are moving into smaller hardware footprint. Um hardware footprint. Um hardware footprint. Um benchmark scores are one thing but also benchmark scores are one thing but also benchmark scores are one thing but also you know a year ago this time a year ago you know a year ago this time a year ago you know a year ago this time a year ago we didn't have any local models that we didn't have any local models that we didn't have any local models that were able to successfully run within clo were able to successfully run within clo were able to successfully run within clo code right it wasn't until GLM 4.5 that code right it wasn't until GLM 4.5 that code right it wasn't until GLM 4.5 that came out in late July and GLM 4.5 air came out in late July and GLM 4.5 air came out in late July and GLM 4.5 air required at least four RTX uh 3090s or required at least four RTX uh 3090s or required at least four RTX uh 3090s or an RTX Pro 6000. Now that footprint for an RTX Pro 6000. Now that footprint for an RTX Pro 6000. Now that footprint for hardware is not needed anymore. All that hardware is not needed anymore. All that hardware is not needed anymore. All that you need is a single RTX 1390 1590 and you need is a single RTX 1390 1590 and you need is a single RTX 1390 1590 and you have something much more capable, you have something much more capable, you have something much more capable, much more intelligent. So is this trend much more intelligent. So is this trend much more intelligent. So is this trend just random or is there more to it?

  4. just random or is there more to it? just random or is there more to it? That's a question that everyone should That's a question that everyone should That's a question that everyone should ask. ask. ask. Um, is it just by by random chance that Um, is it just by by random chance that Um, is it just by by random chance that we've gotten this far from models that we've gotten this far from models that we've gotten this far from models that were weren't able to sustain more than were weren't able to sustain more than were weren't able to sustain more than 4,000 4,000 4,000 um tokens in terms of context lengths um tokens in terms of context lengths um tokens in terms of context lengths and now we have things that are million and now we have things that are million and now we have things that are million parame million uh tokens locally on your parame million uh tokens locally on your parame million uh tokens locally on your hardware that you own. It's not by hardware that you own. It's not by hardware that you own. It's not by chance, you know, it's not just a chance, you know, it's not just a chance, you know, it's not just a coincidence that we got here. Um, there coincidence that we got here. Um, there coincidence that we got here. Um, there is research being done. There is is research being done. There is is research being done. There is efficiency gains to be made. There are efficiency gains to be made. There are efficiency gains to be made. There are architecture hacks that compound and architecture hacks that compound and architecture hacks that compound and they will continue to compound. they will continue to compound. they will continue to compound. And I think I like this line. It's not And I think I like this line. It's not And I think I like this line. It's not that small models are beating big that small models are beating big that small models are beating big models. It's that newer, more efficient models. It's that newer, more efficient models. It's that newer, more efficient models are beating older, less efficient models are beating older, less efficient models are beating older, less efficient ones. ones. ones. Uh so yeah capability density uh is you Uh so yeah capability density uh is you Uh so yeah capability density uh is you know the literature I back this up with know the literature I back this up with know the literature I back this up with uh uh nature machine intelligence uh uh uh nature machine intelligence uh uh uh nature machine intelligence uh calls this pattern densing law and uh calls this pattern densing law and uh calls this pattern densing law and uh basically um you know every three and a basically um you know every three and a basically um you know every three and a half months we are having 50% fear half months we are having 50% fear half months we are having 50% fear parameters whether that's in dense or parameters whether that's in dense or parameters whether that's in dense or activated that's a different story but activated that's a different story but activated that's a different story but we're getting way more intelligence out we're getting way more intelligence out we're getting way more intelligence out of the models that we're running.

  5. So, you know, right now where we're at, So, you know, right now where we're at, it's uh GLM 5.2. That's uh our, you it's uh GLM 5.2. That's uh our, you it's uh GLM 5.2. That's uh our, you know, biggest player and um it's 744 know, biggest player and um it's 744 know, biggest player and um it's 744 billion parameters total with only 40 billion parameters total with only 40 billion parameters total with only 40 billion parameter activated. And that billion parameter activated. And that billion parameter activated. And that supports up to 1 million contexts. You supports up to 1 million contexts. You supports up to 1 million contexts. You can run this in MVFP4 on a machine uh on can run this in MVFP4 on a machine uh on can run this in MVFP4 on a machine uh on a GGX station or on a server with eight a GGX station or on a server with eight a GGX station or on a server with eight RTX Pro 6000. That's something that you RTX Pro 6000. That's something that you RTX Pro 6000. That's something that you like a GX station is something that you like a GX station is something that you like a GX station is something that you can sit under your desk and it's running can sit under your desk and it's running can sit under your desk and it's running this kind of frontier intelligence. this kind of frontier intelligence. this kind of frontier intelligence. Whether you know it it's on one Whether you know it it's on one Whether you know it it's on one benchmark it actually beats GBT 5.5 benchmark it actually beats GBT 5.5 benchmark it actually beats GBT 5.5 extra high. Doesn't that mean that we're extra high. Doesn't that mean that we're extra high. Doesn't that mean that we're getting somewhere with local and open getting somewhere with local and open getting somewhere with local and open source models that we can compete with source models that we can compete with source models that we can compete with the frontier that we're not that far off the frontier that we're not that far off the frontier that we're not that far off from the best that you can get from the from the best that you can get from the from the best that you can get from the cloud? cloud? cloud? We also have Neatron 3 ultra which We also have Neatron 3 ultra which We also have Neatron 3 ultra which proved that NVFB4 training more proved that NVFB4 training more proved that NVFB4 training more efficient training can be done on uh on efficient training can be done on uh on efficient training can be done on uh on hardware right that's that's very hardware right that's that's very hardware right that's that's very important that means that the footprint important that means that the footprint important that means that the footprint even for training these models for even for training these models for even for training these models for fine-tuning them for making small and fine-tuning them for making small and fine-tuning them for making small and specialized models as I was talking specialized models as I was talking specialized models as I was talking earlier could be more efficient could be earlier could be more efficient could be earlier could be more efficient could be done cheaper and could be you know could done cheaper and could be you know could done cheaper and could be you know could deliver you value in terms of economics deliver you value in terms of economics deliver you value in terms of economics way sooner or you know for much less way sooner or you know for much less way sooner or you know for much less money than you used Yeah. So, you know, again, Lama 2, uh, Yeah. So, you know, again, Lama 2, uh, that was a 70 billion parameter model.

  6. that was a 70 billion parameter model. that was a 70 billion parameter model. If you try to run that right now, you're If you try to run that right now, you're If you try to run that right now, you're you'd laugh at it, right? That used to you'd laugh at it, right? That used to you'd laugh at it, right? That used to take eight RTX 1390s to load up and it take eight RTX 1390s to load up and it take eight RTX 1390s to load up and it those same eight RTX3090s could run those same eight RTX3090s could run those same eight RTX3090s could run something like 15 parallel agents right something like 15 parallel agents right something like 15 parallel agents right now with Quen 3.5 27. That's that's a now with Quen 3.5 27. That's that's a now with Quen 3.5 27. That's that's a massive jump in terms of performance massive jump in terms of performance massive jump in terms of performance gains. Um so the densing law basically gains. Um so the densing law basically gains. Um so the densing law basically means that we have similar or better um means that we have similar or better um means that we have similar or better um capabilities with significantly fewer capabilities with significantly fewer capabilities with significantly fewer parameters. That's the impact per parameters. That's the impact per parameters. That's the impact per parameter. As I was saying I want parameter. As I was saying I want parameter. As I was saying I want everybody to live here thinking about everybody to live here thinking about everybody to live here thinking about this term and you know thinking where this term and you know thinking where this term and you know thinking where are we going to get a year from today as are we going to get a year from today as are we going to get a year from today as I was saying earlier everyone here has a I was saying earlier everyone here has a I was saying earlier everyone here has a phone. I'm assuming raise your hand if phone. I'm assuming raise your hand if phone. I'm assuming raise your hand if you have a phone. you have a phone. you have a phone. If you didn't raise your hand, we know If you didn't raise your hand, we know If you didn't raise your hand, we know you lie about other things as well. So, you lie about other things as well. So, you lie about other things as well. So, come on, guys. come on, guys. come on, guys. So, [clears throat] So, [clears throat] So, [clears throat] you know, you can run you can now run you know, you can run you can now run you know, you can run you can now run GBT40 quality on your iPhone. That GBT40 quality on your iPhone. That GBT40 quality on your iPhone. That that's massive. That thing require data that's massive. That thing require data that's massive. That thing require data centers to serve. So, why wouldn't you centers to serve. So, why wouldn't you centers to serve. So, why wouldn't you invest, you know, in sovereign [snorts] invest, you know, in sovereign [snorts] invest, you know, in sovereign [snorts] AI? Why wouldn't you as a consumer, as AI? Why wouldn't you as a consumer, as AI? Why wouldn't you as a consumer, as an individual, as a smalls size an individual, as a smalls size an individual, as a smalls size business, middlesiz business, business, middlesiz business, business, middlesiz business, enterprise, why wouldn't you want to be enterprise, why wouldn't you want to be enterprise, why wouldn't you want to be in control of the models that you're on?

  7. in control of the models that you're on? in control of the models that you're on? Why wouldn't you want to Why wouldn't you want to Why wouldn't you want to make sure that nothing gets taken away make sure that nothing gets taken away make sure that nothing gets taken away from you? That every little thing can be from you? That every little thing can be from you? That every little thing can be optimized for you later on, that the optimized for you later on, that the optimized for you later on, that the performance gains can be made specially performance gains can be made specially performance gains can be made specially and specifically for your use cases, and and specifically for your use cases, and and specifically for your use cases, and that you can save more money that way in that you can save more money that way in that you can save more money that way in the long run. the long run. the long run. and you know ODS for consumers it's and you know ODS for consumers it's and you know ODS for consumers it's basically the way that we support basically the way that we support basically the way that we support individuals but enterprises also and I individuals but enterprises also and I individuals but enterprises also and I think that there is something that we think that there is something that we think that there is something that we like as a community we need to think like as a community we need to think like as a community we need to think about deeply we need enterprises for about deeply we need enterprises for about deeply we need enterprises for open source AI to win we need these open source AI to win we need these open source AI to win we need these people that are using the cloud right people that are using the cloud right people that are using the cloud right now that are basically supporting data now that are basically supporting data now that are basically supporting data centers being built for cloud providers centers being built for cloud providers centers being built for cloud providers to come on this side to own their own to come on this side to own their own to come on this side to own their own hardware to own the stack fully end to hardware to own the stack fully end to hardware to own the stack fully end to end so that we can keep delivering open end so that we can keep delivering open end so that we can keep delivering open source models so that there is an source models so that there is an source models so that there is an incentive for open source providers to incentive for open source providers to incentive for open source providers to actually come up with models so that we actually come up with models so that we actually come up with models so that we can come up with new licenses that can come up with new licenses that can come up with new licenses that allows open source to thrive and so again um open weight and the frontier so again um open weight and the frontier I think I yeah sorry that was a misclick I think I yeah sorry that was a misclick I think I yeah sorry that was a misclick um [clears throat] you know so smaller um [clears throat] you know so smaller um [clears throat] you know so smaller models started bunching above the weight models started bunching above the weight models started bunching above the weight after lama 2 with mistral 7B one of my after lama 2 with mistral 7B one of my after lama 2 with mistral 7B one of my favorite models if you try to build that favorite models if you try to build that favorite models if you try to build that model right now and um clo code or oven model right now and um clo code or oven model right now and um clo code or oven code it's not going to work but it used

  8. code it's not going to work but it used code it's not going to work but it used to take so much in terms of hardware to take so much in terms of hardware to take so much in terms of hardware right that you would now get from a 9b right that you would now get from a 9b right that you would now get from a 9b mill model that I can run with telegram mill model that I can run with telegram mill model that I can run with telegram with or with hermes for example and uh with or with hermes for example and uh with or with hermes for example and uh do a lot of stuff with so we've come a do a lot of stuff with so we've come a do a lot of stuff with so we've come a long Okay. Uh we had that we had mixed long Okay. Uh we had that we had mixed long Okay. Uh we had that we had mixed trial 8 by 7B which you know everybody trial 8 by 7B which you know everybody trial 8 by 7B which you know everybody knows is an MOE. Then the progression knows is an MOE. Then the progression knows is an MOE. Then the progression went uh from that to Lamas 3 you know went uh from that to Lamas 3 you know went uh from that to Lamas 3 you know Lamas 3 8B was one of my favorites still Lamas 3 8B was one of my favorites still Lamas 3 8B was one of my favorites still is um it had unique identity in my is um it had unique identity in my is um it had unique identity in my opinion. Uh then we had like the 70 opinion. Uh then we had like the 70 opinion. Uh then we had like the 70 billion which was like the the thing billion which was like the the thing billion which was like the the thing that I would run basically on my 8 that I would run basically on my 8 that I would run basically on my 8 RTX3090s at home. Then there was like RTX3090s at home. Then there was like RTX3090s at home. Then there was like the 405 the 400 billion plus parameter the 405 the 400 billion plus parameter the 405 the 400 billion plus parameter lamas which again required a lot of lamas which again required a lot of lamas which again required a lot of hardware and if you put it now against hardware and if you put it now against hardware and if you put it now against 3.5 the 27 billion parameter would lose 3.5 the 27 billion parameter would lose 3.5 the 27 billion parameter would lose against it that's in the span of what against it that's in the span of what against it that's in the span of what two years two years and some no I think two years two years and some no I think two years two years and some no I think I think less than two years that's I think less than two years that's I think less than two years that's summer 2024 to March 2026 that's uh summer 2024 to March 2026 that's uh summer 2024 to March 2026 that's uh that's about 21 months and the next big that's about 21 months and the next big that's about 21 months and the next big thing in my opinion gamma 27V and then thing in my opinion gamma 27V and then thing in my opinion gamma 27V and then we had the Quinn 2.5 and that that was we had the Quinn 2.5 and that that was we had the Quinn 2.5 and that that was the moment that I was like okay we the moment that I was like okay we the moment that I was like okay we actually are making progress and the gap actually are making progress and the gap actually are making progress and the gap was shrinking between open source models was shrinking between open source models was shrinking between open source models and uh the frontier uh really lamas 3 and uh the frontier uh really lamas 3 and uh the frontier uh really lamas 3 saved like you know it really helped us saved like you know it really helped us saved like you know it really helped us a lot and then um Gwen 2.5 delivered a a lot and then um Gwen 2.5 delivered a a lot and then um Gwen 2.5 delivered a massive improvement and there was a lot

  9. massive improvement and there was a lot massive improvement and there was a lot of fine-tuning and experiments that of fine-tuning and experiments that of fine-tuning and experiments that could be done on that one there was could be done on that one there was could be done on that one there was amazing papers and um they helped the amazing papers and um they helped the amazing papers and um they helped the community immensely community immensely community immensely in my opinion. in my opinion. in my opinion. Then the next big thing was Deepseek R1 Then the next big thing was Deepseek R1 Then the next big thing was Deepseek R1 in my opinion and uh reasoning becoming in my opinion and uh reasoning becoming in my opinion and uh reasoning becoming something that you can run at home. That something that you can run at home. That something that you can run at home. That was a massive MOE almost 700 billion was a massive MOE almost 700 billion was a massive MOE almost 700 billion parameters. Um you know you had to have parameters. Um you know you had to have parameters. Um you know you had to have like a very beefy server to actually get like a very beefy server to actually get like a very beefy server to actually get it up and running. Um and then you know it up and running. Um and then you know it up and running. Um and then you know the improvements that came from just the improvements that came from just the improvements that came from just more training on that one and Deepseek more training on that one and Deepseek more training on that one and Deepseek R1 that was released in May last year R1 that was released in May last year R1 that was released in May last year made massive jump again. So it showed made massive jump again. So it showed made massive jump again. So it showed that post training could deliver more that post training could deliver more that post training could deliver more improvements on the same on the same improvements on the same on the same improvements on the same on the same checkpoints. Then GBT open source like GBT OSS 12B. Then GBT open source like GBT OSS 12B. Anyone remembers that one from last Anyone remembers that one from last Anyone remembers that one from last summer? Yeah. summer? Yeah. summer? Yeah. Nobody here used it. Come on, guys. I I Nobody here used it. Come on, guys. I I Nobody here used it. Come on, guys. I I need some help here. [laughter] need some help here. [laughter] need some help here. [laughter] Uh it was it was it was one of the first Uh it was it was it was one of the first Uh it was it was it was one of the first open source models that were able to open source models that were able to open source models that were able to successfully do tool calling. Um and uh successfully do tool calling. Um and uh successfully do tool calling. Um and uh it was a step forward. It showed us that it was a step forward. It showed us that it was a step forward. It showed us that we can do more with uh with the hardware we can do more with uh with the hardware we can do more with uh with the hardware that we have at running at home. Uh that that we have at running at home. Uh that that we have at running at home. Uh that was a footprint shift right from like was a footprint shift right from like was a footprint shift right from like you know that massive 700 billion you know that massive 700 billion you know that massive 700 billion parameters deepseek uh R1 that was yeah parameters deepseek uh R1 that was yeah parameters deepseek uh R1 that was yeah 671 billion parameters to something that

  10. 671 billion parameters to something that 671 billion parameters to something that was 1/5 of its size in GBTSS with was 1/5 of its size in GBTSS with was 1/5 of its size in GBTSS with comparable maybe better more agentic comparable maybe better more agentic comparable maybe better more agentic performance. Um [snorts] then the moment of uh quinc 3.5 that's then the moment of uh quinc 3.5 that's 397 that's three that's 397 billion 397 that's three that's 397 billion 397 that's three that's 397 billion parameters uh that's uh that's a BVOE parameters uh that's uh that's a BVOE parameters uh that's uh that's a BVOE and uh you know what's funny is that and uh you know what's funny is that and uh you know what's funny is that about uh it's about 15 times the size of about uh it's about 15 times the size of about uh it's about 15 times the size of uh the Quinn 3.6 and I'm here I'm uh the Quinn 3.6 and I'm here I'm uh the Quinn 3.6 and I'm here I'm comparing 3.5 to 3.6 6 of the dense 27 comparing 3.5 to 3.6 6 of the dense 27 comparing 3.5 to 3.6 6 of the dense 27 billion parameter model and that dense billion parameter model and that dense billion parameter model and that dense model beats it and that dense model has model beats it and that dense model has model beats it and that dense model has 40% higher number of activated 40% higher number of activated 40% higher number of activated parameters. So it's not that far off. parameters. So it's not that far off. parameters. So it's not that far off. That's massive amount of performance That's massive amount of performance That's massive amount of performance gains and a very small amount of time gains and a very small amount of time gains and a very small amount of time with massively different footprint in with massively different footprint in with massively different footprint in terms of hardware uh requirements. And terms of hardware uh requirements. And terms of hardware uh requirements. And that trend happened in like what two that trend happened in like what two that trend happened in like what two three months. So you know how far could three months. So you know how far could three months. So you know how far could we go from here? Um how far before we we go from here? Um how far before we we go from here? Um how far before we get to you know uh a recent model that get to you know uh a recent model that get to you know uh a recent model that there was some news about you know that there was some news about you know that there was some news about you know that is uh finally relaunch the game. How far is uh finally relaunch the game. How far is uh finally relaunch the game. How far before open source delivers something of before open source delivers something of before open source delivers something of that quality that you could run on your that quality that you could run on your that quality that you could run on your own hardware and you can control and own hardware and you can control and own hardware and you can control and will not be taken away from you and will will not be taken away from you and will will not be taken away from you and will not refuse a request from you.

  11. So again these are just some benchmarks So again these are just some benchmarks where you can see that an iteration on where you can see that an iteration on where you can see that an iteration on the 27 billion parameter model a little the 27 billion parameter model a little the 27 billion parameter model a little bit more post training bit more post training bit more post training proved it across all benchmarks and made proved it across all benchmarks and made proved it across all benchmarks and made it one against a model that is almost 15 it one against a model that is almost 15 it one against a model that is almost 15 size for 15 times its size. And again, remember this is 27 billion And again, remember this is 27 billion parameters activated versus 17 billion parameters activated versus 17 billion parameters activated versus 17 billion parameter activated. It's still parameter activated. It's still parameter activated. It's still massively the same amount like you know massively the same amount like you know massively the same amount like you know it's it's only 40% less in terms of the it's it's only 40% less in terms of the it's it's only 40% less in terms of the amount of time it would take to process amount of time it would take to process amount of time it would take to process things but it's 15 times smaller. That's things but it's 15 times smaller. That's things but it's 15 times smaller. That's a lot. a lot. a lot. So again um how long until the So again um how long until the So again um how long until the prediction I made earlier becomes prediction I made earlier becomes prediction I made earlier becomes plausible when I said that we're going plausible when I said that we're going plausible when I said that we're going to have the equivalent of GLM 5.2 2 to have the equivalent of GLM 5.2 2 to have the equivalent of GLM 5.2 2 running on an RTX uh 1590. running on an RTX uh 1590. running on an RTX uh 1590. This is the mass 17 months and this is a This is the mass 17 months and this is a This is the mass 17 months and this is a conservative math. Uh earlier this year conservative math. Uh earlier this year conservative math. Uh earlier this year in December, I had a a very viral post in December, I had a a very viral post in December, I had a a very viral post that I predicted that we're going to that I predicted that we're going to that I predicted that we're going to have the quality of OBS 4.5 running have the quality of OBS 4.5 running have the quality of OBS 4.5 running locally at home on a single RTX uh Pro locally at home on a single RTX uh Pro locally at home on a single RTX uh Pro 6000. That happened by March.

  12. So a question So a question um hardware purchase today does it get um hardware purchase today does it get um hardware purchase today does it get more valuable as models become more more valuable as models become more more valuable as models become more efficient and smaller in size? efficient and smaller in size? efficient and smaller in size? That's a good question. So why are you That's a good question. So why are you That's a good question. So why are you funding other people to build data funding other people to build data funding other people to build data centers so that you can subscribe to centers so that you can subscribe to centers so that you can subscribe to them and pay subsidized tokens and then them and pay subsidized tokens and then them and pay subsidized tokens and then later on get those subsidies are going later on get those subsidies are going later on get those subsidies are going to go away and you're not going to be to go away and you're not going to be to go away and you're not going to be able to run those models and they will able to run those models and they will able to run those models and they will have so many limitations. So might as have so many limitations. So might as have so many limitations. So might as well ask yourself why not own the well ask yourself why not own the well ask yourself why not own the hardware yourself and be in control. hardware yourself and be in control. hardware yourself and be in control. Um, so yeah, the forwardl lookinging Um, so yeah, the forwardl lookinging Um, so yeah, the forwardl lookinging question is basically what will a DGX question is basically what will a DGX question is basically what will a DGX station be able to run in three, six, station be able to run in three, six, station be able to run in three, six, 12, 18 months from now? That's something 12, 18 months from now? That's something 12, 18 months from now? That's something that there is a reason that I'm not that there is a reason that I'm not that there is a reason that I'm not selling any of my RTX3090s if you follow selling any of my RTX3090s if you follow selling any of my RTX3090s if you follow me. And I have a lot of hardware, guys. me. And I have a lot of hardware, guys. me. And I have a lot of hardware, guys. U but I'm interested in seeing what I U but I'm interested in seeing what I U but I'm interested in seeing what I could do with them in a year or two from could do with them in a year or two from could do with them in a year or two from now more than in the amount of money I now more than in the amount of money I now more than in the amount of money I would get for them today. would get for them today. would get for them today. This is not a financial advice by the This is not a financial advice by the This is not a financial advice by the way. Like let me make that very clear.

  13. way. Like let me make that very clear. way. Like let me make that very clear. Um so yeah um the disk side frontier Um so yeah um the disk side frontier Um so yeah um the disk side frontier potential um you know an Nvidia GX potential um you know an Nvidia GX potential um you know an Nvidia GX station could run a lot of uh today it station could run a lot of uh today it station could run a lot of uh today it could run GLM 5.2 to what will it be could run GLM 5.2 to what will it be could run GLM 5.2 to what will it be able to run tomorrow 6 months 18 months able to run tomorrow 6 months 18 months able to run tomorrow 6 months 18 months two years from today we know that you two years from today we know that you two years from today we know that you know RTX3090 [snorts] know RTX3090 [snorts] know RTX3090 [snorts] is the amber uh architecture from 2020 is the amber uh architecture from 2020 is the amber uh architecture from 2020 sells at higher value than MSRP today sells at higher value than MSRP today sells at higher value than MSRP today and it's still being utilized for a lot and it's still being utilized for a lot and it's still being utilized for a lot of use cases so what well at GGX station of use cases so what well at GGX station of use cases so what well at GGX station the actively developed blackwell the actively developed blackwell the actively developed blackwell architecture will be able to run in a architecture will be able to run in a architecture will be able to run in a few months a couple of years that's a few months a couple of years that's a few months a couple of years that's a Good question. Good question. Good question. So the question you have to ask yourself So the question you have to ask yourself So the question you have to ask yourself um if an RTX30590 with 32 GB of VRAM um if an RTX30590 with 32 GB of VRAM um if an RTX30590 with 32 GB of VRAM runs in the equivalent of a GLM 5.2 and runs in the equivalent of a GLM 5.2 and runs in the equivalent of a GLM 5.2 and 18 months and this is the question that 18 months and this is the question that 18 months and this is the question that everybody should be asking themselves everybody should be asking themselves everybody should be asking themselves and I want you all to be looking at the and I want you all to be looking at the and I want you all to be looking at the screen taking this very seriously. Okay.

  14. screen taking this very seriously. Okay. screen taking this very seriously. Okay. Should you buy a GPU?

Summary

The presentation explores the rapid advancement of local and open-source AI models, contrasting them with larger frontier intelligence models. Key references include specific model benchmarks like GLM 5.2 and hardware like RTX 5090, highlighting impressive "impact per parameter" gains. The takeaway is that AI capabilities are becoming exponentially more efficient and accessible, significantly shrinking the gap with cutting-edge systems.

View original episode ↗