I Didn’t Expect Local AI to Go This Far on a Laptop
Read full transcript 13 segments
-
This is one of the most powerful gaming This is one of the most powerful gaming laptops in the [music] world. And yes, laptops in the [music] world. And yes, laptops in the [music] world. And yes, of course it can game, but what you're of course it can game, but what you're of course it can game, but what you're watching right now is not a game. That's watching right now is not a game. That's watching right now is not a game. That's a video generated by an AI model running a video generated by an AI model running a video generated by an AI model running locally on this laptop, in case you locally on this laptop, in case you locally on this laptop, in case you couldn't tell. Oh, and there are two of couldn't tell. Oh, and there are two of couldn't tell. Oh, and there are two of them. them. them. Identical twins. Well, almost. This Identical twins. Well, almost. This Identical twins. Well, almost. This one's got an Nvidia RTX 5080 in it with one's got an Nvidia RTX 5080 in it with one's got an Nvidia RTX 5080 in it with 16 gigs of VRAM, and this one's got a 16 gigs of VRAM, and this one's got a 16 gigs of VRAM, and this one's got a 5090 in it with 24 gigs of VRAM. The 5090 in it with 24 gigs of VRAM. The 5090 in it with 24 gigs of VRAM. The difference between these machines isn't difference between these machines isn't difference between these machines isn't speed, they got the same processor, same speed, they got the same processor, same speed, they got the same processor, same RAM, it's VRAM. And that turns out to be RAM, it's VRAM. And that turns out to be RAM, it's VRAM. And that turns out to be the whole story of local AI. But before the whole story of local AI. But before the whole story of local AI. But before any of that, these are the MSI Raider 16 any of that, these are the MSI Raider 16 any of that, these are the MSI Raider 16 Max HX. And of course, they're gaming Max HX. And of course, they're gaming Max HX. And of course, they're gaming laptops first, so I had to check that laptops first, so I had to check that laptops first, so I had to check that out first. So in case you're not out first. So in case you're not out first. So in case you're not familiar, these are brand new and familiar, these are brand new and familiar, these are brand new and they've got Intel's Core Ultra 9, 24 they've got Intel's Core Ultra 9, 24 they've got Intel's Core Ultra 9, 24 cores of it, 32 gigs of RAM on each one cores of it, 32 gigs of RAM on each one cores of it, 32 gigs of RAM on each one of these. And the whole reason we're of these. And the whole reason we're of these. And the whole reason we're here today is that sweet sweet Nvidia here today is that sweet sweet Nvidia here today is that sweet sweet Nvidia RTX 5090, which is going to show us what RTX 5090, which is going to show us what RTX 5090, which is going to show us what it's capable of in this laptop format.
-
it's capable of in this laptop format. it's capable of in this laptop format. And of course, it's 5080 twin as well. And of course, it's 5080 twin as well. And of course, it's 5080 twin as well. Like I said before, everything is Like I said before, everything is Like I said before, everything is exactly the same, even this gorgeous exactly the same, even this gorgeous exactly the same, even this gorgeous OLED screen. Which uh if you're a OLED screen. Which uh if you're a OLED screen. Which uh if you're a developer and you're staring at a screen developer and you're staring at a screen developer and you're staring at a screen all day, um all day, um all day, um yeah, it's going to be nice for you. And yeah, it's going to be nice for you. And yeah, it's going to be nice for you. And developers play games on their time off developers play games on their time off developers play games on their time off sometimes, or sometimes during work when sometimes, or sometimes during work when sometimes, or sometimes during work when they shouldn't be. And check this out. they shouldn't be. And check this out. they shouldn't be. And check this out. There's actually a quick access panel. There's actually a quick access panel. There's actually a quick access panel. You just pop that open, and right there You just pop that open, and right there You just pop that open, and right there you have access to your RAM, bam, and you have access to your RAM, bam, and you have access to your RAM, bam, and the SSD. Probably shouldn't have done the SSD. Probably shouldn't have done the SSD. Probably shouldn't have done that that that with the laptop on. I was just so with the laptop on. I was just so with the laptop on. I was just so excited. I need to turn this off ASAP. excited. I need to turn this off ASAP. excited. I need to turn this off ASAP. Bad Alex. So yeah, that 32 gigs that it Bad Alex. So yeah, that 32 gigs that it Bad Alex. So yeah, that 32 gigs that it ships with, you don't have to uh stick ships with, you don't have to uh stick ships with, you don't have to uh stick with that. You can get something like with that. You can get something like with that. You can get something like this, and uh this, and uh this, and uh yeah, increase that memory. yeah, increase that memory. yeah, increase that memory. And of course, because this thing is And of course, because this thing is And of course, because this thing is built for gaming, we've got some serious built for gaming, we've got some serious built for gaming, we've got some serious cooling headroom here. These fans on the cooling headroom here. These fans on the cooling headroom here. These fans on the bottom are crazy. This thing is designed bottom are crazy. This thing is designed bottom are crazy. This thing is designed to run flat out for hours without to run flat out for hours without to run flat out for hours without breaking a sweat. Of course, keep the breaking a sweat. Of course, keep the breaking a sweat. Of course, keep the 400 W power brick handy cuz you're going 400 W power brick handy cuz you're going 400 W power brick handy cuz you're going to need that. Also, RGB everywhere, but to need that. Also, RGB everywhere, but to need that. Also, RGB everywhere, but I prefer to turn that off. That's just I prefer to turn that off. That's just I prefer to turn that off. That's just not my thing.
-
Please turn on. Please turn on. Oh, yes. Please turn on. Please turn on. Oh, yes. Signs of life. Just before we kick off Signs of life. Just before we kick off Signs of life. Just before we kick off the AI test here, which is what this the AI test here, which is what this the AI test here, which is what this video is about, I just want to mention video is about, I just want to mention video is about, I just want to mention that this is really cleverly laid out. that this is really cleverly laid out. that this is really cleverly laid out. We got USB-A ports here on the right We got USB-A ports here on the right We got USB-A ports here on the right along with a headphone jack. I've got along with a headphone jack. I've got along with a headphone jack. I've got two Thunderbolt ports on the left along two Thunderbolt ports on the left along two Thunderbolt ports on the left along with an SD card reader exactly where I with an SD card reader exactly where I with an SD card reader exactly where I would need them. And the stuff I don't would need them. And the stuff I don't would need them. And the stuff I don't normally need is on the back. HDMI, you normally need is on the back. HDMI, you normally need is on the back. HDMI, you plug that in and you're done. Ethernet plug that in and you're done. Ethernet plug that in and you're done. Ethernet port is back there. The power plug is port is back there. The power plug is port is back there. The power plug is back there. They're not taking up side back there. They're not taking up side back there. They're not taking up side real estate. Kind of like that. I daily real estate. Kind of like that. I daily real estate. Kind of like that. I daily drive a MacBook Pro and I wish maybe it drive a MacBook Pro and I wish maybe it drive a MacBook Pro and I wish maybe it would adopt some of that, but I'm sure would adopt some of that, but I'm sure would adopt some of that, but I'm sure we'll have some comments about that in we'll have some comments about that in we'll have some comments about that in the comment section. All right, let's the comment section. All right, let's the comment section. All right, let's climb this ladder. Of course, you have climb this ladder. Of course, you have climb this ladder. Of course, you have the ability to easily run Ollama on the ability to easily run Ollama on the ability to easily run Ollama on this, which is a tool that's pretty this, which is a tool that's pretty this, which is a tool that's pretty familiar to all you by now. And right familiar to all you by now. And right familiar to all you by now. And right now I'm running Ollama pull Qwen 3.5 4B. now I'm running Ollama pull Qwen 3.5 4B. now I'm running Ollama pull Qwen 3.5 4B. By default, Ollama is going to grab the By default, Ollama is going to grab the By default, Ollama is going to grab the 4-bit quantization, which means it's 4-bit quantization, which means it's 4-bit quantization, which means it's only going to be What is that? 3.4 gigs only going to be What is that? 3.4 gigs only going to be What is that? 3.4 gigs in size. Pretty small model. And for in size. Pretty small model. And for in size. Pretty small model. And for those of you that are not familiar, those of you that are not familiar, those of you that are not familiar, quantization, when you quantize a model, quantization, when you quantize a model, quantization, when you quantize a model, you squeeze it. So, you're basically you squeeze it. So, you're basically you squeeze it. So, you're basically taking a 16-bit model weights and you're taking a 16-bit model weights and you're taking a 16-bit model weights and you're compressing it down to 4 bits in this compressing it down to 4 bits in this compressing it down to 4 bits in this case. And we have 16 GB of VRAM on this case. And we have 16 GB of VRAM on this case. And we have 16 GB of VRAM on this machine. We have 24 on this one. So, machine. We have 24 on this one. So, machine. We have 24 on this one. So, you're going to need to be able to fit you're going to need to be able to fit you're going to need to be able to fit your models within that space. As I'll your models within that space. As I'll your models within that space. As I'll show you, it'll run larger models, but show you, it'll run larger models, but show you, it'll run larger models, but speed is going to suffer at point. We'll speed is going to suffer at point. We'll speed is going to suffer at point. We'll get into that momentarily. Ollama run get into that momentarily. Ollama run get into that momentarily. Ollama run verbose and go. Now, this is a small verbose and go. Now, this is a small verbose and go. Now, this is a small model, so you're most likely going to be
-
model, so you're most likely going to be model, so you're most likely going to be just chatting with it. I know you're all just chatting with it. I know you're all just chatting with it. I know you're all going to love this, but I'm just going going to love this, but I'm just going going to love this, but I'm just going to say hi to it. And boom, that is to say hi to it. And boom, that is to say hi to it. And boom, that is incredibly fast and it's fast for a few incredibly fast and it's fast for a few incredibly fast and it's fast for a few reasons. Let's be clear about that. reasons. Let's be clear about that. reasons. Let's be clear about that. These are Nvidia RTX GPUs in there and These are Nvidia RTX GPUs in there and These are Nvidia RTX GPUs in there and they're running CUDA, which is a they're running CUDA, which is a they're running CUDA, which is a software stack by Nvidia that allows you software stack by Nvidia that allows you software stack by Nvidia that allows you to run AI on machines like this as well to run AI on machines like this as well to run AI on machines like this as well as huge data centers with server racks. as huge data centers with server racks. as huge data centers with server racks. And the CUDA stack excels at running AI. And the CUDA stack excels at running AI. And the CUDA stack excels at running AI. 151 tokens per second here for the token 151 tokens per second here for the token 151 tokens per second here for the token generation on the 5080, 146 on the 5090. generation on the 5080, 146 on the 5090. generation on the 5080, 146 on the 5090. Interesting. Now, if you're chatting Interesting. Now, if you're chatting Interesting. Now, if you're chatting here, you read about five words per here, you read about five words per here, you read about five words per second. So, this is about like uh 30 second. So, this is about like uh 30 second. So, this is about like uh 30 times faster than you can read, not a times faster than you can read, not a times faster than you can read, not a problem. Now, besides Ollama, another problem. Now, besides Ollama, another problem. Now, besides Ollama, another really popular tool is LM Studio, which really popular tool is LM Studio, which really popular tool is LM Studio, which I've shown on the channel many times I've shown on the channel many times I've shown on the channel many times before. This model, Gemma 4 12B, also a before. This model, Gemma 4 12B, also a before. This model, Gemma 4 12B, also a 4-bit quantization, is 7 GB on disk. 4-bit quantization, is 7 GB on disk. 4-bit quantization, is 7 GB on disk. We're going a little bit larger now. We're going a little bit larger now. We're going a little bit larger now. We're still under that 16 GB threshold. We're still under that 16 GB threshold. We're still under that 16 GB threshold. We don't want to cross that. Now, that We don't want to cross that. Now, that We don't want to cross that. Now, that includes the model size plus the includes the model size plus the includes the model size plus the context. The context is the your entire context. The context is the your entire context. The context is the your entire conversation and anything that's conversation and anything that's conversation and anything that's included in it. For example, if you're included in it. For example, if you're included in it. For example, if you're using this as a coding assistant, you're using this as a coding assistant, you're using this as a coding assistant, you're going to be sending it along with your going to be sending it along with your going to be sending it along with your prompt, you're going to be sending it prompt, you're going to be sending it prompt, you're going to be sending it the code base or some of the code base, the code base or some of the code base, the code base or some of the code base, whatever your agent decides is whatever your agent decides is whatever your agent decides is appropriate. But for now, we're just appropriate. But for now, we're just appropriate. But for now, we're just chatting. And for the kind of chat that chatting. And for the kind of chat that chatting. And for the kind of chat that I'm having with it, I don't need a huge I'm having with it, I don't need a huge I'm having with it, I don't need a huge context. I'm just looking to see how context. I'm just looking to see how context. I'm just looking to see how fast it goes. So, we're using 16 GB out fast it goes. So, we're using 16 GB out fast it goes. So, we're using 16 GB out of 32 available on the system and of 32 available on the system and of 32 available on the system and initially we load into the system RAM
-
initially we load into the system RAM initially we load into the system RAM and then we go into the GPU right here. and then we go into the GPU right here. and then we go into the GPU right here. So, here's the 5090 and there's the So, here's the 5090 and there's the So, here's the 5090 and there's the dedicated GPU memory and you can see dedicated GPU memory and you can see dedicated GPU memory and you can see we've loaded about 10 out of the 24 we've loaded about 10 out of the 24 we've loaded about 10 out of the 24 available. Similar story on the 5080. available. Similar story on the 5080. available. Similar story on the 5080. Here's the GPU and we are loading about Here's the GPU and we are loading about Here's the GPU and we are loading about 9.1 out of 16 that's available there. 9.1 out of 16 that's available there. 9.1 out of 16 that's available there. And go. This is a thinking model, so And go. This is a thinking model, so And go. This is a thinking model, so initially it's going to do a little initially it's going to do a little initially it's going to do a little thinking and then it's going to spit out thinking and then it's going to spit out thinking and then it's going to spit out the text for me. Funny that both of the text for me. Funny that both of the text for me. Funny that both of these stories are about some kind of these stories are about some kind of these stories are about some kind of shop, but they're clearly different shop, but they're clearly different shop, but they're clearly different stories here. So, we've got 61 tokens stories here. So, we've got 61 tokens stories here. So, we've got 61 tokens per second here on the 5080 and 64 per second here on the 5080 and 64 per second here on the 5080 and 64 tokens per second on the 5090. Not a tokens per second on the 5090. Not a tokens per second on the 5090. Not a huge difference, but still pretty huge difference, but still pretty huge difference, but still pretty ridiculous speed cuz if you're chatting, ridiculous speed cuz if you're chatting, ridiculous speed cuz if you're chatting, that's more than you need. Now, by the that's more than you need. Now, by the that's more than you need. Now, by the way, if you're familiar with O llama and way, if you're familiar with O llama and way, if you're familiar with O llama and LM Studio, they run on an engine called LM Studio, they run on an engine called LM Studio, they run on an engine called llama.cpp, which is a command line tool llama.cpp, which is a command line tool llama.cpp, which is a command line tool that you can run and use it as a server. that you can run and use it as a server. that you can run and use it as a server. You can chat with it also. So, I ran a You can chat with it also. So, I ran a You can chat with it also. So, I ran a bunch of models on pure llama.cpp, so bunch of models on pure llama.cpp, so bunch of models on pure llama.cpp, so there's no performance implications. And there's no performance implications. And there's no performance implications. And I got some scores here. I'm looking over I got some scores here. I'm looking over I got some scores here. I'm looking over here because that's where I have the here because that's where I have the here because that's where I have the scores.
-
scores. scores. Qwen 3.5 4B, Gemma 4 12B, we already Qwen 3.5 4B, Gemma 4 12B, we already Qwen 3.5 4B, Gemma 4 12B, we already seen these. Slightly different numbers, seen these. Slightly different numbers, seen these. Slightly different numbers, but within the margin of error. And you but within the margin of error. And you but within the margin of error. And you can see the difference between RTX 5080 can see the difference between RTX 5080 can see the difference between RTX 5080 and 5090. Not huge. GPT-OSS 20B is also and 5090. Not huge. GPT-OSS 20B is also and 5090. Not huge. GPT-OSS 20B is also one of those popular models, even though one of those popular models, even though one of those popular models, even though it's already getting pretty long in the it's already getting pretty long in the it's already getting pretty long in the tooth, but it's fast. It's a mixture of tooth, but it's fast. It's a mixture of tooth, but it's fast. It's a mixture of experts model. The other two on this experts model. The other two on this experts model. The other two on this chart are dense models, which means all chart are dense models, which means all chart are dense models, which means all the parameters run. So, the 12 billion the parameters run. So, the 12 billion the parameters run. So, the 12 billion parameter model, 12 billion parameters parameter model, 12 billion parameters parameter model, 12 billion parameters are active. 4 billion parameter model, 4 are active. 4 billion parameter model, 4 are active. 4 billion parameter model, 4 billion is active. But in 20B, I think billion is active. But in 20B, I think billion is active. But in 20B, I think three or four are active. I'm sure three or four are active. I'm sure three or four are active. I'm sure somebody in the comments will correct me somebody in the comments will correct me somebody in the comments will correct me there, but not all the parameters are there, but not all the parameters are there, but not all the parameters are active, and that's why it's so much active, and that's why it's so much active, and that's why it's so much faster. Now, that's token generation faster. Now, that's token generation faster. Now, that's token generation speed from scratch. All those fit in 16 speed from scratch. All those fit in 16 speed from scratch. All those fit in 16 GB, easy. So, extra VRAM isn't buying a GB, easy. So, extra VRAM isn't buying a GB, easy. So, extra VRAM isn't buying a speed, it's just that the 5090 has more speed, it's just that the 5090 has more speed, it's just that the 5090 has more CUDA cores than on the 5080. So, it has CUDA cores than on the 5080. So, it has CUDA cores than on the 5080. So, it has more cores on the chip on that GPU. The more cores on the chip on that GPU. The more cores on the chip on that GPU. The numbers are not that far off, and that's numbers are not that far off, and that's numbers are not that far off, and that's because I gave it really short prompts, because I gave it really short prompts, because I gave it really short prompts, so there's not much computation so there's not much computation so there's not much computation happening there on the GPU. Now, what happening there on the GPU. Now, what happening there on the GPU. Now, what happens after you've had a little bit of happens after you've had a little bit of happens after you've had a little bit of a conversation and now you have a lot of a conversation and now you have a lot of a conversation and now you have a lot of context built up. Maybe not a lot, but I context built up. Maybe not a lot, but I context built up. Maybe not a lot, but I also ran this with an 8K context. That's also ran this with an 8K context. That's also ran this with an 8K context. That's 8,000 tokens of history already in the 8,000 tokens of history already in the 8,000 tokens of history already in the chat that gets repeated back into the chat that gets repeated back into the chat that gets repeated back into the GPU, sent back to the GPU as your GPU, sent back to the GPU as your GPU, sent back to the GPU as your context, so it has to process more.
-
context, so it has to process more. context, so it has to process more. Well, it slows down just a little bit, Well, it slows down just a little bit, Well, it slows down just a little bit, but basically we have pretty much the but basically we have pretty much the but basically we have pretty much the same pattern. There's one more very same pattern. There's one more very same pattern. There's one more very important number to consider, and that's important number to consider, and that's important number to consider, and that's the prompt processing that happens the prompt processing that happens the prompt processing that happens actually before the generation. So, I actually before the generation. So, I actually before the generation. So, I maybe I should have done that first. maybe I should have done that first. maybe I should have done that first. That's when you send the input and it That's when you send the input and it That's when you send the input and it gets processed by the GPU. That's the gets processed by the GPU. That's the gets processed by the GPU. That's the computation. And this is much, much computation. And this is much, much computation. And this is much, much faster. We're talking about 5,104 faster. We're talking about 5,104 faster. We're talking about 5,104 tokens per second for Qwen 4B and up to tokens per second for Qwen 4B and up to tokens per second for Qwen 4B and up to 6,983 6,983 6,983 on the RTX 5090 for GPT-OSS 20B. Now, on the RTX 5090 for GPT-OSS 20B. Now, on the RTX 5090 for GPT-OSS 20B. Now, here's a thing you actually need to here's a thing you actually need to here's a thing you actually need to understand about local AI, and it's one understand about local AI, and it's one understand about local AI, and it's one of the reasons I have these two laptops of the reasons I have these two laptops of the reasons I have these two laptops here. The difference in VRAM. VRAM here. The difference in VRAM. VRAM here. The difference in VRAM. VRAM decides which models you can run at all. decides which models you can run at all. decides which models you can run at all. You can think of it like doors. 16 gigs You can think of it like doors. 16 gigs You can think of it like doors. 16 gigs opens every door we just walked through opens every door we just walked through opens every door we just walked through so far. 24 gigs will open up a few more so far. 24 gigs will open up a few more so far. 24 gigs will open up a few more doors at the top. So, everyday chat, doors at the top. So, everyday chat, doors at the top. So, everyday chat, brainstorming, the stuff most people brainstorming, the stuff most people brainstorming, the stuff most people will pay 20 bucks a month for, you know will pay 20 bucks a month for, you know will pay 20 bucks a month for, you know I'm talking about? Yeah, that's covered. I'm talking about? Yeah, that's covered. I'm talking about? Yeah, that's covered. And local never leaves the machine, but And local never leaves the machine, but And local never leaves the machine, but chat is easy. Can you actually use this chat is easy. Can you actually use this chat is easy. Can you actually use this for work? I jumped right into an oldie for work? I jumped right into an oldie for work? I jumped right into an oldie model, but a goodie model, because this model, but a goodie model, because this model, but a goodie model, because this model is very good for coding, or at model is very good for coding, or at model is very good for coding, or at least it used to be. But it works very least it used to be. But it works very least it used to be. But it works very well with code editors. Like here I've well with code editors. Like here I've well with code editors. Like here I've got VS Code running, and I'm using the got VS Code running, and I'm using the got VS Code running, and I'm using the Continue plugin, and it's Qwen 3 Coder Continue plugin, and it's Qwen 3 Coder Continue plugin, and it's Qwen 3 Coder 30B running locally in Llama.cpp. Let's 30B running locally in Llama.cpp. Let's 30B running locally in Llama.cpp. Let's actually pop open the task manager so actually pop open the task manager so actually pop open the task manager so you can see what's going on. And there you can see what's going on. And there you can see what's going on. And there is that dedicated GPU. The memory is is that dedicated GPU. The memory is is that dedicated GPU. The memory is 23.3 GB out of 24 with 131,000 context 23.3 GB out of 24 with 131,000 context 23.3 GB out of 24 with 131,000 context window. So, yeah, we're at that limit. I
-
window. So, yeah, we're at that limit. I window. So, yeah, we're at that limit. I wonder what's going to happen on the 16 wonder what's going to happen on the 16 wonder what's going to happen on the 16 gig machine, but first let's try it gig machine, but first let's try it gig machine, but first let's try it here. You can see that GPU is not active here. You can see that GPU is not active here. You can see that GPU is not active right now. I'm going to select that right now. I'm going to select that right now. I'm going to select that model and just say hi to it. I know, I model and just say hi to it. I know, I model and just say hi to it. I know, I know, you all hate that, but it's it's know, you all hate that, but it's it's know, you all hate that, but it's it's proof of concept, okay? And you can see proof of concept, okay? And you can see proof of concept, okay? And you can see that little spike on the GPU right that little spike on the GPU right that little spike on the GPU right there. That's where it answered me. there. That's where it answered me. there. That's where it answered me. Okay, fine. I'll give it a longer Okay, fine. I'll give it a longer Okay, fine. I'll give it a longer prompt. I'll give this architecture prompt. I'll give this architecture prompt. I'll give this architecture prompt, which I sometimes uh prompt, which I sometimes uh prompt, which I sometimes uh Design a scalable web application Design a scalable web application Design a scalable web application architecture for an e-commerce platform, architecture for an e-commerce platform, architecture for an e-commerce platform, and so on. You can read it. Pause the and so on. You can read it. Pause the and so on. You can read it. Pause the video if you want to. This is a good video if you want to. This is a good video if you want to. This is a good prompt because not only does this have prompt because not only does this have prompt because not only does this have to process the prompt, it's not that to process the prompt, it's not that to process the prompt, it's not that long, okay? But, what this does is long, okay? But, what this does is long, okay? But, what this does is generate a very, very long response for generate a very, very long response for generate a very, very long response for me. And that's pretty much how fast it's me. And that's pretty much how fast it's me. And that's pretty much how fast it's going on this machine. It's a big model, going on this machine. It's a big model, going on this machine. It's a big model, so it's not going to go super fast. But, so it's not going to go super fast. But, so it's not going to go super fast. But, there's that GPU being very, very busy. there's that GPU being very, very busy. there's that GPU being very, very busy. 99 to 100% utilization. Now, Continue's 99 to 100% utilization. Now, Continue's 99 to 100% utilization. Now, Continue's agent also has tool use capabilities agent also has tool use capabilities agent also has tool use capabilities here. So, it's going to create new here. So, it's going to create new here. So, it's going to create new documents, and it's going to put the documents, and it's going to put the documents, and it's going to put the architecture.md architecture.md architecture.md file and everything related to it into file and everything related to it into file and everything related to it into that document instead of spitting it out that document instead of spitting it out that document instead of spitting it out here. Hey, that whatever works. I do here. Hey, that whatever works. I do here. Hey, that whatever works. I do hear the fans spinning up on this thing, hear the fans spinning up on this thing, hear the fans spinning up on this thing, so it's working hard. Let's see what so it's working hard. Let's see what so it's working hard. Let's see what happens over here. I have the exact same happens over here. I have the exact same happens over here. I have the exact same setup here. Select Qwen Code 30B, and setup here. Select Qwen Code 30B, and setup here. Select Qwen Code 30B, and let's do a little "Hi" here.
-
let's do a little "Hi" here. let's do a little "Hi" here. Uh, this one's taking a little bit Uh, this one's taking a little bit Uh, this one's taking a little bit longer to answer. Now, this model is longer to answer. Now, this model is longer to answer. Now, this model is pretty popular, but it's 61 GB full-size pretty popular, but it's 61 GB full-size pretty popular, but it's 61 GB full-size model. And we're running a 4-bit model. And we're running a 4-bit model. And we're running a 4-bit quantization here. Here's an example of quantization here. Here's an example of quantization here. Here's an example of one. The 4-bit quant is 18.6 one. The 4-bit quant is 18.6 one. The 4-bit quant is 18.6 GB in this particular case, but usually GB in this particular case, but usually GB in this particular case, but usually it's going to be around there. So, it's going to be around there. So, it's going to be around there. So, that's just outside of what this GPU has that's just outside of what this GPU has that's just outside of what this GPU has available. I finally got an answer here, available. I finally got an answer here, available. I finally got an answer here, but that took a while. That's because it but that took a while. That's because it but that took a while. That's because it started spilling over to system memory started spilling over to system memory started spilling over to system memory and the CPU in order to run this model. and the CPU in order to run this model. and the CPU in order to run this model. Now, 30B is a bit of an older model. Now, 30B is a bit of an older model. Now, 30B is a bit of an older model. There are some newer models that are There are some newer models that are There are some newer models that are around now that are taking over for the around now that are taking over for the around now that are taking over for the development space, especially in this development space, especially in this development space, especially in this smaller model range, smaller to smaller model range, smaller to smaller model range, smaller to medium-size model range. And we're going medium-size model range. And we're going medium-size model range. And we're going up right against that limit right here up right against that limit right here up right against that limit right here on the 5090. The 5080 can do it, but on the 5090. The 5080 can do it, but on the 5090. The 5080 can do it, but it's having a little bit of a time it's having a little bit of a time it's having a little bit of a time keeping up. So, we've got Qwen 3.5 27B, keeping up. So, we've got Qwen 3.5 27B, keeping up. So, we've got Qwen 3.5 27B, Code 30B, the one we just ran here in Code 30B, the one we just ran here in Code 30B, the one we just ran here in the editor and quant 3.635b.
-
the editor and quant 3.635b. the editor and quant 3.635b. By the way, the 27b, that's a dense By the way, the 27b, that's a dense By the way, the 27b, that's a dense model. All 27 billion parameters, that's model. All 27 billion parameters, that's model. All 27 billion parameters, that's the largest densest model that I've run. the largest densest model that I've run. the largest densest model that I've run. We're getting nine tokens per second on We're getting nine tokens per second on We're getting nine tokens per second on the 5080. The other models are mixture the 5080. The other models are mixture the 5080. The other models are mixture of experts models. And let's take a look of experts models. And let's take a look of experts models. And let's take a look at the 5090. You can see at the 5090. You can see at the 5090. You can see quite a difference there. More than quite a difference there. More than quite a difference there. More than double, even more than triple in certain double, even more than triple in certain double, even more than triple in certain cases. That 27 billion parameter model cases. That 27 billion parameter model cases. That 27 billion parameter model definitely has some CPU and system definitely has some CPU and system definitely has some CPU and system memory spillover, so it slows down quite memory spillover, so it slows down quite memory spillover, so it slows down quite a bit. At that point, you're looking at a bit. At that point, you're looking at a bit. At that point, you're looking at the 5090 machine, uh nine tokens a the 5090 machine, uh nine tokens a the 5090 machine, uh nine tokens a second is pretty unusable at that point. second is pretty unusable at that point. second is pretty unusable at that point. Although coder 30b still looking pretty Although coder 30b still looking pretty Although coder 30b still looking pretty decent on both of these machines, decent on both of these machines, decent on both of these machines, although it's going to be two times although it's going to be two times although it's going to be two times faster on the 5090 machine. And finally, faster on the 5090 machine. And finally, faster on the 5090 machine. And finally, the mixture of experts 35 billion the mixture of experts 35 billion the mixture of experts 35 billion parameter model quant 3.6. Yeah, same parameter model quant 3.6. Yeah, same parameter model quant 3.6. Yeah, same story there. Same goes for the 8k story there. Same goes for the 8k story there. Same goes for the 8k context. We're down to eight tokens per context. We're down to eight tokens per context. We're down to eight tokens per second here, which is actually not a second here, which is actually not a second here, which is actually not a terrible drop there. terrible drop there. terrible drop there. All these bigger models have slower All these bigger models have slower All these bigger models have slower prompt processing than GPT-OSS or the prompt processing than GPT-OSS or the prompt processing than GPT-OSS or the small model because there's just so much small model because there's just so much small model because there's just so much to chug through. There's always this to chug through. There's always this to chug through. There's always this balancing act that you have to work balancing act that you have to work balancing act that you have to work with. The larger models will obviously with. The larger models will obviously with. The larger models will obviously be heavier and they'll take up more be heavier and they'll take up more be heavier and they'll take up more memory.
-
memory. memory. And you want to use the larger models And you want to use the larger models And you want to use the larger models cuz they're smarter with uh you know, cuz they're smarter with uh you know, cuz they're smarter with uh you know, with your code editors, but you have to with your code editors, but you have to with your code editors, but you have to balance it out. That's why I still do balance it out. That's why I still do balance it out. That's why I still do coder 30b because it looks pretty nice coder 30b because it looks pretty nice coder 30b because it looks pretty nice here, doesn't it? Here I've got flux, a here, doesn't it? Here I've got flux, a here, doesn't it? Here I've got flux, a 12 billion parameter image model. It's 12 billion parameter image model. It's 12 billion parameter image model. It's running inside comfy UI, everybody's running inside comfy UI, everybody's running inside comfy UI, everybody's favorite image and video generation favorite image and video generation favorite image and video generation tool. tool. tool. It's powerful, okay? And go. This one's It's powerful, okay? And go. This one's It's powerful, okay? And go. This one's pretty quick. Video generation is pretty quick. Video generation is pretty quick. Video generation is another story. Oh, well, this is all another story. Oh, well, this is all another story. Oh, well, this is all local. Did I mention that? There's no local. Did I mention that? There's no local. Did I mention that? There's no like you only have two generations like you only have two generations like you only have two generations remaining messages or errors. No sweat. remaining messages or errors. No sweat. remaining messages or errors. No sweat. You can do this as many times as you You can do this as many times as you You can do this as many times as you want. Done. It's about 45 seconds or so want. Done. It's about 45 seconds or so want. Done. It's about 45 seconds or so to make this image. Now, images are to make this image. Now, images are to make this image. Now, images are pretty easy thing. What about video pretty easy thing. What about video pretty easy thing. What about video generation? Here's one 2.2 and I took generation? Here's one 2.2 and I took generation? Here's one 2.2 and I took that same image that got generated and that same image that got generated and that same image that got generated and plugged it in here. So, we're doing plugged it in here. So, we're doing plugged it in here. So, we're doing image to video. And you've kind of image to video. And you've kind of image to video. And you've kind of already seen a preview for this because already seen a preview for this because already seen a preview for this because you saw the beginning of the video, you saw the beginning of the video, you saw the beginning of the video, maybe. Well, that was one example. maybe. Well, that was one example. maybe. Well, that was one example. Here's another one. It's a game sim Here's another one. It's a game sim Here's another one. It's a game sim racing and something happened uh to that racing and something happened uh to that racing and something happened uh to that speedometer after the first frame. Don't speedometer after the first frame. Don't speedometer after the first frame. Don't ask me what, but that's how things are ask me what, but that's how things are ask me what, but that's how things are these days. It's not perfect. This took these days. It's not perfect. This took these days. It's not perfect. This took 3 and 1/2 minutes, 3 seconds, 81 frames.
-
3 and 1/2 minutes, 3 seconds, 81 frames. 3 and 1/2 minutes, 3 seconds, 81 frames. That speedometer though. That speedometer though. That speedometer though. What happened to What happened to What happened to kilometers per hour? Is it really that kilometers per hour? Is it really that kilometers per hour? Is it really that hard to spell out kilometers per hour? hard to spell out kilometers per hour? hard to spell out kilometers per hour? Come on. Love it. And by the way, the Come on. Love it. And by the way, the Come on. Love it. And by the way, the smaller twin can render all the exact smaller twin can render all the exact smaller twin can render all the exact same stuff, but 4 minutes instead of 3 same stuff, but 4 minutes instead of 3 same stuff, but 4 minutes instead of 3 and 1/2. That's the whole difference. and 1/2. That's the whole difference. and 1/2. That's the whole difference. Not huge. So, how are gaming laptops Not huge. So, how are gaming laptops Not huge. So, how are gaming laptops able to do all this? Well, it's the able to do all this? Well, it's the able to do all this? Well, it's the Nvidia GPUs inside with their tensor Nvidia GPUs inside with their tensor Nvidia GPUs inside with their tensor cores, dedicated GDDR7 VRAM, the Intel cores, dedicated GDDR7 VRAM, the Intel cores, dedicated GDDR7 VRAM, the Intel Core Ultra 9 inside. That's usually you Core Ultra 9 inside. That's usually you Core Ultra 9 inside. That's usually you have your agents working away. The have your agents working away. The have your agents working away. The agents do most of their work on the CPU, agents do most of their work on the CPU, agents do most of their work on the CPU, handing things off to the GPU to the handing things off to the GPU to the handing things off to the GPU to the models. Who would have thought that models. Who would have thought that models. Who would have thought that gaming hardware like this would have gaming hardware like this would have gaming hardware like this would have been so perfectly suited for AI been so perfectly suited for AI been so perfectly suited for AI in today's world. So, not only are you in today's world. So, not only are you in today's world. So, not only are you getting a good gaming machine if you get getting a good gaming machine if you get getting a good gaming machine if you get one of these, but you're also getting a one of these, but you're also getting a one of these, but you're also getting a top-of-the-line Nvidia CUDA stack GPU top-of-the-line Nvidia CUDA stack GPU top-of-the-line Nvidia CUDA stack GPU inside of one of these. And if you have inside of one of these. And if you have inside of one of these. And if you have a requirement for running local video a requirement for running local video a requirement for running local video and image generation or running local and image generation or running local and image generation or running local models, your code doesn't leave your models, your code doesn't leave your models, your code doesn't leave your system. It doesn't get any better than system. It doesn't get any better than system. It doesn't get any better than this because there is no RTX Pro 6000 this because there is no RTX Pro 6000 this because there is no RTX Pro 6000 laptop. Well, not yet. And I don't think laptop. Well, not yet. And I don't think laptop. Well, not yet. And I don't think there will be. So, this is kind of like there will be. So, this is kind of like there will be. So, this is kind of like the top-of-the-line. And of course, the top-of-the-line. And of course, the top-of-the-line. And of course, you're on the CUDA stack. Best in class.
-
you're on the CUDA stack. Best in class. you're on the CUDA stack. Best in class. Anyway, it's kind of a pleasure to run Anyway, it's kind of a pleasure to run Anyway, it's kind of a pleasure to run laptops that are really solidly built laptops that are really solidly built laptops that are really solidly built and present absolutely zero hiccups and present absolutely zero hiccups and present absolutely zero hiccups while being used. I like that. I'm used while being used. I like that. I'm used while being used. I like that. I'm used to that with my MacBook Pro. to that with my MacBook Pro. to that with my MacBook Pro. >> [laughter] >> [laughter] >> [laughter] >> And I really appreciate being able to >> And I really appreciate being able to >> And I really appreciate being able to run the Nvidia hardware in machines like run the Nvidia hardware in machines like run the Nvidia hardware in machines like this. If you enjoyed this video, you're this. If you enjoyed this video, you're this. If you enjoyed this video, you're probably going to like this one next. probably going to like this one next. probably going to like this one next. Thanks for watching and [music] I'll see Thanks for watching and [music] I'll see Thanks for watching and [music] I'll see you next time.
Summary
The main theme is the powerful capabilities of MSI Raider 16 Max HX gaming laptops for running AI models locally, particularly the impact of VRAM. Key subjects include the Nvidia RTX 5080 and 5090 GPUs, Intel Core Ultra 9 processors, and the significance of ample VRAM for AI tasks. The practical takeaway is that high VRAM is the crucial factor determining local AI performance on these machines, making them suitable for demanding AI workloads beyond gaming.