← Back
Aaron Zisk August 26, 2026 13m

I Plugged RTX Pro 6000 Into a Strix Halo... Didn’t Expect This

Read full transcript 11 segments
  1. 224 224 GB. That's how much graphics memory is GB. That's how much graphics memory is GB. That's how much graphics memory is sitting right here on my desk in a tiny sitting right here on my desk in a tiny sitting right here on my desk in a tiny little package. And these two are little package. And these two are little package. And these two are working together to run a single 122 working together to run a single 122 working together to run a single 122 billion parameter model. But that's not billion parameter model. But that's not billion parameter model. But that's not even the crazy thing. The crazy thing is even the crazy thing. The crazy thing is even the crazy thing. The crazy thing is that this is an Nvidia GPU and it's that this is an Nvidia GPU and it's that this is an Nvidia GPU and it's plugged into a mini PC with an AMD Strix plugged into a mini PC with an AMD Strix plugged into a mini PC with an AMD Strix Halo. AMD and Nvidia forced to do the Halo. AMD and Nvidia forced to do the Halo. AMD and Nvidia forced to do the one thing they aren't meant to, run the one thing they aren't meant to, run the one thing they aren't meant to, run the same AI model at the same time. Boom. same AI model at the same time. Boom. same AI model at the same time. Boom. Boom. Boom. Boom. Look at that. They're both working. I Look at that. They're both working. I Look at that. They're both working. I can hear it. 33, 34 tokens per second can hear it. 33, 34 tokens per second can hear it. 33, 34 tokens per second from both of them combined. There's the from both of them combined. There's the from both of them combined. There's the split and I'll get into this in a bit. split and I'll get into this in a bit. split and I'll get into this in a bit. But a little backstory first. Last year But a little backstory first. Last year But a little backstory first. Last year GMK Tech dropped the Evo X2 and it was GMK Tech dropped the Evo X2 and it was GMK Tech dropped the Evo X2 and it was one of the first, if not the first mini one of the first, if not the first mini one of the first, if not the first mini PC on the planet to run AMD Strix Halo, PC on the planet to run AMD Strix Halo, PC on the planet to run AMD Strix Halo, which for the uninitiated is a APU. which for the uninitiated is a APU. which for the uninitiated is a APU. APU's GPU and CPU together and they APU's GPU and CPU together and they APU's GPU and CPU together and they share a common pool of memory that's up share a common pool of memory that's up share a common pool of memory that's up to 128 GB, most of which can be to 128 GB, most of which can be to 128 GB, most of which can be allocated to the GPU to be able to run allocated to the GPU to be able to run allocated to the GPU to be able to run LLMs, image generation, things like LLMs, image generation, things like LLMs, image generation, things like that. They beat everybody to it, even that. They beat everybody to it, even that. They beat everybody to it, even Nvidia's DGX Spark. And now GMK Tech has Nvidia's DGX Spark. And now GMK Tech has Nvidia's DGX Spark. And now GMK Tech has done it again. First to market with a done it again. First to market with a done it again. First to market with a brand new trick for the Strix Halo brand new trick for the Strix Halo brand new trick for the Strix Halo machines. It's a door in the back that machines. It's a door in the back that machines. It's a door in the back that changes everything. All right, all changes everything. All right, all changes everything. All right, all right, I'll stop being all dramatic.

  2. right, I'll stop being all dramatic. right, I'll stop being all dramatic. It's an OCuLink port, okay? Nobody else It's an OCuLink port, okay? Nobody else It's an OCuLink port, okay? Nobody else has it, but it opens up so many has it, but it opens up so many has it, but it opens up so many possibilities. Basically, it's a PCIe possibilities. Basically, it's a PCIe possibilities. Basically, it's a PCIe slot that has an external port. 63 slot that has an external port. 63 slot that has an external port. 63 gigabit per second speed. It's the same gigabit per second speed. It's the same gigabit per second speed. It's the same kind of slot that a graphics card plugs kind of slot that a graphics card plugs kind of slot that a graphics card plugs into inside a full-size desktop, except into inside a full-size desktop, except into inside a full-size desktop, except it's on a mini PC. Which means for the it's on a mini PC. Which means for the it's on a mini PC. Which means for the first time we can take a real desktop first time we can take a real desktop first time we can take a real desktop GPU and plug it into a Strix Halo GPU and plug it into a Strix Halo GPU and plug it into a Strix Halo machine. So I'm going to start out with machine. So I'm going to start out with machine. So I'm going to start out with something small and then get into something small and then get into something small and then get into something ridiculous. But before I go something ridiculous. But before I go something ridiculous. But before I go and bolt a $10,000 GPU into a mini PC, I and bolt a $10,000 GPU into a mini PC, I and bolt a $10,000 GPU into a mini PC, I want to check something first. There you want to check something first. There you want to check something first. There you go, little card. Sit and wait for me go, little card. Sit and wait for me go, little card. Sit and wait for me right there. The other day I was working right there. The other day I was working right there. The other day I was working from a hotel and the Wi-Fi was open to from a hotel and the Wi-Fi was open to from a hotel and the Wi-Fi was open to basically anyone in the building. That's basically anyone in the building. That's basically anyone in the building. That's great for convenience, but not where I great for convenience, but not where I great for convenience, but not where I want my traffic floating around. That's want my traffic floating around. That's want my traffic floating around. That's why I connect to Surfshark before I get why I connect to Surfshark before I get why I connect to Surfshark before I get to work. My Git commits, container to work. My Git commits, container to work. My Git commits, container downloads, and terminal sessions are all downloads, and terminal sessions are all downloads, and terminal sessions are all sent through an AES-256 encrypted sent through an AES-256 encrypted sent through an AES-256 encrypted tunnel. CleanWeb handles the annoying tunnel. CleanWeb handles the annoying tunnel. CleanWeb handles the annoying stuff, too, including ads and trackers stuff, too, including ads and trackers stuff, too, including ads and trackers that clutter up websites. They have an that clutter up websites. They have an that clutter up websites. They have an independently audited no-logs policy, independently audited no-logs policy, independently audited no-logs policy, plus RAM-only servers that wipe plus RAM-only servers that wipe plus RAM-only servers that wipe themselves on reboot. My favorite themselves on reboot. My favorite themselves on reboot. My favorite feature is there's still no device feature is there's still no device feature is there's still no device limit. One plan covers my MacBook, my limit. One plan covers my MacBook, my limit. One plan covers my MacBook, my phone, and basically any other device I phone, and basically any other device I phone, and basically any other device I managed to collect. I feel much better managed to collect. I feel much better managed to collect. I feel much better connecting back to my lab when I'm connecting back to my lab when I'm connecting back to my lab when I'm working away from the office. Use working away from the office. Use working away from the office. Use surfshark.com/alexis surfshark.com/alexis surfshark.com/alexis king to get four extra months for free, king to get four extra months for free, king to get four extra months for free, and you'll also have 30 days to try it and you'll also have 30 days to try it and you'll also have 30 days to try it risk-free. There's a link down below.

  3. risk-free. There's a link down below. risk-free. There's a link down below. Protect your connection, keep building, Protect your connection, keep building, Protect your connection, keep building, and now back to the video. and now back to the video. and now back to the video. And what I want to check is if all the And what I want to check is if all the And what I want to check is if all the stuff I tested last year with the X2 stuff I tested last year with the X2 stuff I tested last year with the X2 still holds up on the X3. still holds up on the X3. still holds up on the X3. So, I grabbed the model I already tested So, I grabbed the model I already tested So, I grabbed the model I already tested on the X2 last year. It's older, yeah, on the X2 last year. It's older, yeah, on the X2 last year. It's older, yeah, but I wanted to make sure. Qwen 2.5 32 but I wanted to make sure. Qwen 2.5 32 but I wanted to make sure. Qwen 2.5 32 billion parameters. On the old X2, it billion parameters. On the old X2, it billion parameters. On the old X2, it ran about 11 tokens a second. On the X3, ran about 11 tokens a second. On the X3, ran about 11 tokens a second. On the X3, 11.2. Same chip, same speed, and that's 11.2. Same chip, same speed, and that's 11.2. Same chip, same speed, and that's a good thing, which makes the rest of a good thing, which makes the rest of a good thing, which makes the rest of the testing reliable. One quick note for the testing reliable. One quick note for the testing reliable. One quick note for the nerds, all my tests are going to be the nerds, all my tests are going to be the nerds, all my tests are going to be using llama.cpp and not ollama, even using llama.cpp and not ollama, even using llama.cpp and not ollama, even though I did test ollama last year. though I did test ollama last year. though I did test ollama last year. Ollama's convenient, but on this Ollama's convenient, but on this Ollama's convenient, but on this hardware, it's slower. The numbers check hardware, it's slower. The numbers check hardware, it's slower. The numbers check out. Let's get to the fun part. out. Let's get to the fun part. out. Let's get to the fun part. >> [music] >> [music] >> [music] >> This is the RTX 5080, 16 gigabytes of >> This is the RTX 5080, 16 gigabytes of >> This is the RTX 5080, 16 gigabytes of very fast memory. Plug it into that very fast memory. Plug it into that very fast memory. Plug it into that OCuLink port, and on a model that fits, OCuLink port, and on a model that fits, OCuLink port, and on a model that fits, it's about three times faster than the it's about three times faster than the it's about three times faster than the built-in GPU. And I say on the model built-in GPU. And I say on the model built-in GPU. And I say on the model that fits because 16 gigabytes is not that fits because 16 gigabytes is not that fits because 16 gigabytes is not going to fit that many models, not like going to fit that many models, not like going to fit that many models, not like the 128 gigs that can fit big models the 128 gigs that can fit big models the 128 gigs that can fit big models inside here. But we'll get to that. The inside here. But we'll get to that. The inside here. But we'll get to that. The moment the model gets too big, yeah, moment the model gets too big, yeah, moment the model gets too big, yeah, look at that. We're using all the look at that. We're using all the look at that. We're using all the memory. We're using 100% of the GPU, and memory. We're using 100% of the GPU, and memory. We're using 100% of the GPU, and nothing is being spit out at all, but it nothing is being spit out at all, but it nothing is being spit out at all, but it is being spit out. It's just very, very is being spit out. It's just very, very is being spit out. It's just very, very slow. NVTop is showing that we are using slow. NVTop is showing that we are using slow. NVTop is showing that we are using the GPU to the fullest here. Look at the GPU to the fullest here. Look at the GPU to the fullest here. Look at that. We have a couple of tokens, ladies that. We have a couple of tokens, ladies that. We have a couple of tokens, ladies and gentlemen. I'm not going to make you and gentlemen. I'm not going to make you and gentlemen. I'm not going to make you sit through it because it's about 1.6 sit through it because it's about 1.6 sit through it because it's about 1.6 tokens per second. The 5080 falls off a tokens per second. The 5080 falls off a tokens per second. The 5080 falls off a cliff here. We're at 11.2 tokens per

  4. cliff here. We're at 11.2 tokens per cliff here. We're at 11.2 tokens per second on the 8060 that's inside the second on the 8060 that's inside the second on the 8060 that's inside the Strix Halo. So, the eGPU wins big right Strix Halo. So, the eGPU wins big right Strix Halo. So, the eGPU wins big right until the model doesn't fit. Oh, and until the model doesn't fit. Oh, and until the model doesn't fit. Oh, and just a little aside here, this comes just a little aside here, this comes just a little aside here, this comes with Windows and it works fine. It with Windows and it works fine. It with Windows and it works fine. It actually runs both GPUs exactly the same actually runs both GPUs exactly the same actually runs both GPUs exactly the same way that it does on Linux. I installed way that it does on Linux. I installed way that it does on Linux. I installed Linux right away on there in a dual boot Linux right away on there in a dual boot Linux right away on there in a dual boot scenario, but I did do this test first. scenario, but I did do this test first. scenario, but I did do this test first. And start. And this is doing both at the And start. And this is doing both at the And start. And this is doing both at the same time. This is using my scalable web same time. This is using my scalable web same time. This is using my scalable web architecture prompt. architecture prompt. architecture prompt. It's actually doing it on both GPUs. It's actually doing it on both GPUs. It's actually doing it on both GPUs. This is what it looks like in NVTop. You This is what it looks like in NVTop. You This is what it looks like in NVTop. You can see that spike right there on the can see that spike right there on the can see that spike right there on the first GPU, that's the Nvidia one. And first GPU, that's the Nvidia one. And first GPU, that's the Nvidia one. And here it is, the same spike or a little here it is, the same spike or a little here it is, the same spike or a little bit longer on the AMD GPU. And overall, bit longer on the AMD GPU. And overall, bit longer on the AMD GPU. And overall, we got 87 tokens per second on the we got 87 tokens per second on the we got 87 tokens per second on the Nvidia, 24 tokens per second on the AMD. Nvidia, 24 tokens per second on the AMD. Nvidia, 24 tokens per second on the AMD. Now, this is only for one agent. What Now, this is only for one agent. What Now, this is only for one agent. What happens if we do like 10 agent or chat, happens if we do like 10 agent or chat, happens if we do like 10 agent or chat, whatever you want to call it. Let's do whatever you want to call it. Let's do whatever you want to call it. Let's do 10 and print a thousand tokens out. So, 10 and print a thousand tokens out. So, 10 and print a thousand tokens out. So, start. This will considerably slow down start. This will considerably slow down start. This will considerably slow down each one, but it'll show what these GPUs each one, but it'll show what these GPUs each one, but it'll show what these GPUs are capable of on multi-throughput are capable of on multi-throughput are capable of on multi-throughput scenarios. Definitely a lot more work scenarios. Definitely a lot more work scenarios. Definitely a lot more work being done. You can tell here in NVTop.

  5. being done. You can tell here in NVTop. being done. You can tell here in NVTop. Oh, yeah. It's spinning up, making a lot Oh, yeah. It's spinning up, making a lot Oh, yeah. It's spinning up, making a lot more noise now. It's not that much more noise now. It's not that much more noise now. It's not that much noise. But, you can see we're almost noise. But, you can see we're almost noise. But, you can see we're almost 10,000 tokens there on this one, which 10,000 tokens there on this one, which 10,000 tokens there on this one, which it should be around 10,000 at the end it should be around 10,000 at the end it should be around 10,000 at the end when we're done here. Nvidia one is when we're done here. Nvidia one is when we're done here. Nvidia one is done. AMD is still running. And here's done. AMD is still running. And here's done. AMD is still running. And here's the breakdown now. 70 tokens per second the breakdown now. 70 tokens per second the breakdown now. 70 tokens per second for this one, 70 for this one, 72, 71, for this one, 70 for this one, 72, 71, for this one, 70 for this one, 72, 71, 71, 79, 69. So, the Nvidia one is 71, 79, 69. So, the Nvidia one is 71, 79, 69. So, the Nvidia one is handling 10 agents pretty much no handling 10 agents pretty much no handling 10 agents pretty much no problem here. Completely done. 21 watts problem here. Completely done. 21 watts problem here. Completely done. 21 watts was used. What's going on with AMD? was used. What's going on with AMD? was used. What's going on with AMD? Still running. Still running. Still running. Are we done with all the agents? Nope. Are we done with all the agents? Nope. Are we done with all the agents? Nope. Some of them are just starting out now. Some of them are just starting out now. Some of them are just starting out now. And we're about 22 21 tokens a second And we're about 22 21 tokens a second And we're about 22 21 tokens a second there. This kind of shows you how it there. This kind of shows you how it there. This kind of shows you how it scales out to multiple agents if you scales out to multiple agents if you scales out to multiple agents if you were to be interested in doing that, were to be interested in doing that, were to be interested in doing that, which a lot of you are. So, the 5080 which a lot of you are. So, the 5080 which a lot of you are. So, the 5080 works, and that's cool, but 16 gigs is works, and that's cool, but 16 gigs is works, and that's cool, but 16 gigs is kind of a warm-up. I wanted to know what kind of a warm-up. I wanted to know what kind of a warm-up. I wanted to know what happens if we strap on the biggest GPU happens if we strap on the biggest GPU happens if we strap on the biggest GPU that Nvidia makes. Well, they make that Nvidia makes. Well, they make that Nvidia makes. Well, they make bigger GPUs, but this is the one that'll bigger GPUs, but this is the one that'll bigger GPUs, but this is the one that'll actually sell you for 10 to 11 to 14,000 actually sell you for 10 to 11 to 14,000 actually sell you for 10 to 11 to 14,000 dollars.

  6. That's 96 gigabytes of VRAM on a single That's 96 gigabytes of VRAM on a single GPU. That 32 billion model that killed GPU. That 32 billion model that killed GPU. That 32 billion model that killed the 5080, well, the Pro 6000 does it at the 5080, well, the Pro 6000 does it at the 5080, well, the Pro 6000 does it at 69 tokens a second using almost 600 69 tokens a second using almost 600 69 tokens a second using almost 600 watts. That's almost a full potential watts. That's almost a full potential watts. That's almost a full potential that it can go. 19.9 out of 96 gigabytes that it can go. 19.9 out of 96 gigabytes that it can go. 19.9 out of 96 gigabytes used. And that is six times faster than used. And that is six times faster than used. And that is six times faster than the built-in chip. But here's the fun the built-in chip. But here's the fun the built-in chip. But here's the fun part. Look at that. [laughter] I loaded part. Look at that. [laughter] I loaded part. Look at that. [laughter] I loaded the 122 billion parameter model Q4 the 122 billion parameter model Q4 the 122 billion parameter model Q4 quantization. That's 76.5 gigabytes on quantization. That's 76.5 gigabytes on quantization. That's 76.5 gigabytes on disk. NVTop actually shows the Nvidia disk. NVTop actually shows the Nvidia disk. NVTop actually shows the Nvidia side, but it doesn't show side, but it doesn't show side, but it doesn't show the AMD side cuz it's not hooked up to the AMD side cuz it's not hooked up to the AMD side cuz it's not hooked up to work with NVTop. And it ran at 121 work with NVTop. And it ran at 121 work with NVTop. And it ran at 121 tokens per second, which is way faster tokens per second, which is way faster tokens per second, which is way faster than the 32 billion parameter model. The than the 32 billion parameter model. The than the 32 billion parameter model. The bigger model was faster. How? Well, the bigger model was faster. How? Well, the bigger model was faster. How? Well, the 122 billion is a mixture of experts 122 billion is a mixture of experts 122 billion is a mixture of experts model. It's got 122 billion parameters model. It's got 122 billion parameters model. It's got 122 billion parameters total, but only 10 are active at one total, but only 10 are active at one total, but only 10 are active at one time. So, when you're looking at the time. So, when you're looking at the time. So, when you're looking at the model sheet on Hugging Face, for model sheet on Hugging Face, for model sheet on Hugging Face, for example, you see that A10B, that means example, you see that A10B, that means example, you see that A10B, that means active [music] 10. So, at any moment in active [music] 10. So, at any moment in active [music] 10. So, at any moment in time, only 10 are actually loaded. The time, only 10 are actually loaded. The time, only 10 are actually loaded. The model is huge, but it's also light on model is huge, but it's also light on model is huge, but it's also light on its feet. That same trick also helps the its feet. That same trick also helps the its feet. That same trick also helps the little chip inside here, too, the 8060.

  7. little chip inside here, too, the 8060. little chip inside here, too, the 8060. A modern 35 billion expert model like A modern 35 billion expert model like A modern 35 billion expert model like this Qwen 3.6 runs almost six times this Qwen 3.6 runs almost six times this Qwen 3.6 runs almost six times faster than an old-school model of size. faster than an old-school model of size. faster than an old-school model of size. Old-school being dense, new-school being Old-school being dense, new-school being Old-school being dense, new-school being MoE, mixture of experts. Oh, by the way, MoE, mixture of experts. Oh, by the way, MoE, mixture of experts. Oh, by the way, uh when I first plugged this GPU in, the uh when I first plugged this GPU in, the uh when I first plugged this GPU in, the computer refused to turn on altogether. computer refused to turn on altogether. computer refused to turn on altogether. Black screen. The firmware just could Black screen. The firmware just could Black screen. The firmware just could not handle a card this big at startup. not handle a card this big at startup. not handle a card this big at startup. So, I had to boot the machine without So, I had to boot the machine without So, I had to boot the machine without it, and this is a big no-no when it it, and this is a big no-no when it it, and this is a big no-no when it comes to eGPUs, and especially Oculink. comes to eGPUs, and especially Oculink. comes to eGPUs, and especially Oculink. You're not supposed to hot plug them. You're not supposed to hot plug them. You're not supposed to hot plug them. But, that's in fact what I had to do in But, that's in fact what I had to do in But, that's in fact what I had to do in order to get this working. Booted into order to get this working. Booted into order to get this working. Booted into Linux, then powered on the card, and Linux, then powered on the card, and Linux, then powered on the card, and snuck it in through the back [music] snuck it in through the back [music] snuck it in through the back [music] door. Actually, I had Claude Code write door. Actually, I had Claude Code write door. Actually, I had Claude Code write me a script to be able to do that. But, me a script to be able to do that. But, me a script to be able to do that. But, it worked. So, this thing runs basically it worked. So, this thing runs basically it worked. So, this thing runs basically everything on its own, which means we're everything on its own, which means we're everything on its own, which means we're done, right? Yeah, I don't think so. I done, right? Yeah, I don't think so. I done, right? Yeah, I don't think so. I want to see if we can throw a model want to see if we can throw a model want to see if we can throw a model that's even too big for either one of that's even too big for either one of that's even too big for either one of these GPUs to handle. I decided to take these GPUs to handle. I decided to take these GPUs to handle. I decided to take that 122 billion parameter model, but that 122 billion parameter model, but that 122 billion parameter model, but >> [laughter] >> [laughter] >> [laughter] >> there is this Q8 version of it, which is >> there is this Q8 version of it, which is >> there is this Q8 version of it, which is supposed to be higher quality because supposed to be higher quality because supposed to be higher quality because there's less quantization. The less you there's less quantization. The less you there's less quantization. The less you quantize, the less information you're quantize, the less information you're quantize, the less information you're taking away, and therefore you're taking away, and therefore you're taking away, and therefore you're retaining more. So, it's supposed to be retaining more. So, it's supposed to be retaining more. So, it's supposed to be better quality, right? I haven't better quality, right? I haven't better quality, right? I haven't actually tested this one, but I tested actually tested this one, but I tested actually tested this one, but I tested different quantizations, what kind of different quantizations, what kind of different quantizations, what kind of output they produce. I did a whole output they produce. I did a whole output they produce. I did a whole ladder in a different video of it. I'll ladder in a different video of it. I'll ladder in a different video of it. I'll leave a link to it somewhere up here and leave a link to it somewhere up here and leave a link to it somewhere up here and down below. Anyway, this is the big one, down below. Anyway, this is the big one, down below. Anyway, this is the big one, and it's 130 GB on disk, which means it and it's 130 GB on disk, which means it and it's 130 GB on disk, which means it can't fit on either one of those. So, I can't fit on either one of those. So, I can't fit on either one of those. So, I had to split it in half. Part of it

  8. had to split it in half. Part of it had to split it in half. Part of it loads on the Nvidia GPU, and part of it loads on the Nvidia GPU, and part of it loads on the Nvidia GPU, and part of it loads on the AMD GPU. You can call me loads on the AMD GPU. You can call me loads on the AMD GPU. You can call me crazy in the comments, crazy in the comments, crazy in the comments, but you know, we're here to do these but you know, we're here to do these but you know, we're here to do these kinds of crazy things. That's what we do kinds of crazy things. That's what we do kinds of crazy things. That's what we do on this channel. One model running on this channel. One model running on this channel. One model running across two graphics cards from two across two graphics cards from two across two graphics cards from two companies that built completely companies that built completely companies that built completely different software stacks to do different software stacks to do different software stacks to do inference. And it works. There it goes. inference. And it works. There it goes. inference. And it works. There it goes. It's not exactly a 50/50 split. It's It's not exactly a 50/50 split. It's It's not exactly a 50/50 split. It's more of a balanced split. Nvidia got 88 more of a balanced split. Nvidia got 88 more of a balanced split. Nvidia got 88 GB, AMD got 34, and we're getting a nice GB, AMD got 34, and we're getting a nice GB, AMD got 34, and we're getting a nice 37 tokens per second. 38. Nvidia using 37 tokens per second. 38. Nvidia using 37 tokens per second. 38. Nvidia using about 150 watts, about 70 watts on the about 150 watts, about 70 watts on the about 150 watts, about 70 watts on the Radeon. Pretty cool. Told you. But, Radeon. Pretty cool. Told you. But, Radeon. Pretty cool. Told you. But, you're probably wondering, well, how did you're probably wondering, well, how did you're probably wondering, well, how did I do that? The problem is Nvidia and AMD I do that? The problem is Nvidia and AMD I do that? The problem is Nvidia and AMD speak different languages. They're speak different languages. They're speak different languages. They're different software stacks. Nvidia uses different software stacks. Nvidia uses different software stacks. Nvidia uses something called CUDA, which is been something called CUDA, which is been something called CUDA, which is been around for a long time and it's pretty around for a long time and it's pretty around for a long time and it's pretty well developed. And AMD uses something well developed. And AMD uses something well developed. And AMD uses something called ROCm. Not rock'em, but more like called ROCm. Not rock'em, but more like called ROCm. Not rock'em, but more like ROCm. Radeon Open Compute Platform. I ROCm. Radeon Open Compute Platform. I ROCm. Radeon Open Compute Platform. I didn't [music] know that. Now I do. Now didn't [music] know that. Now I do. Now didn't [music] know that. Now I do. Now you know, too. Those two don't mix, but you know, too. Those two don't mix, but you know, too. Those two don't mix, but Vulcan works on both. And I've talked Vulcan works on both. And I've talked Vulcan works on both. And I've talked about Vulcan before. Vulcan is about Vulcan before. Vulcan is about Vulcan before. Vulcan is cross-platform. So, it's meant and cross-platform. So, it's meant and cross-platform. So, it's meant and designed for to be used across different designed for to be used across different designed for to be used across different platforms. That's a very bad sentence, platforms. That's a very bad sentence, platforms. That's a very bad sentence, but you get the idea. When you're using but you get the idea. When you're using but you get the idea. When you're using something like LM Studio, you have the something like LM Studio, you have the something like LM Studio, you have the choice between using CUDA or Vulcan. You choice between using CUDA or Vulcan. You choice between using CUDA or Vulcan. You can pick it right there in LM Studio can pick it right there in LM Studio can pick it right there in LM Studio under the runtime. They even have under the runtime. They even have under the runtime. They even have different performance characteristics.

  9. different performance characteristics. different performance characteristics. Some models run better on Vulcan, some Some models run better on Vulcan, some Some models run better on Vulcan, some run better on ROCm. Vulcan and ROCm also run better on ROCm. Vulcan and ROCm also run better on ROCm. Vulcan and ROCm also work on Windows. Well, mostly Vulcan. work on Windows. Well, mostly Vulcan. work on Windows. Well, mostly Vulcan. Sometimes ROCm, but Sometimes ROCm, but Sometimes ROCm, but yeah, Vulcan. That's the whole trick. yeah, Vulcan. That's the whole trick. yeah, Vulcan. That's the whole trick. Second problem is the model is too big Second problem is the model is too big Second problem is the model is too big for one GPU. So, you have to slice it up for one GPU. So, you have to slice it up for one GPU. So, you have to slice it up like an assembly line. Model weights are like an assembly line. Model weights are like an assembly line. Model weights are made up of layers. Yeah, about 70% live made up of layers. Yeah, about 70% live made up of layers. Yeah, about 70% live on the Nvidia GPU and the rest go on the on the Nvidia GPU and the rest go on the on the Nvidia GPU and the rest go on the AMD GPU. And every single word, every AMD GPU. And every single word, every AMD GPU. And every single word, every token passes through both. We're kind of token passes through both. We're kind of token passes through both. We're kind of limited by that OcuLink GPU. If we had a limited by that OcuLink GPU. If we had a limited by that OcuLink GPU. If we had a faster connection, it would be even faster connection, it would be even faster connection, it would be even faster than this. Another problem that I faster than this. Another problem that I faster than this. Another problem that I ran into is you can configure the AMD ran into is you can configure the AMD ran into is you can configure the AMD GPU in the BIOS on this machine. Not all GPU in the BIOS on this machine. Not all GPU in the BIOS on this machine. Not all the Strix Halo machines give you that the Strix Halo machines give you that the Strix Halo machines give you that control, but here you can and you can control, but here you can and you can control, but here you can and you can allocate the amount of memory you want allocate the amount of memory you want allocate the amount of memory you want to go to that chip, the GPU. And here to go to that chip, the GPU. And here to go to that chip, the GPU. And here I've selected 96 GB. If left on auto, it I've selected 96 GB. If left on auto, it I've selected 96 GB. If left on auto, it not connect with this GPU. So, you have not connect with this GPU. So, you have not connect with this GPU. So, you have to actually specify it. I also was to actually specify it. I also was to actually specify it. I also was curious about concurrency because Nvidia curious about concurrency because Nvidia curious about concurrency because Nvidia GPUs are really good at serving or GPUs are really good at serving or GPUs are really good at serving or specifically this Pro 6000 is really specifically this Pro 6000 is really specifically this Pro 6000 is really good at serving multiple requests, good at serving multiple requests, good at serving multiple requests, multiple people at the same time, good multiple people at the same time, good multiple people at the same time, good with agents. So, I threw eight requests with agents. So, I threw eight requests with agents. So, I threw eight requests at the same time and of course it scales at the same time and of course it scales at the same time and of course it scales pretty nicely. Total speed climbed to pretty nicely. Total speed climbed to pretty nicely. Total speed climbed to about 80 tokens per second here. Oh, and about 80 tokens per second here. Oh, and about 80 tokens per second here. Oh, and to be clear this is together not just to be clear this is together not just to be clear this is together not just the Pro 6000. Now, another nice thing the Pro 6000. Now, another nice thing the Pro 6000. Now, another nice thing here is that while this is a 600 watt here is that while this is a 600 watt here is that while this is a 600 watt card, it only draws about 130 to 150 card, it only draws about 130 to 150 card, it only draws about 130 to 150 watts while doing this thing. And about watts while doing this thing. And about watts while doing this thing. And about 76 watts from the AMD machine. Both of

  10. 76 watts from the AMD machine. Both of 76 watts from the AMD machine. Both of these machines are actually very cool these machines are actually very cool these machines are actually very cool even right now. I'm even surprised by even right now. I'm even surprised by even right now. I'm even surprised by the RTX 6000, how cool it is. the RTX 6000, how cool it is. the RTX 6000, how cool it is. Uh, don't don't put your fingers in the Uh, don't don't put your fingers in the Uh, don't don't put your fingers in the fan, folks. Uh, it's just not a good fan, folks. Uh, it's just not a good fan, folks. Uh, it's just not a good idea. So, that brings me to another idea. So, that brings me to another idea. So, that brings me to another question. Should you buy one of these or question. Should you buy one of these or question. Should you buy one of these or the whole the whole the whole system? Well, if you're going to buy one system? Well, if you're going to buy one system? Well, if you're going to buy one of these, you might as well get the eGPU of these, you might as well get the eGPU of these, you might as well get the eGPU dock. And if you get the eGPU dock, you dock. And if you get the eGPU dock, you dock. And if you get the eGPU dock, you might as well get an eGPU. [music] might as well get an eGPU. [music] might as well get an eGPU. [music] Or, should you Or, should you Or, should you grab [clears throat] grab [clears throat] grab [clears throat] last year's model, which has the exact last year's model, which has the exact last year's model, which has the exact same chip as this if you don't need to same chip as this if you don't need to same chip as this if you don't need to expand it with an eGPU? And looking at expand it with an eGPU? And looking at expand it with an eGPU? And looking at the price of this machine, it's less the price of this machine, it's less the price of this machine, it's less than most Strix Halo machines out there than most Strix Halo machines out there than most Strix Halo machines out there with this new capability, which is with this new capability, which is with this new capability, which is incredible. And what's even more incredible. And what's even more incredible. And what's even more incredible that last year's model, the incredible that last year's model, the incredible that last year's model, the X2, is $3699. It's actually a little bit X2, is $3699. It's actually a little bit X2, is $3699. It's actually a little bit cheaper, $100 cheaper than the new one. cheaper, $100 cheaper than the new one. cheaper, $100 cheaper than the new one. Good luck finding anything under $4000 Good luck finding anything under $4000 Good luck finding anything under $4000 that has a Strix Halo chip in it. These that has a Strix Halo chip in it. These that has a Strix Halo chip in it. These two have the same exact chip, the same two have the same exact chip, the same two have the same exact chip, the same memory, 128 gigs. What's the extra 100 memory, 128 gigs. What's the extra 100 memory, 128 gigs. What's the extra 100 bucks get you? That port. But, you're bucks get you? That port. But, you're bucks get you? That port. But, you're not only getting stuff, you're also kind not only getting stuff, you're also kind not only getting stuff, you're also kind of losing something. First of all, of losing something. First of all, of losing something. First of all, stability. Actually, this thing's pretty stability. Actually, this thing's pretty stability. Actually, this thing's pretty stable with that base. But, it is a stable with that base. But, it is a stable with that base. But, it is a little awkward. You can probably knock little awkward. You can probably knock little awkward. You can probably knock it over with your hand easily when it's it over with your hand easily when it's it over with your hand easily when it's on your desk.

  11. on your desk. on your desk. Ooh, and you can't lay it down. This one Ooh, and you can't lay it down. This one Ooh, and you can't lay it down. This one you can at least lay it down like this. you can at least lay it down like this. you can at least lay it down like this. However, what the X2 still gives you is However, what the X2 still gives you is However, what the X2 still gives you is a ton of ports. If you don't need the a ton of ports. If you don't need the a ton of ports. If you don't need the OcuLink port, this is the one to get. OcuLink port, this is the one to get. OcuLink port, this is the one to get. The X3 is very cool. It's got that The X3 is very cool. It's got that The X3 is very cool. It's got that OcuLink, but it doesn't have all the OcuLink, but it doesn't have all the OcuLink, but it doesn't have all the ports that the X2 has. If you haven't ports that the X2 has. If you haven't ports that the X2 has. If you haven't already caught my X2 video, watch it already caught my X2 video, watch it already caught my X2 video, watch it right over here. Thanks for watching and right over here. Thanks for watching and right over here. Thanks for watching and I'll see you next time. [music]

Summary

This tech analysis focuses on the unprecedented integration of AMD and Nvidia GPUs on a mini PC running a massive AI model, highlighted by the OCuLink port that allows external desktop GPUs. The practical takeaway is that this setup unlocks significantly more AI processing power in a compact form factor, with the added benefit of enhanced security through Surfshark for network traffic.

View original episode ↗