← Back
Aaron Zisk September 16, 2026 26m

8 RTX Pro 6000’s Wasn’t What I Expected

Read full transcript 22 segments
  1. So, I recently built this machine with So, I recently built this machine with four RTX Pro 6000s, but who needs four four RTX Pro 6000s, but who needs four four RTX Pro 6000s, but who needs four when you can have eight? This is when you can have eight? This is when you can have eight? This is definitely sold last week. Right before definitely sold last week. Right before definitely sold last week. Right before I dropped the box on my face, I got I dropped the box on my face, I got I dropped the box on my face, I got something new. Oh, yeah. something new. Oh, yeah. something new. Oh, yeah. This is the Camino Grando, and it's 768 This is the Camino Grando, and it's 768 This is the Camino Grando, and it's 768 GB of VRAM all in one box. Now, the box GB of VRAM all in one box. Now, the box GB of VRAM all in one box. Now, the box is not that much bigger than the other is not that much bigger than the other is not that much bigger than the other box. Actually, it's quite small for box. Actually, it's quite small for box. Actually, it's quite small for having eight of these things in here, having eight of these things in here, having eight of these things in here, but it's heavy. This thing weighs almost but it's heavy. This thing weighs almost but it's heavy. This thing weighs almost as much as me. Eight RTX Pro 6000s are as much as me. Eight RTX Pro 6000s are as much as me. Eight RTX Pro 6000s are stacked in here. And yeah, they're close stacked in here. And yeah, they're close stacked in here. And yeah, they're close together. That's because this whole together. That's because this whole together. That's because this whole thing is liquid cooled. A little secret, thing is liquid cooled. A little secret, thing is liquid cooled. A little secret, I've actually been testing this thing I've actually been testing this thing I've actually been testing this thing for a few months. I was so excited about for a few months. I was so excited about for a few months. I was so excited about this thing. And there are a few things this thing. And there are a few things this thing. And there are a few things that I'd like to point out that are not that I'd like to point out that are not that I'd like to point out that are not great besides the obvious amazing thing great besides the obvious amazing thing great besides the obvious amazing thing that it's going to be super fast. And of that it's going to be super fast. And of that it's going to be super fast. And of course, it's going to hold gigantic course, it's going to hold gigantic course, it's going to hold gigantic models and run them all quickly. But for models and run them all quickly. But for models and run them all quickly. But for whom? Who needs something like this? whom? Who needs something like this? whom? Who needs something like this? Sure, one developer with a dozen coding Sure, one developer with a dozen coding Sure, one developer with a dozen coding agents could use it all at once or a agents could use it all at once or a agents could use it all at once or a whole team of developers can use it in whole team of developers can use it in whole team of developers can use it in the office at the same time. For the office at the same time. For the office at the same time. For example, a 400 GBTE LLM with over example, a 400 GBTE LLM with over example, a 400 GBTE LLM with over 400,000 token context window. Yeah, I 400,000 token context window. Yeah, I 400,000 token context window. Yeah, I ran that all that one box, no cloud. So, ran that all that one box, no cloud. So, ran that all that one box, no cloud. So, of course, I had to put it through its of course, I had to put it through its of course, I had to put it through its paces.

  2. So, here's the thing. Everybody talks So, here's the thing. Everybody talks about running local LLMs and usually about running local LLMs and usually about running local LLMs and usually that means one person, one laptop, one that means one person, one laptop, one that means one person, one laptop, one model, you get 30, 40 tokens per second model, you get 30, 40 tokens per second model, you get 30, 40 tokens per second and you're happy, right? But that's not and you're happy, right? But that's not and you're happy, right? But that's not how developers work anymore. That was a how developers work anymore. That was a how developers work anymore. That was a year ago and now it's completely year ago and now it's completely year ago and now it's completely different. Even if you're solo, you're different. Even if you're solo, you're different. Even if you're solo, you're not running one chat window. You got a not running one chat window. You got a not running one chat window. You got a coding agent on this repo, you got coding agent on this repo, you got coding agent on this repo, you got another one writing tests over there and another one writing tests over there and another one writing tests over there and another one reviewing and every one of another one reviewing and every one of another one reviewing and every one of them sends 50,000 tokens. not your them sends 50,000 tokens. not your them sends 50,000 tokens. not your prompt, but the whole context because prompt, but the whole context because prompt, but the whole context because yeah, you need to include all that. yeah, you need to include all that. yeah, you need to include all that. That's a team's worth of load from one That's a team's worth of load from one That's a team's worth of load from one person. Uh, I probably should rephrase person. Uh, I probably should rephrase person. Uh, I probably should rephrase that next time. And if you're actually a that next time. And if you're actually a that next time. And if you're actually a team, you multiply that by 10. team, you multiply that by 10. team, you multiply that by 10. Today, I'm looking at the Camino Grando. Today, I'm looking at the Camino Grando. Today, I'm looking at the Camino Grando. Camino has been around for a while, and Camino has been around for a while, and Camino has been around for a while, and their specialty is to build these liquid their specialty is to build these liquid their specialty is to build these liquid cool GPU workstations and servers. I did cool GPU workstations and servers. I did cool GPU workstations and servers. I did not buy this one. They let me borrow it not buy this one. They let me borrow it not buy this one. They let me borrow it for a little bit. And at current prices, for a little bit. And at current prices, for a little bit. And at current prices, just the GPUs alone are about 15 grand a just the GPUs alone are about 15 grand a just the GPUs alone are about 15 grand a piece and there's eight of them. So piece and there's eight of them. So piece and there's eight of them. So yeah, you better put these to good use. yeah, you better put these to good use. yeah, you better put these to good use. This thing comes as a 4 unitit chassis.

  3. This thing comes as a 4 unitit chassis. This thing comes as a 4 unitit chassis. You can rack mount it or you can desktop You can rack mount it or you can desktop You can rack mount it or you can desktop mount it, which is what I did. And it's mount it, which is what I did. And it's mount it, which is what I did. And it's 121 lbs just the machine, but it came in 121 lbs just the machine, but it came in 121 lbs just the machine, but it came in a big box. And for insurance purposes, I a big box. And for insurance purposes, I a big box. And for insurance purposes, I did not carry it alone. I had help. did not carry it alone. I had help. did not carry it alone. I had help. Right inside, we got a single AMD epic Right inside, we got a single AMD epic Right inside, we got a single AMD epic 9474F, 9474F, 9474F, 48 cores, 96 threads. It's an epic CPU. 48 cores, 96 threads. It's an epic CPU. 48 cores, 96 threads. It's an epic CPU. What can I say? 512 gigs of DDR5 is in What can I say? 512 gigs of DDR5 is in What can I say? 512 gigs of DDR5 is in there, too, which is an insane amount of there, too, which is an insane amount of there, too, which is an insane amount of RAM right now. I understand. But then RAM right now. I understand. But then RAM right now. I understand. But then there's also the reason we're all here, there's also the reason we're all here, there's also the reason we're all here, eight Nvidia RTX Pro 6000 Blackwell eight Nvidia RTX Pro 6000 Blackwell eight Nvidia RTX Pro 6000 Blackwell Server Edition. Each one has 96 GB of Server Edition. Each one has 96 GB of Server Edition. Each one has 96 GB of VRAM. That's 768 GB of VRAM total. And VRAM. That's 768 GB of VRAM total. And VRAM. That's 768 GB of VRAM total. And did I try to run GLM 5.2 on it? Yes, I did I try to run GLM 5.2 on it? Yes, I did I try to run GLM 5.2 on it? Yes, I did. Did it succeed? Sort of. Now, for did. Did it succeed? Sort of. Now, for did. Did it succeed? Sort of. Now, for scale, the RTX 5090 has 32 gigs of VRAM. scale, the RTX 5090 has 32 gigs of VRAM. scale, the RTX 5090 has 32 gigs of VRAM. So, this is like 24 50s worth of memory So, this is like 24 50s worth of memory So, this is like 24 50s worth of memory in one box. That's a good title for this in one box. That's a good title for this in one box. That's a good title for this video. How do you manage to fit eight of video. How do you manage to fit eight of video. How do you manage to fit eight of these in a 4unit box when each one is a these in a 4unit box when each one is a these in a 4unit box when each one is a two slot card? You can take that cooler two slot card? You can take that cooler two slot card? You can take that cooler off. Every GPU has a custom copper water off. Every GPU has a custom copper water off. Every GPU has a custom copper water block on it. It covers the die, the block on it. It covers the die, the block on it. It covers the die, the memory, and the VRM. And that turns the memory, and the VRM. And that turns the memory, and the VRM. And that turns the two slot card into just a single slot.

  4. two slot card into just a single slot. two slot card into just a single slot. Every pair of cards has its own little Every pair of cards has its own little Every pair of cards has its own little manifold. And the fittings are dripless manifold. And the fittings are dripless manifold. And the fittings are dripless quick disconnects, color-coded red and quick disconnects, color-coded red and quick disconnects, color-coded red and blue, so you can pull one GPU out blue, so you can pull one GPU out blue, so you can pull one GPU out without draining the loop. And the whole without draining the loop. And the whole without draining the loop. And the whole thing is one shared loop for CPUs and thing is one shared loop for CPUs and thing is one shared loop for CPUs and GPUs with a 450ml reservoir with pumps GPUs with a 450ml reservoir with pumps GPUs with a 450ml reservoir with pumps built into it. Now, you may be wondering built into it. Now, you may be wondering built into it. Now, you may be wondering what it takes to power something like what it takes to power something like what it takes to power something like this. this. this. Yeah, 6 1/2 kW. That's like having four Yeah, 6 1/2 kW. That's like having four Yeah, 6 1/2 kW. That's like having four space heaters running all plugged into space heaters running all plugged into space heaters running all plugged into the same wall and it's not going to the same wall and it's not going to the same wall and it's not going to work. You need something special for work. You need something special for work. You need something special for that. Now, the Grondo comes with four that. Now, the Grondo comes with four that. Now, the Grondo comes with four PSUs and each one of them has a regular PSUs and each one of them has a regular PSUs and each one of them has a regular standard plug. So, you can run it into standard plug. So, you can run it into standard plug. So, you can run it into different outlets at the same time. They different outlets at the same time. They different outlets at the same time. They all combine the power, but make sure all combine the power, but make sure all combine the power, but make sure they're on different circuits. Now, they're on different circuits. Now, they're on different circuits. Now, here's one thing that people often here's one thing that people often here's one thing that people often overlook with GPUs, with multiple GPUs, overlook with GPUs, with multiple GPUs, overlook with GPUs, with multiple GPUs, is PCIe lanes. This is why we need the is PCIe lanes. This is why we need the is PCIe lanes. This is why we need the Thread Ripper or the Epic. The Epic has Thread Ripper or the Epic. The Epic has Thread Ripper or the Epic. The Epic has 128 lanes. So each GPU gets a certain 128 lanes. So each GPU gets a certain 128 lanes. So each GPU gets a certain set of lanes, 16 to be precise, or at set of lanes, 16 to be precise, or at set of lanes, 16 to be precise, or at least most of them. Seven of them get 16 least most of them. Seven of them get 16 least most of them. Seven of them get 16 lanes and one of them gets eight lanes.

  5. lanes and one of them gets eight lanes. lanes and one of them gets eight lanes. Now, because GPUs are using up most of Now, because GPUs are using up most of Now, because GPUs are using up most of the lanes, you only have two M.2 slots, the lanes, you only have two M.2 slots, the lanes, you only have two M.2 slots, so you have to use them wisely. And I I so you have to use them wisely. And I I so you have to use them wisely. And I I did run into issues where I had to swap did run into issues where I had to swap did run into issues where I had to swap out models because I ran out of space. out models because I ran out of space. out models because I ran out of space. There's also no Envy Link, so all these There's also no Envy Link, so all these There's also no Envy Link, so all these cards have to talk to each other via cards have to talk to each other via cards have to talk to each other via PCIe. There's four hot swappable 200 PCIe. There's four hot swappable 200 PCIe. There's four hot swappable 200 watt power supplies. That's 8 kW of watt power supplies. That's 8 kW of watt power supplies. That's 8 kW of capacity total. You probably don't want capacity total. You probably don't want capacity total. You probably don't want to use all eight. Yeah, it's going to to use all eight. Yeah, it's going to to use all eight. Yeah, it's going to destroy your stuff. 8 RTX Pro 6000 destroy your stuff. 8 RTX Pro 6000 destroy your stuff. 8 RTX Pro 6000 running at their full 600 watt running at their full 600 watt running at their full 600 watt potential. That's a total of 4,800 watts potential. That's a total of 4,800 watts potential. That's a total of 4,800 watts just in GPUs. Now, I don't have a 8 kW just in GPUs. Now, I don't have a 8 kW just in GPUs. Now, I don't have a 8 kW circuit in here. So, I had two PSUs circuit in here. So, I had two PSUs circuit in here. So, I had two PSUs plugged into a 240 volt 20 amp circuit, plugged into a 240 volt 20 amp circuit, plugged into a 240 volt 20 amp circuit, one into a regular 120 volt 15 amp one into a regular 120 volt 15 amp one into a regular 120 volt 15 amp outlet, and another one in a Jackaryi outlet, and another one in a Jackaryi outlet, and another one in a Jackaryi battery. Yeah, it's on a jackeri. Hey, battery. Yeah, it's on a jackeri. Hey, battery. Yeah, it's on a jackeri. Hey, I'm only human. Okay, I'm only human. Okay, I'm only human. Okay, now just a quick bit of housekeeping. I now just a quick bit of housekeeping. I now just a quick bit of housekeeping. I ran the GPUs capped at 300 watts each ran the GPUs capped at 300 watts each ran the GPUs capped at 300 watts each instead of the default 600. Don't leave instead of the default 600. Don't leave instead of the default 600. Don't leave yet. Don't leave. I'll tell you why it yet. Don't leave. I'll tell you why it yet. Don't leave. I'll tell you why it worked fine. Partly because of my wiring worked fine. Partly because of my wiring worked fine. Partly because of my wiring situation. partly because when I pushed situation. partly because when I pushed situation. partly because when I pushed it at 600, the machine actually dropped it at 600, the machine actually dropped it at 600, the machine actually dropped on me a couple times overnight while I on me a couple times overnight while I on me a couple times overnight while I was doing my long runs. You're probably was doing my long runs. You're probably was doing my long runs. You're probably thinking, half the power, half the thinking, half the power, half the thinking, half the power, half the speed, right? Well, that turned out to speed, right? Well, that turned out to speed, right? Well, that turned out to be one of the more interesting finds in be one of the more interesting finds in be one of the more interesting finds in this project. Whenever I'm working away this project. Whenever I'm working away this project. Whenever I'm working away from home, I end up connecting to from home, I end up connecting to from home, I end up connecting to networks I basically know nothing about.

  6. networks I basically know nothing about. networks I basically know nothing about. A network I completely trust. Not A network I completely trust. Not A network I completely trust. Not really. So, before I start working, I really. So, before I start working, I really. So, before I start working, I connect to Surf Shark. My terminal connect to Surf Shark. My terminal connect to Surf Shark. My terminal sessions, repo traffic, and container sessions, repo traffic, and container sessions, repo traffic, and container downloads all travel through an AES 256 downloads all travel through an AES 256 downloads all travel through an AES 256 encrypted tunnel, and Clean Web blocks a encrypted tunnel, and Clean Web blocks a encrypted tunnel, and Clean Web blocks a lot of the tracking and advertising junk lot of the tracking and advertising junk lot of the tracking and advertising junk that I don't want. There's also an that I don't want. There's also an that I don't want. There's also an independently audited no logs policy. independently audited no logs policy. independently audited no logs policy. Plus, RAM only servers that get wiped Plus, RAM only servers that get wiped Plus, RAM only servers that get wiped every time they restart. And with the every time they restart. And with the every time they restart. And with the amount of hardware I have and I travel amount of hardware I have and I travel amount of hardware I have and I travel with, unlimited devices is a pretty big with, unlimited devices is a pretty big with, unlimited devices is a pretty big deal. It covers my MacBook, my phone, deal. It covers my MacBook, my phone, deal. It covers my MacBook, my phone, and pretty much any other device I have and pretty much any other device I have and pretty much any other device I have in my backpack that connects to the in my backpack that connects to the in my backpack that connects to the internet. And when I'm checking a server internet. And when I'm checking a server internet. And when I'm checking a server or sshing back to the office, that extra or sshing back to the office, that extra or sshing back to the office, that extra layer makes a lot of sense. So head over layer makes a lot of sense. So head over layer makes a lot of sense. So head over to surfshark.com/alexiscin to surfshark.com/alexiscin to surfshark.com/alexiscin for four extra months free. And if it's for four extra months free. And if it's for four extra months free. And if it's not for you, there's a 30-day money back not for you, there's a 30-day money back not for you, there's a 30-day money back guarantee. Protect your connection. Get guarantee. Protect your connection. Get guarantee. Protect your connection. Get your work done. Now, back to the video. your work done. Now, back to the video. your work done. Now, back to the video. Let's talk about noise. The Camino says Let's talk about noise. The Camino says Let's talk about noise. The Camino says 39 to 70 dB depending on which fans you 39 to 70 dB depending on which fans you 39 to 70 dB depending on which fans you get because you can customize that.

  7. get because you can customize that. get because you can customize that. There's a 6200 RPM version and a 3,000 There's a 6200 RPM version and a 3,000 There's a 6200 RPM version and a 3,000 RPM version. But let me tell you this, RPM version. But let me tell you this, RPM version. But let me tell you this, you're not going to have this on your you're not going to have this on your you're not going to have this on your desk next to you. >> Even at its quietest setting, under full >> Even at its quietest setting, under full AGPU load, it gets pretty loud. Yeah, it AGPU load, it gets pretty loud. Yeah, it AGPU load, it gets pretty loud. Yeah, it could get pretty loud if you're sitting could get pretty loud if you're sitting could get pretty loud if you're sitting next to it. I'd move away a little bit next to it. I'd move away a little bit next to it. I'd move away a little bit more. Maybe even more. more. Maybe even more. more. Maybe even more. Can you still hear it? Yeah. You need to Can you still hear it? Yeah. You need to Can you still hear it? Yeah. You need to be far away from it. Now, the cooling is be far away from it. Now, the cooling is be far away from it. Now, the cooling is a different story. The hottest GPU I saw a different story. The hottest GPU I saw a different story. The hottest GPU I saw was 62 Celsius and the CPU peaked at 65. was 62 Celsius and the CPU peaked at 65. was 62 Celsius and the CPU peaked at 65. So, the liquid cooling here is not a So, the liquid cooling here is not a So, the liquid cooling here is not a gimmick. That's the reason this box gimmick. That's the reason this box gimmick. That's the reason this box exists. It's really good and efficient. exists. It's really good and efficient. exists. It's really good and efficient. All right, let's go over some model All right, let's go over some model All right, let's go over some model results. Ubuntu 24.04, Nvidia Open results. Ubuntu 24.04, Nvidia Open results. Ubuntu 24.04, Nvidia Open Driver 610, CUDA 13, and VLM Nightly. I Driver 610, CUDA 13, and VLM Nightly. I Driver 610, CUDA 13, and VLM Nightly. I kept it updated as I was doing my kept it updated as I was doing my kept it updated as I was doing my testing. And we're using Tensor testing. And we're using Tensor testing. And we're using Tensor Parallelism across all eight GPUs.

  8. Parallelism across all eight GPUs. Parallelism across all eight GPUs. There's GLM 5.2 right on top. 433 GB of There's GLM 5.2 right on top. 433 GB of There's GLM 5.2 right on top. 433 GB of weights. And then a few other new ones weights. And then a few other new ones weights. And then a few other new ones came out, so I had to do those. Here's a came out, so I had to do those. Here's a came out, so I had to do those. Here's a list of five models I ran mostly in list of five models I ran mostly in list of five models I ran mostly in NVFP4 format, which is the 4-bit format, NVFP4 format, which is the 4-bit format, NVFP4 format, which is the 4-bit format, floating point that Blackwell supports. floating point that Blackwell supports. floating point that Blackwell supports. It's what it was built for. The big boy It's what it was built for. The big boy It's what it was built for. The big boy GLM 5.2. This one was served with 49,600 GLM 5.2. This one was served with 49,600 GLM 5.2. This one was served with 49,600 token context window. Quen 3 235B. token context window. Quen 3 235B. token context window. Quen 3 235B. Getting a little bit long in the tooth Getting a little bit long in the tooth Getting a little bit long in the tooth that model, but it's pretty big. It's a that model, but it's pretty big. It's a that model, but it's pretty big. It's a mixture of experts model. 22 billion mixture of experts model. 22 billion mixture of experts model. 22 billion parameters active. 144 GB on disk. This parameters active. 144 GB on disk. This parameters active. 144 GB on disk. This size right here, 144 to about 185 for size right here, 144 to about 185 for size right here, 144 to about 185 for GLM Flash 5.3. This is kind of like the GLM Flash 5.3. This is kind of like the GLM Flash 5.3. This is kind of like the sweet spot for this type of machine. sweet spot for this type of machine. sweet spot for this type of machine. Sure, it can run bigger models, but then Sure, it can run bigger models, but then Sure, it can run bigger models, but then you're going to do context. You're going you're going to do context. You're going you're going to do context. You're going to have multiple sessions, multiple to have multiple sessions, multiple to have multiple sessions, multiple agents hitting it at the same time, so agents hitting it at the same time, so agents hitting it at the same time, so you want to probably stick around here. you want to probably stick around here. you want to probably stick around here. GLM 5.2 loads in about 4 1/2 minutes and GLM 5.2 loads in about 4 1/2 minutes and GLM 5.2 loads in about 4 1/2 minutes and lands at 738 GB of the 784 available.

  9. lands at 738 GB of the 784 available. lands at 738 GB of the 784 available. That's about 95% of the VRAM on this That's about 95% of the VRAM on this That's about 95% of the VRAM on this machine just gone for one model. But it machine just gone for one model. But it machine just gone for one model. But it fits. That's the whole point. There's fits. That's the whole point. There's fits. That's the whole point. There's basically no other single box you can basically no other single box you can basically no other single box you can put on a desk that does this. There's a put on a desk that does this. There's a put on a desk that does this. There's a DJX station, which I recently did a DJX station, which I recently did a DJX station, which I recently did a video on. That one has 252 GB of high video on. That one has 252 GB of high video on. That one has 252 GB of high bandwidth memory, which is way faster bandwidth memory, which is way faster bandwidth memory, which is way faster than this memory, but it doesn't have a than this memory, but it doesn't have a than this memory, but it doesn't have a total capacity to hold this model all in total capacity to hold this model all in total capacity to hold this model all in VRAM. By the way, the DJX Station is VRAM. By the way, the DJX Station is VRAM. By the way, the DJX Station is also really good at the smaller models. also really good at the smaller models. also really good at the smaller models. And I want to compare how the smaller And I want to compare how the smaller And I want to compare how the smaller models act on this machine versus the models act on this machine versus the models act on this machine versus the DJX station. Stay tuned for that. DJX station. Stay tuned for that. DJX station. Stay tuned for that. Let's start where most of developers Let's start where most of developers Let's start where most of developers live. One developer, one agent. I know, live. One developer, one agent. I know, live. One developer, one agent. I know, I know you're using more than that now, I know you're using more than that now, I know you're using more than that now, but let's start at the baseline. 2048 but let's start at the baseline. 2048 but let's start at the baseline. 2048 token prompt, 128 tokens out. On a token prompt, 128 tokens out. On a token prompt, 128 tokens out. On a single RTX Pro 6000, a 30 billion single RTX Pro 6000, a 30 billion single RTX Pro 6000, a 30 billion parameter class model is around 100 parameter class model is around 100 parameter class model is around 100 tokens per second. So, what happens when tokens per second. So, what happens when tokens per second. So, what happens when you spread big models across eight cards you spread big models across eight cards you spread big models across eight cards over PCIe? Ah, GLM 5.2 48 tokens per over PCIe? Ah, GLM 5.2 48 tokens per over PCIe? Ah, GLM 5.2 48 tokens per second generation. Okay, that's a 433 second generation. Okay, that's a 433 second generation. Okay, that's a 433 gig model doing 48. That's pretty good.

  10. gig model doing 48. That's pretty good. gig model doing 48. That's pretty good. Not blazing, but it's fast enough. We'll Not blazing, but it's fast enough. We'll Not blazing, but it's fast enough. We'll come back to that. Quen 3 235B 85. Wow. come back to that. Quen 3 235B 85. Wow. come back to that. Quen 3 235B 85. Wow. Okay. Now, here are some of the more Okay. Now, here are some of the more Okay. Now, here are some of the more modern models. Deepseek v4 flash 102 modern models. Deepseek v4 flash 102 modern models. Deepseek v4 flash 102 tokens per second. GLM 5.3 104 and Quen tokens per second. GLM 5.3 104 and Quen tokens per second. GLM 5.3 104 and Quen 3.8 wins this one. 126 tokens per 3.8 wins this one. 126 tokens per 3.8 wins this one. 126 tokens per second. Quen 3.8 flash next by the way. second. Quen 3.8 flash next by the way. second. Quen 3.8 flash next by the way. All right. Quen 4 architecture. Boom. All right. Quen 4 architecture. Boom. All right. Quen 4 architecture. Boom. Oh, if you don't know what I'm talking Oh, if you don't know what I'm talking Oh, if you don't know what I'm talking about, Quen 3.8 models that came out about, Quen 3.8 models that came out about, Quen 3.8 models that came out earlier are actually Quen 3 model earlier are actually Quen 3 model earlier are actually Quen 3 model architecture or 3.8 and Quen 3.8 Flash. architecture or 3.8 and Quen 3.8 Flash. architecture or 3.8 and Quen 3.8 Flash. Next has Quen 4 architecture. Don't ask Next has Quen 4 architecture. Don't ask Next has Quen 4 architecture. Don't ask me why. I don't know why they did that. me why. I don't know why they did that. me why. I don't know why they did that. Okay. The other number you should be Okay. The other number you should be Okay. The other number you should be interested in prompt processing because interested in prompt processing because interested in prompt processing because this number matters a lot especially this number matters a lot especially this number matters a lot especially when you're using agents in a code when you're using agents in a code when you're using agents in a code editor. For example, GLM 5.2 about 2200 editor. For example, GLM 5.2 about 2200 editor. For example, GLM 5.2 about 2200 tokens per second. Quen 3235B we're up a tokens per second. Quen 3235B we're up a tokens per second. Quen 3235B we're up a lot to 4779 tokens per second. GLM 5.3 lot to 4779 tokens per second. GLM 5.3 lot to 4779 tokens per second. GLM 5.3 Flash 8,300. Deepseek V4 flash 8700 and Flash 8,300. Deepseek V4 flash 8700 and Flash 8,300. Deepseek V4 flash 8700 and Quen 3.8 Flash. Next, holy cow 12,600.

  11. Quen 3.8 Flash. Next, holy cow 12,600. Quen 3.8 Flash. Next, holy cow 12,600. That's fast. Now, here's what it looks That's fast. Now, here's what it looks That's fast. Now, here's what it looks like in practice for a single user for like in practice for a single user for like in practice for a single user for time to first token on a 2,00 token time to first token on a 2,00 token time to first token on a 2,00 token prompt. A quarter of a second on flash prompt. A quarter of a second on flash prompt. A quarter of a second on flash next up to almost a second for GLM 5.2 next up to almost a second for GLM 5.2 next up to almost a second for GLM 5.2 and everything else is in between there. and everything else is in between there. and everything else is in between there. When we look back at that 12,700 tokens When we look back at that 12,700 tokens When we look back at that 12,700 tokens per second of prompt processing, that per second of prompt processing, that per second of prompt processing, that means a 2,00 token prompt is basically means a 2,00 token prompt is basically means a 2,00 token prompt is basically instant. instant. instant. Now, I gave it a 128,000 token prompt. Now, I gave it a 128,000 token prompt. Now, I gave it a 128,000 token prompt. That's basically a whole codebase. a That's basically a whole codebase. a That's basically a whole codebase. a small code base on a Mac, on a mini PC, small code base on a Mac, on a mini PC, small code base on a Mac, on a mini PC, or pretty much anything else I've or pretty much anything else I've or pretty much anything else I've tested. A 128,000 token prompt is kind tested. A 128,000 token prompt is kind tested. A 128,000 token prompt is kind of a let's go make coffee situation. of a let's go make coffee situation. of a let's go make coffee situation. >> Oh, okay. Coffee can wait. I'm still >> Oh, okay. Coffee can wait. I'm still >> Oh, okay. Coffee can wait. I'm still waiting for the M5 Ultras, which are waiting for the M5 Ultras, which are waiting for the M5 Ultras, which are about to come out. That might change about to come out. That might change about to come out. That might change things a little bit, but on everything things a little bit, but on everything things a little bit, but on everything older, you're going to be waiting older, you're going to be waiting older, you're going to be waiting minutes. As the prompt length increases, minutes. As the prompt length increases, minutes. As the prompt length increases, your time to first token also increases, your time to first token also increases, your time to first token also increases, sometimes quite a bit. Gwen 3.8 8 flash. sometimes quite a bit. Gwen 3.8 8 flash. sometimes quite a bit. Gwen 3.8 8 flash. Next, we're at 14 1/2 seconds to first Next, we're at 14 1/2 seconds to first Next, we're at 14 1/2 seconds to first token at 128,000 tokens. For GLM 5.3 token at 128,000 tokens. For GLM 5.3 token at 128,000 tokens. For GLM 5.3 flash, we're at about 18 and 12.

  12. flash, we're at about 18 and 12. flash, we're at about 18 and 12. Deepseek V4 flash, we're at 26. And in Deepseek V4 flash, we're at 26. And in Deepseek V4 flash, we're at 26. And in GLM 5.2, wo, 50 seconds. That's the big GLM 5.2, wo, 50 seconds. That's the big GLM 5.2, wo, 50 seconds. That's the big boy, right? So, yeah, to be expected. boy, right? So, yeah, to be expected. boy, right? So, yeah, to be expected. Still though, under a minute for the Still though, under a minute for the Still though, under a minute for the biggest model and about 15 seconds for biggest model and about 15 seconds for biggest model and about 15 seconds for the fastest one. Quen 3 235B didn't make the fastest one. Quen 3 235B didn't make the fastest one. Quen 3 235B didn't make it that far. We only got 32 and uh then it that far. We only got 32 and uh then it that far. We only got 32 and uh then it kind of crashed on me for the rest of it kind of crashed on me for the rest of it kind of crashed on me for the rest of the times. Sorry. So this is what eight the times. Sorry. So this is what eight the times. Sorry. So this is what eight GPUs earning their keep looks like GPUs earning their keep looks like GPUs earning their keep looks like because prompt processing that's the because prompt processing that's the because prompt processing that's the first stage of inference before token first stage of inference before token first stage of inference before token generation happens. This is all generation happens. This is all generation happens. This is all computebound. So everything happens on computebound. So everything happens on computebound. So everything happens on the GPU chip. But then the second stage the GPU chip. But then the second stage the GPU chip. But then the second stage is token generation. So what happens to is token generation. So what happens to is token generation. So what happens to generation speed once all that context generation speed once all that context generation speed once all that context is sitting in memory? Huh? Every line is is sitting in memory? Huh? Every line is is sitting in memory? Huh? Every line is pretty much flat here, huh? Nothing pretty much flat here, huh? Nothing pretty much flat here, huh? Nothing happens. Flash Next is still generating happens. Flash Next is still generating happens. Flash Next is still generating at about 120 to 122 tokens per second at about 120 to 122 tokens per second at about 120 to 122 tokens per second even at 128,000 token context. GLM 5.2 even at 128,000 token context. GLM 5.2 even at 128,000 token context. GLM 5.2 pretty steady around what is that 40 50 pretty steady around what is that 40 50 pretty steady around what is that 40 50 none of them slow down. That's pretty none of them slow down. That's pretty none of them slow down. That's pretty incredible.

  13. incredible. incredible. Now, quick aside here because this Now, quick aside here because this Now, quick aside here because this matters to agents. I'm using VLM here to matters to agents. I'm using VLM here to matters to agents. I'm using VLM here to serve the models and VLM has prefix serve the models and VLM has prefix serve the models and VLM has prefix caching. If your conversation has about caching. If your conversation has about caching. If your conversation has about 16,000 tokens in it and you send it 16,000 tokens in it and you send it 16,000 tokens in it and you send it another message without caching, GLM 5.2 another message without caching, GLM 5.2 another message without caching, GLM 5.2 takes 5.9 seconds to first token. And takes 5.9 seconds to first token. And takes 5.9 seconds to first token. And with caching, 0.83 with caching, 0.83 with caching, 0.83 seconds, 7 times faster. And you kind of seconds, 7 times faster. And you kind of seconds, 7 times faster. And you kind of see the same pattern all throughout. see the same pattern all throughout. see the same pattern all throughout. You'll say, "Alex, that's obvious. Turn You'll say, "Alex, that's obvious. Turn You'll say, "Alex, that's obvious. Turn on caching, right?" Well, especially on caching, right?" Well, especially on caching, right?" Well, especially with Agentic Flows, you want to have with Agentic Flows, you want to have with Agentic Flows, you want to have that option on. that option on. that option on. Okay, remember the power cap? I ran the Okay, remember the power cap? I ran the Okay, remember the power cap? I ran the same test at 300 watts and then 600 same test at 300 watts and then 600 same test at 300 watts and then 600 watts. And you said 300 is going to be watts. And you said 300 is going to be watts. And you said 300 is going to be slower than 600. Well, actually, I'm the slower than 600. Well, actually, I'm the slower than 600. Well, actually, I'm the one that said it, but you were thinking one that said it, but you were thinking one that said it, but you were thinking it, weren't you? GLM 5.2 48 tokens per it, weren't you? GLM 5.2 48 tokens per it, weren't you? GLM 5.2 48 tokens per second at 300 watts. Quen 3235b 85. GLM second at 300 watts. Quen 3235b 85. GLM second at 300 watts. Quen 3235b 85. GLM 5.2, this is 32 users now, by the way. 5.2, this is 32 users now, by the way. 5.2, this is 32 users now, by the way. We got 111 tokens per second at 300 We got 111 tokens per second at 300 We got 111 tokens per second at 300 watts. And Quen 3235B watts. And Quen 3235B watts. And Quen 3235B 64 users 254 tokens per second. What 64 users 254 tokens per second. What 64 users 254 tokens per second. What does this look like at 600 watts? Boom.

  14. does this look like at 600 watts? Boom. does this look like at 600 watts? Boom. It's a wash. I wouldn't even call that It's a wash. I wouldn't even call that It's a wash. I wouldn't even call that close to being different. Really? Well, close to being different. Really? Well, close to being different. Really? Well, actually 254 and 252. Yeah, I know. actually 254 and 252. Yeah, I know. actually 254 and 252. Yeah, I know. Okay, calm down. It's close enough. And Okay, calm down. It's close enough. And Okay, calm down. It's close enough. And look at the actual GPU power draw. At look at the actual GPU power draw. At look at the actual GPU power draw. At 300 watt caps, the eight GPU together 300 watt caps, the eight GPU together 300 watt caps, the eight GPU together pull about 1,550 watts during inference. pull about 1,550 watts during inference. pull about 1,550 watts during inference. This is for GLM 5.2. This is for GLM 5.2. This is for GLM 5.2. >> That's just one of the plugs and it's >> That's just one of the plugs and it's >> That's just one of the plugs and it's drawing 126 watts idle. drawing 126 watts idle. drawing 126 watts idle. >> And at 600, they pull about 1,700 W per >> And at 600, they pull about 1,700 W per >> And at 600, they pull about 1,700 W per card. The peak I ever saw was 268 watts, card. The peak I ever saw was 268 watts, card. The peak I ever saw was 268 watts, which means that these cards are memory which means that these cards are memory which means that these cards are memory bandwidth bound during inference. and bandwidth bound during inference. and bandwidth bound during inference. and they're talking to each other over PCIe, they're talking to each other over PCIe, they're talking to each other over PCIe, so they never get anywhere near the 600 so they never get anywhere near the 600 so they never get anywhere near the 600 watt limit. Doubling the power limit watt limit. Doubling the power limit watt limit. Doubling the power limit bought me nothing more than heat and a bought me nothing more than heat and a bought me nothing more than heat and a less stable box. Which brings me to a less stable box. Which brings me to a less stable box. Which brings me to a question that I've had for a while. The question that I've had for a while. The question that I've had for a while. The workstation edition, which has 600 watt workstation edition, which has 600 watt workstation edition, which has 600 watt cap versus the Max Q edition, and that's cap versus the Max Q edition, and that's cap versus the Max Q edition, and that's going to be another video, I think. So, going to be another video, I think. So, going to be another video, I think. So, stay tuned for that. Make sure you stay tuned for that. Make sure you stay tuned for that. Make sure you subscribe. By the way, just sitting subscribe. By the way, just sitting subscribe. By the way, just sitting there with a model loaded and nothing there with a model loaded and nothing there with a model loaded and nothing happening, the GPUs are pulling about happening, the GPUs are pulling about happening, the GPUs are pulling about 700 watts idle. just sitting there. It's 700 watts idle. just sitting there. It's 700 watts idle. just sitting there. It's doing nothing expensively.

  15. doing nothing expensively. doing nothing expensively. So, make sure you pay the power bill. So, make sure you pay the power bill. So, make sure you pay the power bill. Don't get me wrong, this is not magic. Don't get me wrong, this is not magic. Don't get me wrong, this is not magic. Every time those AGPUs sync up, they go Every time those AGPUs sync up, they go Every time those AGPUs sync up, they go over PCIe and you feel it. It's not Envy over PCIe and you feel it. It's not Envy over PCIe and you feel it. It's not Envy Link. I would love to test some Envy Link. I would love to test some Envy Link. I would love to test some Envy Link on this channel. So, if you're Link on this channel. So, if you're Link on this channel. So, if you're listening and you have some of those listening and you have some of those listening and you have some of those laying around, let me know. Get in laying around, let me know. Get in laying around, let me know. Get in touch. A100s, anyone? H200s, Quen 3 235B touch. A100s, anyone? H200s, Quen 3 235B touch. A100s, anyone? H200s, Quen 3 235B on four GPUs with tensor parallel 4 did on four GPUs with tensor parallel 4 did on four GPUs with tensor parallel 4 did 58 tokens per second in an earlier test. 58 tokens per second in an earlier test. 58 tokens per second in an earlier test. I did on all eight GPUs TP8 it did 37. I did on all eight GPUs TP8 it did 37. I did on all eight GPUs TP8 it did 37. So slower more GPUs slower because of So slower more GPUs slower because of So slower more GPUs slower because of the all reduce over PCIe costs more than the all reduce over PCIe costs more than the all reduce over PCIe costs more than extra compute gives you. There's the extra compute gives you. There's the extra compute gives you. There's the balance that you need to maintain and balance that you need to maintain and balance that you need to maintain and you need to tweak it. So for a smaller you need to tweak it. So for a smaller you need to tweak it. So for a smaller model that would fit on four cards, not model that would fit on four cards, not model that would fit on four cards, not all eight, you would actually probably all eight, you would actually probably all eight, you would actually probably better off running two copies on four better off running two copies on four better off running two copies on four cards each and serve twice the people, cards each and serve twice the people, cards each and serve twice the people, not one copy on eight. GLM 5.2 at 48 not one copy on eight. GLM 5.2 at 48 not one copy on eight. GLM 5.2 at 48 tokens per second is with the plain tokens per second is with the plain tokens per second is with the plain official recipe. There's no spec official recipe. There's no spec official recipe. There's no spec decoding here. If you're curious about decoding here. If you're curious about decoding here. If you're curious about speculative decoding, I made a separate speculative decoding, I made a separate speculative decoding, I made a separate video on that. I'll link to it down video on that. I'll link to it down video on that. I'll link to it down below. And with MTP turned on, I got below. And with MTP turned on, I got below. And with MTP turned on, I got about 100 tokens per second in earlier about 100 tokens per second in earlier about 100 tokens per second in earlier testing. But yeah, that config was a testing. But yeah, that config was a testing. But yeah, that config was a little bit fussier to make. Now, the little bit fussier to make. Now, the little bit fussier to make. Now, the ecosystem is still catching up to these ecosystem is still catching up to these ecosystem is still catching up to these GPUs. These are still considered pretty GPUs. These are still considered pretty GPUs. These are still considered pretty new, even though they've been around for new, even though they've been around for new, even though they've been around for now almost 2 years in some form or now almost 2 years in some form or now almost 2 years in some form or another. For example, GLM 5.3 Flash another. For example, GLM 5.3 Flash another. For example, GLM 5.3 Flash would not start on any stock VLM image.

  16. would not start on any stock VLM image. would not start on any stock VLM image. I tried six different configurations, I tried six different configurations, I tried six different configurations, got the same kernel error every time got the same kernel error every time got the same kernel error every time because the model uses an attention because the model uses an attention because the model uses an attention variant the Blackwell workstation kernel variant the Blackwell workstation kernel variant the Blackwell workstation kernel doesn't handle yet. It's too new. So, doesn't handle yet. It's too new. So, doesn't handle yet. It's too new. So, once VLM upstreams it and supports it, once VLM upstreams it and supports it, once VLM upstreams it and supports it, you might get slightly different you might get slightly different you might get slightly different numbers. It may be better even. numbers. It may be better even. numbers. It may be better even. All right, let's try this out. All right, let's try this out. All right, let's try this out. This is going to be nuts. All four of This is going to be nuts. All four of This is going to be nuts. All four of these machines are going to ping the these machines are going to ping the these machines are going to ping the Grandondo, which is obviously not in Grandondo, which is obviously not in Grandondo, which is obviously not in this room, but you might be able to hear this room, but you might be able to hear this room, but you might be able to hear it. Yeah, it's spinning up. This is Maya. She's doing a test suite This is Maya. She's doing a test suite for orders API. This is Ravi. He's for orders API. This is Ravi. He's for orders API. This is Ravi. He's working on a metrics dashboard. Lena is working on a metrics dashboard. Lena is working on a metrics dashboard. Lena is hardening the inventory report script. hardening the inventory report script. hardening the inventory report script. Yeah, I didn't write this. Okay. But Yeah, I didn't write this. Okay. But Yeah, I didn't write this. Okay. But it's still a test. And Tom over there. it's still a test. And Tom over there. it's still a test. And Tom over there. Tom containerizing the Q worker. Yeah. Tom containerizing the Q worker. Yeah. Tom containerizing the Q worker. Yeah. So, they're all busy. Okay. I just So, they're all busy. Okay. I just So, they're all busy. Okay. I just That's the whole point here. I need to That's the whole point here. I need to That's the whole point here. I need to keep them busy.

  17. keep them busy. keep them busy. All right. Now, we're cooking. Each one All right. Now, we're cooking. Each one All right. Now, we're cooking. Each one of these is doing about a bunch. A of these is doing about a bunch. A of these is doing about a bunch. A bunch. Whatever the agent needs. Each bunch. Whatever the agent needs. Each bunch. Whatever the agent needs. Each one of these launched multiple agents one of these launched multiple agents one of these launched multiple agents and sub agents all pointing at the grand and sub agents all pointing at the grand and sub agents all pointing at the grand all at the same time. all at the same time. all at the same time. Some of these require permission. Let's Some of these require permission. Let's Some of these require permission. Let's do always. Boom. And confirm. He must be do always. Boom. And confirm. He must be do always. Boom. And confirm. He must be doing something dangerous there. Maya or doing something dangerous there. Maya or doing something dangerous there. Maya or Lena. This is showing the utilization Lena. This is showing the utilization Lena. This is showing the utilization and the VRAMm usage on each of the GPUs and the VRAMm usage on each of the GPUs and the VRAMm usage on each of the GPUs running on the Grando. So, we're 100% running on the Grando. So, we're 100% running on the Grando. So, we're 100% utilization, jumping up to about 185 to utilization, jumping up to about 185 to utilization, jumping up to about 185 to 190 watts per GPU and 90 out of 96 GB 190 watts per GPU and 90 out of 96 GB 190 watts per GPU and 90 out of 96 GB used on each of those RTX Pro 6000s. And used on each of those RTX Pro 6000s. And used on each of those RTX Pro 6000s. And this is running DeepSync V4 Flash. It's this is running DeepSync V4 Flash. It's this is running DeepSync V4 Flash. It's going pretty fast. It's going decently going pretty fast. It's going decently going pretty fast. It's going decently fast. And all these agents and sub fast. And all these agents and sub fast. And all these agents and sub agents are all getting their share. It's agents are all getting their share. It's agents are all getting their share. It's plenty. We're gonna go with 32 agents on plenty. We're gonna go with 32 agents on plenty. We're gonna go with 32 agents on each machine to start with. And boom.

  18. each machine to start with. And boom. each machine to start with. And boom. Three, two, one, and go. Three, two, one, and go. Three, two, one, and go. Woo. That's spinning up. You might be Woo. That's spinning up. You might be Woo. That's spinning up. You might be able to hear it from the other room. able to hear it from the other room. able to hear it from the other room. Nice. So, we got a total of 128 Nice. So, we got a total of 128 Nice. So, we got a total of 128 concurrent agents. Each one has 32. And concurrent agents. Each one has 32. And concurrent agents. Each one has 32. And it's chugging away about 3,000 tokens it's chugging away about 3,000 tokens it's chugging away about 3,000 tokens per second. All right, let's stop that. per second. All right, let's stop that. per second. All right, let's stop that. Let's go to 128 Let's go to 128 Let's go to 128 agents. per machine. Oh my gosh, this is going per machine. Oh my gosh, this is going to be crazy. Launch all. Oh boy, it's to be crazy. Launch all. Oh boy, it's to be crazy. Launch all. Oh boy, it's doing it. 512 concurrent agents. Wow. doing it. 512 concurrent agents. Wow. doing it. 512 concurrent agents. Wow. Utilization is about the same. 100%. It Utilization is about the same. 100%. It Utilization is about the same. 100%. It better be. Finally. Let's do 256. And better be. Finally. Let's do 256. And better be. Finally. Let's do 256. And that's 256 agents for Maya, Ravi, Lena, that's 256 agents for Maya, Ravi, Lena, that's 256 agents for Maya, Ravi, Lena, and Tom. and Tom. and Tom. The four ninjas over here. The four ninjas over here. The four ninjas over here. Let's go. Launch all. Let's go. Launch all. Let's go. Launch all. Three, two, one. Boom. Well, Three, two, one. Boom. Well, Three, two, one. Boom. Well, it's not liking it. It's definitely not it's not liking it. It's definitely not it's not liking it. It's definitely not liking it. Oh boy. It's trying. It's liking it. Oh boy. It's trying. It's liking it. Oh boy. It's trying. It's trying. 1,000 concurrent agents. Hey, it trying. 1,000 concurrent agents. Hey, it trying. 1,000 concurrent agents. Hey, it did it. We're at 6,000 to 7,000 tokens did it. We're at 6,000 to 7,000 tokens did it. We're at 6,000 to 7,000 tokens per second per second per second and it's actually doing it. It had to and it's actually doing it. It had to and it's actually doing it. It had to spin up. Wow. KV cash 41% and zero spin up. Wow. KV cash 41% and zero spin up. Wow. KV cash 41% and zero errors. Time to first token is pretty errors. Time to first token is pretty errors. Time to first token is pretty high there. And it's not super happy, high there. And it's not super happy, high there. And it's not super happy, but it's completing the work. Since I've

  19. but it's completing the work. Since I've but it's completing the work. Since I've been talking here, we we've done over been talking here, we we've done over been talking here, we we've done over 630,000 630,000 630,000 tokens total. Not bad. Not bad. Job well tokens total. Not bad. Not bad. Job well tokens total. Not bad. Not bad. Job well done, humans. You can all go home. done, humans. You can all go home. done, humans. You can all go home. Bye, Maya. I'll miss you. Bye, Maya. I'll miss you. Bye, Maya. I'll miss you. So I ran 1 2 4 8 16 32 64 concurrent So I ran 1 2 4 8 16 32 64 concurrent So I ran 1 2 4 8 16 32 64 concurrent users against all five models. Same 448 users against all five models. Same 448 users against all five models. Same 448 token prompt. No queue and every user token prompt. No queue and every user token prompt. No queue and every user actually in flight at the same time. actually in flight at the same time. actually in flight at the same time. Quen 3.8 flash. Next one user gets 126 Quen 3.8 flash. Next one user gets 126 Quen 3.8 flash. Next one user gets 126 tokens per second. And as we add more tokens per second. And as we add more tokens per second. And as we add more users, the total throughput goes up users, the total throughput goes up users, the total throughput goes up because we get more tokens per second because we get more tokens per second because we get more tokens per second generated totally by the machine. But generated totally by the machine. But generated totally by the machine. But but the number of tokens per second per but the number of tokens per second per but the number of tokens per second per user goes down. So with four users we user goes down. So with four users we user goes down. So with four users we get 90. Eight users 82. And by the way, get 90. Eight users 82. And by the way, get 90. Eight users 82. And by the way, this is either eight users or it could this is either eight users or it could this is either eight users or it could be one developer running eight agents in be one developer running eight agents in be one developer running eight agents in parallel. So here we've got 82 tokens parallel. So here we've got 82 tokens parallel. So here we've got 82 tokens per second for all eight of your agents, per second for all eight of your agents, per second for all eight of your agents, which is still pretty good. That's which is still pretty good. That's which is still pretty good. That's faster per agent than most people get faster per agent than most people get faster per agent than most people get from one agent in a cloud API. 16 users, from one agent in a cloud API. 16 users, from one agent in a cloud API. 16 users, 61, 32 users, 37, and 64 users. We're 61, 32 users, 37, and 64 users. We're 61, 32 users, 37, and 64 users. We're down to 20.6. Not amazing down there.

  20. down to 20.6. Not amazing down there. down to 20.6. Not amazing down there. Let's take a look at these other models Let's take a look at these other models Let's take a look at these other models here. And you can see where we stand. here. And you can see where we stand. here. And you can see where we stand. Quen 3.8 being the fastest, of course. Quen 3.8 being the fastest, of course. Quen 3.8 being the fastest, of course. GLM 5.3 Flash and Deepc4 Flash are about GLM 5.3 Flash and Deepc4 Flash are about GLM 5.3 Flash and Deepc4 Flash are about the same. Here are the numbers for 1. the same. Here are the numbers for 1. the same. Here are the numbers for 1. Here are the numbers for 4 8 and then Here are the numbers for 4 8 and then Here are the numbers for 4 8 and then ridiculous 64. Are you running 64 agents ridiculous 64. Are you running 64 agents ridiculous 64. Are you running 64 agents at the same time? Tell the truth. Come at the same time? Tell the truth. Come at the same time? Tell the truth. Come on. What's the most number of agents you on. What's the most number of agents you on. What's the most number of agents you ran at the same time? Put it down in the ran at the same time? Put it down in the ran at the same time? Put it down in the comments. And the machine is putting out comments. And the machine is putting out comments. And the machine is putting out 530 tokens per second aggregate at that 530 tokens per second aggregate at that 530 tokens per second aggregate at that point with a peak decode or token point with a peak decode or token point with a peak decode or token generation burst up to 2,880. Here are generation burst up to 2,880. Here are generation burst up to 2,880. Here are the numbers for one user, four users, the numbers for one user, four users, the numbers for one user, four users, eight users, and finally we got 64 eight users, and finally we got 64 eight users, and finally we got 64 users. There's that 530 from Quen 3.8 users. There's that 530 from Quen 3.8 users. There's that 530 from Quen 3.8 flash. Next, wo time to first token flash. Next, wo time to first token flash. Next, wo time to first token looks uh pretty interesting, right? looks uh pretty interesting, right? looks uh pretty interesting, right? Especially for GLM 5.2 there. Time to Especially for GLM 5.2 there. Time to Especially for GLM 5.2 there. Time to first token is that wait before anything first token is that wait before anything first token is that wait before anything at all shows up. At eight users, we're at all shows up. At eight users, we're at all shows up. At eight users, we're at about a second and a half for all at about a second and a half for all at about a second and a half for all three Flash models. 5.5 seconds for GLM three Flash models. 5.5 seconds for GLM three Flash models. 5.5 seconds for GLM 5.2. At 16, we're at 2 1/2 seconds and 5.2. At 16, we're at 2 1/2 seconds and 5.2. At 16, we're at 2 1/2 seconds and almost 10 seconds for GLM 5.2. And at almost 10 seconds for GLM 5.2. And at almost 10 seconds for GLM 5.2. And at 64, 64, 64, you know, we're getting to be in a place you know, we're getting to be in a place you know, we're getting to be in a place where you might not want to be, where you might not want to be, where you might not want to be, especially with GLM 5.2. Waiting 33 especially with GLM 5.2. Waiting 33 especially with GLM 5.2. Waiting 33 seconds for time to first open. GH, I'm seconds for time to first open. GH, I'm seconds for time to first open. GH, I'm sorry you have to go through that. We're sorry you have to go through that. We're sorry you have to go through that. We're so impatient these days, aren't we? This so impatient these days, aren't we? This so impatient these days, aren't we? This thing is literally writing our stuff for thing is literally writing our stuff for thing is literally writing our stuff for us. And come on, hurry up. Task masters.

  21. us. And come on, hurry up. Task masters. us. And come on, hurry up. Task masters. Thanks for staying through the whole Thanks for staying through the whole Thanks for staying through the whole video, by the way. And thanks to the video, by the way. And thanks to the video, by the way. And thanks to the members of the channel. I appreciate you members of the channel. I appreciate you members of the channel. I appreciate you all. Short version, with the right all. Short version, with the right all. Short version, with the right models, 16 developers on coding agents models, 16 developers on coding agents models, 16 developers on coding agents or one developer with 16 agents. Every or one developer with 16 agents. Every or one developer with 16 agents. Every one of these models is getting 60 tokens one of these models is getting 60 tokens one of these models is getting 60 tokens per second with a 2 and 1/2 second wait per second with a 2 and 1/2 second wait per second with a 2 and 1/2 second wait time from one box just sitting hopefully time from one box just sitting hopefully time from one box just sitting hopefully in another room. And for reference, in another room. And for reference, in another room. And for reference, storage review, which had the exact same storage review, which had the exact same storage review, which had the exact same machine as me. It got shipped to me from machine as me. It got shipped to me from machine as me. It got shipped to me from them. They did a write up on this. They them. They did a write up on this. They them. They did a write up on this. They ran Claude Code sessions against Miniax ran Claude Code sessions against Miniax ran Claude Code sessions against Miniax on the same machine, and they got 39 on the same machine, and they got 39 on the same machine, and they got 39 tokens per second per user at eight tokens per second per user at eight tokens per second per user at eight sessions, which is about what you get sessions, which is about what you get sessions, which is about what you get from Claude Opus through the API. So, from Claude Opus through the API. So, from Claude Opus through the API. So, would I buy one of these? Well, I mean, would I buy one of these? Well, I mean, would I buy one of these? Well, I mean, uh, I can't, but here's how I think of uh, I can't, but here's how I think of uh, I can't, but here's how I think of it. If you're a solo developer, this is it. If you're a solo developer, this is it. If you're a solo developer, this is the machine where you stop thinking the machine where you stop thinking the machine where you stop thinking about the model as a thing you wait for. about the model as a thing you wait for. about the model as a thing you wait for. Also, if you're a solo developer buying Also, if you're a solo developer buying Also, if you're a solo developer buying one of these, you better be making some one of these, you better be making some one of these, you better be making some decent money off of your gigs or renting decent money off of your gigs or renting decent money off of your gigs or renting it out or something. The main target it out or something. The main target it out or something. The main target audience here is obviously going to be audience here is obviously going to be audience here is obviously going to be businesses, small businesses, medium businesses, small businesses, medium businesses, small businesses, medium businesses, where there's teams of businesses, where there's teams of businesses, where there's teams of people working off of these things.

  22. people working off of these things. people working off of these things. Eight agents at 80 tokens per second Eight agents at 80 tokens per second Eight agents at 80 tokens per second each. A whole code base in the prompt in each. A whole code base in the prompt in each. A whole code base in the prompt in 14 seconds. It's kind of overkill for 14 seconds. It's kind of overkill for 14 seconds. It's kind of overkill for one person if you ask me, but it's one one person if you ask me, but it's one one person if you ask me, but it's one of two boxes that I've recently tested of two boxes that I've recently tested of two boxes that I've recently tested where I don't feel like I'm limited by where I don't feel like I'm limited by where I don't feel like I'm limited by how many agents I could run. And if how many agents I could run. And if how many agents I could run. And if you're one person that builds like a you're one person that builds like a you're one person that builds like a team of agents, hey, it's not that team of agents, hey, it's not that team of agents, hey, it's not that crazy. The other thing you should think crazy. The other thing you should think crazy. The other thing you should think about is this is really kind of new about is this is really kind of new about is this is really kind of new hardware still. So, you're going to need hardware still. So, you're going to need hardware still. So, you're going to need to get comfortable with nightly VLM to get comfortable with nightly VLM to get comfortable with nightly VLM builds. Maybe getting your hands a builds. Maybe getting your hands a builds. Maybe getting your hands a little bit dirty with the little bit dirty with the little bit dirty with the behindthe-scenes activities or if you behindthe-scenes activities or if you behindthe-scenes activities or if you just want to stand it up and run it. The just want to stand it up and run it. The just want to stand it up and run it. The recipes are out there. Nvidia has a recipes are out there. Nvidia has a recipes are out there. Nvidia has a bunch of recipes on their site. VLM has bunch of recipes on their site. VLM has bunch of recipes on their site. VLM has recipes which I've been using as well. recipes which I've been using as well. recipes which I've been using as well. Pretty handy. They don't always list the Pretty handy. They don't always list the Pretty handy. They don't always list the exact thing that you have running. For exact thing that you have running. For exact thing that you have running. For example, uh here is a recipe for RTX Pro example, uh here is a recipe for RTX Pro example, uh here is a recipe for RTX Pro 6000, but four of them. So, you just 6000, but four of them. So, you just 6000, but four of them. So, you just need to modify it to your own size. For need to modify it to your own size. For need to modify it to your own size. For example, tensor parallel size, you'd set example, tensor parallel size, you'd set example, tensor parallel size, you'd set that to eight. Or if it's a small enough that to eight. Or if it's a small enough that to eight. Or if it's a small enough model, you can run a couple of instances model, you can run a couple of instances model, you can run a couple of instances of it. Like I mentioned earlier, I of it. Like I mentioned earlier, I of it. Like I mentioned earlier, I recently did that build with four RTX recently did that build with four RTX recently did that build with four RTX Pro 6000s, and you can watch that right Pro 6000s, and you can watch that right Pro 6000s, and you can watch that right over here next. Thanks for watching, and over here next. Thanks for watching, and over here next. Thanks for watching, and I'll see you next time.

Summary

The main theme is the introduction of the Camino Grando, a powerful liquid-cooled workstation featuring eight RTX Pro 6000 GPUs with 768GB of VRAM. This machine is designed to handle massive local LLMs and complex multi-agent development workflows without relying on the cloud. The practical takeaway is that for developers, especially teams, who need to run advanced AI models and concurrent agent tasks locally, such high-performance, albeit expensive, hardware is becoming essential.

View original episode ↗