192GB of VRAM in One PC… The Cheap Way
Read full transcript 12 segments
-
When you want to run large language When you want to run large language models locally, it comes down to VRAM. models locally, it comes down to VRAM. models locally, it comes down to VRAM. How much VRAM do you have? More of it How much VRAM do you have? More of it How much VRAM do you have? More of it means you can fit bigger models, more means you can fit bigger models, more means you can fit bigger models, more context, more users at once. The context, more users at once. The context, more users at once. The expensive path is Nvidia. This RTX Pro expensive path is Nvidia. This RTX Pro expensive path is Nvidia. This RTX Pro 4000 GPU is $2,500 and it gives you 24 4000 GPU is $2,500 and it gives you 24 4000 GPU is $2,500 and it gives you 24 gigs of VRAM. But if you want to go the gigs of VRAM. But if you want to go the gigs of VRAM. But if you want to go the cheaper route, you have GPUs from Intel, cheaper route, you have GPUs from Intel, cheaper route, you have GPUs from Intel, the B series. The B60 also has 24 gigs the B series. The B60 also has 24 gigs the B series. The B60 also has 24 gigs of VRAM, but it's 650 bucks. Now, if you of VRAM, but it's 650 bucks. Now, if you of VRAM, but it's 650 bucks. Now, if you wanted 48 gigabytes of VRAM, should you wanted 48 gigabytes of VRAM, should you wanted 48 gigabytes of VRAM, should you get two B60s or should you get one card get two B60s or should you get one card get two B60s or should you get one card with two GPUs on it? And also, is there with two GPUs on it? And also, is there with two GPUs on it? And also, is there a performance benefit to that? And yes, a performance benefit to that? And yes, a performance benefit to that? And yes, this one got a little different, so I this one got a little different, so I this one got a little different, so I had to call someone way smarter than me. had to call someone way smarter than me. had to call someone way smarter than me. My phone battery has basically become a My phone battery has basically become a My phone battery has basically become a two-part routine. I use it hard for half two-part routine. I use it hard for half two-part routine. I use it hard for half the day, then I find a charger to get the day, then I find a charger to get the day, then I find a charger to get through the rest of the day. And if through the rest of the day. And if through the rest of the day. And if you're bouncing between office, coffee you're bouncing between office, coffee you're bouncing between office, coffee shops, airports, or just a long day out, shops, airports, or just a long day out, shops, airports, or just a long day out, you know the feeling. This segment is you know the feeling. This segment is you know the feeling. This segment is sponsored by Basis and their Pico Go sponsored by Basis and their Pico Go sponsored by Basis and their Pico Go series. The Pico Go Air is a light daily series. The Pico Go Air is a light daily series. The Pico Go Air is a light daily option, only 27 in thick with NFC smart option, only 27 in thick with NFC smart option, only 27 in thick with NFC smart battery monitoring through the Basis app battery monitoring through the Basis app battery monitoring through the Basis app and Basis says it runs 18° Fahrenheit and Basis says it runs 18° Fahrenheit and Basis says it runs 18° Fahrenheit cooler for close to skin comfort. For my cooler for close to skin comfort. For my cooler for close to skin comfort. For my routine, I'd grab the Pico Gold AM52. I routine, I'd grab the Pico Gold AM52. I routine, I'd grab the Pico Gold AM52. I already carry a battery every day, but already carry a battery every day, but already carry a battery every day, but 10,000 mAh gives me more breathing room 10,000 mAh gives me more breathing room 10,000 mAh gives me more breathing room with a built-in cable and charging for with a built-in cable and charging for with a built-in cable and charging for up to three devices. Air for everyday up to three devices. Air for everyday up to three devices. Air for everyday carry, AM52 for longer days. Check carry, AM52 for longer days. Check carry, AM52 for longer days. Check compatibility, especially if you want to compatibility, especially if you want to compatibility, especially if you want to use that chi 2.25 watt ultraast wireless
-
use that chi 2.25 watt ultraast wireless use that chi 2.25 watt ultraast wireless charging, and hit the link in the charging, and hit the link in the charging, and hit the link in the description to check out the basis Pico description to check out the basis Pico description to check out the basis Pico series. This is something interesting. series. This is something interesting. series. This is something interesting. Excuse me. And the reason is because Excuse me. And the reason is because Excuse me. And the reason is because it's only one card and it's like two of it's only one card and it's like two of it's only one card and it's like two of these. So yeah, these are both B60s from these. So yeah, these are both B60s from these. So yeah, these are both B60s from Intel and I did a whole episode on Intel and I did a whole episode on Intel and I did a whole episode on those. But this is a B60 duel those. But this is a B60 duel those. But this is a B60 duel frankly bag. What's really interesting frankly bag. What's really interesting frankly bag. What's really interesting about this is that it's the same size as about this is that it's the same size as about this is that it's the same size as one of these. And Maxon, they're the one of these. And Maxon, they're the one of these. And Maxon, they're the only manufacturer that's making these. only manufacturer that's making these. only manufacturer that's making these. These are the two GPUs. You can see them These are the two GPUs. You can see them These are the two GPUs. You can see them through the back. So, I'm going to test through the back. So, I'm going to test through the back. So, I'm going to test one against two to see what we can do. one against two to see what we can do. one against two to see what we can do. Did not mean for that to rhyme. And I'm Did not mean for that to rhyme. And I'm Did not mean for that to rhyme. And I'm also going to do some tests against the also going to do some tests against the also going to do some tests against the B70s, too. Put that against AMD and B70s, too. Put that against AMD and B70s, too. Put that against AMD and Nvidia. And we have this picture. Intel Nvidia. And we have this picture. Intel Nvidia. And we have this picture. Intel sits right down here with a B60. They sits right down here with a B60. They sits right down here with a B60. They also got the B70 for about $1,000 with also got the B70 for about $1,000 with also got the B70 for about $1,000 with 32 gigs. And up here, we've got 48 GB 32 gigs. And up here, we've got 48 GB 32 gigs. And up here, we've got 48 GB for two B60s. Or we got this Mac Sun, for two B60s. Or we got this Mac Sun, for two B60s. Or we got this Mac Sun, which is actually more expensive than which is actually more expensive than which is actually more expensive than buying two B60s. So, why why would I get buying two B60s. So, why why would I get buying two B60s. So, why why would I get one of these? Oh, it's heavier. Now, one of these? Oh, it's heavier. Now, one of these? Oh, it's heavier. Now, you're looking at the RTX Pro 5000, 48 you're looking at the RTX Pro 5000, 48 you're looking at the RTX Pro 5000, 48 gigs of memory, and we're at $6,000. If gigs of memory, and we're at $6,000. If gigs of memory, and we're at $6,000. If you want to get 96, you're over here you want to get 96, you're over here you want to get 96, you're over here with a Pro 6000. Kind of in a class on with a Pro 6000. Kind of in a class on with a Pro 6000. Kind of in a class on its own. So, Intel's pitch here is more its own. So, Intel's pitch here is more its own. So, Intel's pitch here is more VRAM per dollar. But cheap VRAM and VRAM per dollar. But cheap VRAM and VRAM per dollar. But cheap VRAM and useful VRAM are different things. So, useful VRAM are different things. So, useful VRAM are different things. So, let's run some things inside this beast let's run some things inside this beast let's run some things inside this beast over here. I ended up running small over here. I ended up running small over here. I ended up running small models, medium models, and larger models, medium models, and larger models, medium models, and larger models, not huge models. So, inside this
-
models, not huge models. So, inside this models, not huge models. So, inside this beast here, which is a fractal case, and beast here, which is a fractal case, and beast here, which is a fractal case, and it's got an Asus Pro board on it, Intel it's got an Asus Pro board on it, Intel it's got an Asus Pro board on it, Intel Xeon CPU. As usual, I like to start with Xeon CPU. As usual, I like to start with Xeon CPU. As usual, I like to start with my small Quen 3 4 billion parameter my small Quen 3 4 billion parameter my small Quen 3 4 billion parameter model. And I know some of you will say, model. And I know some of you will say, model. And I know some of you will say, "Ah, that's just too small." And you're "Ah, that's just too small." And you're "Ah, that's just too small." And you're right, for some things, it is too small, right, for some things, it is too small, right, for some things, it is too small, including what we're trying to do here, including what we're trying to do here, including what we're trying to do here, but it's kind of a baseline. Here's one but it's kind of a baseline. Here's one but it's kind of a baseline. Here's one B70 generating tokens at 92 tokens per B70 generating tokens at 92 tokens per B70 generating tokens at 92 tokens per second. And by the way, this is the FP8 second. And by the way, this is the FP8 second. And by the way, this is the FP8 version of that model. And Intel does version of that model. And Intel does version of that model. And Intel does have support for the FP8 on the hardware have support for the FP8 on the hardware have support for the FP8 on the hardware level. And the B70 beats running two level. And the B70 beats running two level. And the B70 beats running two B60s working together in tensor B60s working together in tensor B60s working together in tensor parallelism. Notice also that one B60 parallelism. Notice also that one B60 parallelism. Notice also that one B60 runs a little bit slower than two B60s. runs a little bit slower than two B60s. runs a little bit slower than two B60s. That's because we are using tensor That's because we are using tensor That's because we are using tensor parallelism. So there's fast parallelism. So there's fast parallelism. So there's fast communication and low latency between communication and low latency between communication and low latency between the two GPUs. But that's not always the the two GPUs. But that's not always the the two GPUs. But that's not always the case. It depends on the model. So for case. It depends on the model. So for case. It depends on the model. So for example here quen 3.5 27 billion integer example here quen 3.5 27 billion integer example here quen 3.5 27 billion integer 4 quantization it's a newer model and we 4 quantization it's a newer model and we 4 quantization it's a newer model and we have slightly different architecture have slightly different architecture have slightly different architecture here 28.3 tokens per second on the B70 here 28.3 tokens per second on the B70 here 28.3 tokens per second on the B70 and now two B60s actually beat it one and now two B60s actually beat it one and now two B60s actually beat it one B60 can't run it at all because this B60 can't run it at all because this B60 can't run it at all because this model is larger even with 4bit model is larger even with 4bit model is larger even with 4bit quantization the model is almost 30 gigs quantization the model is almost 30 gigs quantization the model is almost 30 gigs in size and then you also need a little in size and then you also need a little in size and then you also need a little bit of headroom which this does not bit of headroom which this does not bit of headroom which this does not have. Finally this uh oldie but a have. Finally this uh oldie but a have. Finally this uh oldie but a goodie. Well, maybe not so great goodie. Well, maybe not so great goodie. Well, maybe not so great anymore, but I used to do this in my anymore, but I used to do this in my anymore, but I used to do this in my test, so I included it here. It's a 30 test, so I included it here. It's a 30 test, so I included it here. It's a 30 billion parameter mixture of experts billion parameter mixture of experts billion parameter mixture of experts integer for quantization. B70 45 tokens integer for quantization. B70 45 tokens integer for quantization. B70 45 tokens per second. B60 44.5 very close there.
-
per second. B60 44.5 very close there. per second. B60 44.5 very close there. And then two B60s go down in uh speed And then two B60s go down in uh speed And then two B60s go down in uh speed here. Now, I did a separate video on the here. Now, I did a separate video on the here. Now, I did a separate video on the B70s and the B60s and the B50 last year. B70s and the B60s and the B50 last year. B70s and the B60s and the B50 last year. So, quick reality check on the software So, quick reality check on the software So, quick reality check on the software because it kind of shapes these numbers. because it kind of shapes these numbers. because it kind of shapes these numbers. Intel's hardware is here is good. The Intel's hardware is here is good. The Intel's hardware is here is good. The software is still catching up to where software is still catching up to where software is still catching up to where Nvidia is. Nvidia is way further ahead Nvidia is. Nvidia is way further ahead Nvidia is. Nvidia is way further ahead in software and that's what allows them in software and that's what allows them in software and that's what allows them to charge these kinds of prices. The to charge these kinds of prices. The to charge these kinds of prices. The software stack I'm running is LLM Scaler software stack I'm running is LLM Scaler software stack I'm running is LLM Scaler and it's a fork of VLM that's probably and it's a fork of VLM that's probably and it's a fork of VLM that's probably about a month behind where VLM currently about a month behind where VLM currently about a month behind where VLM currently is. And the supported models are listed is. And the supported models are listed is. And the supported models are listed here in this GitHub repository. And here in this GitHub repository. And here in this GitHub repository. And you'll see that some of the newest you'll see that some of the newest you'll see that some of the newest models that are actually supported here models that are actually supported here models that are actually supported here probably a month or two old which is probably a month or two old which is probably a month or two old which is fine. You're not going to be running the fine. You're not going to be running the fine. You're not going to be running the break neck state-of-the-art stuff, but break neck state-of-the-art stuff, but break neck state-of-the-art stuff, but it's not going to be bleeding edge. it's not going to be bleeding edge. it's not going to be bleeding edge. Bleeding edge. Break neck. What is this? Bleeding edge. Break neck. What is this? Bleeding edge. Break neck. What is this? What? I'm not a violent channel. Stop What? I'm not a violent channel. Stop What? I'm not a violent channel. Stop all these violent terms for latest and all these violent terms for latest and all these violent terms for latest and greatest. Let's call it latest and greatest. Let's call it latest and greatest. Let's call it latest and greatest. Okay, that's better. However, greatest. Okay, that's better. However, greatest. Okay, that's better. However, you'll notice if you watch my video you'll notice if you watch my video you'll notice if you watch my video about the B70s from maybe a month and a about the B70s from maybe a month and a about the B70s from maybe a month and a half ago, the Quen 3.5 wasn't even there half ago, the Quen 3.5 wasn't even there half ago, the Quen 3.5 wasn't even there yet. So, they're definitely working on yet. So, they're definitely working on yet. So, they're definitely working on this all the time, making stuff happen, this all the time, making stuff happen, this all the time, making stuff happen, which is great. Progress being made.
-
which is great. Progress being made. which is great. Progress being made. Finally, it was time to check the Mac Finally, it was time to check the Mac Finally, it was time to check the Mac Sun card. I seated the card. I plugged Sun card. I seated the card. I plugged Sun card. I seated the card. I plugged it in, powered on the machine, and it in, powered on the machine, and it in, powered on the machine, and nothing. Card was not detected. Tool nothing. Card was not detected. Tool nothing. Card was not detected. Tool says no device discovered, and the PCI says no device discovered, and the PCI says no device discovered, and the PCI list was empty. The card was plugged in, list was empty. The card was plugged in, list was empty. The card was plugged in, I swear. But the computer was saying the I swear. But the computer was saying the I swear. But the computer was saying the card doesn't exist. So, this was my card doesn't exist. So, this was my card doesn't exist. So, this was my first time using a single card with two first time using a single card with two first time using a single card with two GPUs on it. And that has to be treated GPUs on it. And that has to be treated GPUs on it. And that has to be treated slightly differently because in this slightly differently because in this slightly differently because in this case, it's a switchless card, which case, it's a switchless card, which case, it's a switchless card, which means the GPUs can't talk directly to means the GPUs can't talk directly to means the GPUs can't talk directly to each other. they have to talk through each other. they have to talk through each other. they have to talk through the CPU. They have to talk through the the CPU. They have to talk through the the CPU. They have to talk through the PCI lane that they're connected to. Each PCI lane that they're connected to. Each PCI lane that they're connected to. Each one of them is kind of treated one of them is kind of treated one of them is kind of treated separately. So the motherboard itself is separately. So the motherboard itself is separately. So the motherboard itself is responsible for splitting that slot, responsible for splitting that slot, responsible for splitting that slot, that PCI slot in half or one per GPU. that PCI slot in half or one per GPU. that PCI slot in half or one per GPU. And that's called bifurcation. And if And that's called bifurcation. And if And that's called bifurcation. And if the board, the motherboard is not set up the board, the motherboard is not set up the board, the motherboard is not set up for it, the GPU is going to be pretty for it, the GPU is going to be pretty for it, the GPU is going to be pretty much invisible. So I couldn't fix that much invisible. So I couldn't fix that much invisible. So I couldn't fix that from the operating system. I needed to from the operating system. I needed to from the operating system. I needed to go to the BIOS. Luckily, a motherboard go to the BIOS. Luckily, a motherboard go to the BIOS. Luckily, a motherboard like this has dual Ethernet ports for like this has dual Ethernet ports for like this has dual Ethernet ports for connectivity into the OS and a separate connectivity into the OS and a separate connectivity into the OS and a separate Ethernet port for management, which Ethernet port for management, which Ethernet port for management, which means you can actually manage the means you can actually manage the means you can actually manage the computer as a server remotely. You don't computer as a server remotely. You don't computer as a server remotely. You don't even need to plug a monitor or a even need to plug a monitor or a even need to plug a monitor or a keyboard in here. So, I can hop into the keyboard in here. So, I can hop into the keyboard in here. So, I can hop into the BIOS through remote connection for my BIOS through remote connection for my BIOS through remote connection for my Mac. Finally, I found the PCI settings Mac. Finally, I found the PCI settings Mac. Finally, I found the PCI settings and slot one was where I put this card and slot one was where I put this card and slot one was where I put this card and I was able to split that PCI card and I was able to split that PCI card and I was able to split that PCI card into two. Rebooted, watched the screen, into two. Rebooted, watched the screen, into two. Rebooted, watched the screen, and it came up. There they are. both of and it came up. There they are. both of and it came up. There they are. both of them, two B60s, full Gen 5, full them, two B60s, full Gen 5, full them, two B60s, full Gen 5, full bandwidth, 24 gigs each, and it's alive.
-
bandwidth, 24 gigs each, and it's alive. bandwidth, 24 gigs each, and it's alive. Now, the question that came up in my Now, the question that came up in my Now, the question that came up in my mind immediately is we now have two GPUs mind immediately is we now have two GPUs mind immediately is we now have two GPUs sharing one slot. Does that cost you sharing one slot. Does that cost you sharing one slot. Does that cost you anything versus two real cards that are anything versus two real cards that are anything versus two real cards that are side by side, each one of them having side by side, each one of them having side by side, each one of them having their own PCI slot? So, here's Quen 34B, their own PCI slot? So, here's Quen 34B, their own PCI slot? So, here's Quen 34B, and it's looking like the two B60s. All and it's looking like the two B60s. All and it's looking like the two B60s. All right, let's try the 27 billion right, let's try the 27 billion right, let's try the 27 billion parameter model. Okay, 26.7 tokens per parameter model. Okay, 26.7 tokens per parameter model. Okay, 26.7 tokens per second average. 25.3 26.6 second average. 25.3 26.6 second average. 25.3 26.6 and we're at prompt processing is 1789 and we're at prompt processing is 1789 and we're at prompt processing is 1789 tokens per second. Oh, up to 2100 token tokens per second. Oh, up to 2100 token tokens per second. Oh, up to 2100 token per second now. 2200 tokens per second. per second now. 2200 tokens per second. per second now. 2200 tokens per second. Now, I know I didn't talk about prompt Now, I know I didn't talk about prompt Now, I know I didn't talk about prompt processing. I'll cover that in another processing. I'll cover that in another processing. I'll cover that in another video, so make sure you don't miss that. video, so make sure you don't miss that. video, so make sure you don't miss that. But just wanted to mention it for this But just wanted to mention it for this But just wanted to mention it for this one. Overall, Max Sun, the dual GPU card one. Overall, Max Sun, the dual GPU card one. Overall, Max Sun, the dual GPU card against two B60s is pretty much against two B60s is pretty much against two B60s is pretty much identical. Sure, there's some margin for identical. Sure, there's some margin for identical. Sure, there's some margin for error, like in this case right here with error, like in this case right here with error, like in this case right here with 4B model FP8, but pretty much every 4B model FP8, but pretty much every 4B model FP8, but pretty much every model squeezing into two GPUs on board, model squeezing into two GPUs on board, model squeezing into two GPUs on board, splitting one slot cost you nothing.
-
splitting one slot cost you nothing. splitting one slot cost you nothing. Same speed as two cards, one slot, Same speed as two cards, one slot, Same speed as two cards, one slot, though. Now, tokens per second isn't though. Now, tokens per second isn't though. Now, tokens per second isn't everything. What else can you do with everything. What else can you do with everything. What else can you do with this? I ran LTX2. This is the video this? I ran LTX2. This is the video this? I ran LTX2. This is the video generation model in Comfy UI. This is a generation model in Comfy UI. This is a generation model in Comfy UI. This is a texttovideo pipeline, and I did texttovideo pipeline, and I did texttovideo pipeline, and I did specifically request Arnold specifically request Arnold specifically request Arnold Schwarzenegger saying, "I'll be back. Schwarzenegger saying, "I'll be back. Schwarzenegger saying, "I'll be back. I'll be back. I'll be back. I'll be back. >> This time it goes the other way. The B70 >> This time it goes the other way. The B70 >> This time it goes the other way. The B70 wins. Of course, we know the B70 is wins. Of course, we know the B70 is wins. Of course, we know the B70 is going to win versus two B60s or one dual going to win versus two B60s or one dual going to win versus two B60s or one dual B60 because video generation is B60 because video generation is B60 because video generation is computebound. It's uh not memory bound, computebound. It's uh not memory bound, computebound. It's uh not memory bound, raw horsepower. And the B70 has more raw horsepower. And the B70 has more raw horsepower. And the B70 has more horsepower than a B60. Also, by the way, horsepower than a B60. Also, by the way, horsepower than a B60. Also, by the way, LTX is uh bound to one GPU. So, in both LTX is uh bound to one GPU. So, in both LTX is uh bound to one GPU. So, in both cases, it ran on only one B60 and only cases, it ran on only one B60 and only cases, it ran on only one B60 and only one B60 out of the dual card, which is one B60 out of the dual card, which is one B60 out of the dual card, which is something you can do. You can use those something you can do. You can use those something you can do. You can use those individually if you want to. You can individually if you want to. You can individually if you want to. You can even have different models running on even have different models running on even have different models running on each B60 if you wanted to. So, in this each B60 if you wanted to. So, in this each B60 if you wanted to. So, in this case, all that VRAM that we had didn't case, all that VRAM that we had didn't case, all that VRAM that we had didn't really help much. LTX still runs and LTX really help much. LTX still runs and LTX really help much. LTX still runs and LTX can run on smaller hardware, but you can can run on smaller hardware, but you can can run on smaller hardware, but you can see the difference in the time that it see the difference in the time that it see the difference in the time that it took to generate. And you can see that took to generate. And you can see that took to generate. And you can see that we're very close on one B60 versus the we're very close on one B60 versus the we're very close on one B60 versus the one GPU B60. I'm going to confuse myself one GPU B60. I'm going to confuse myself one GPU B60. I'm going to confuse myself and I'm sounding confusing. If I say and I'm sounding confusing. If I say and I'm sounding confusing. If I say B60, I mean one of these. And if I say B60, I mean one of these. And if I say B60, I mean one of these. And if I say the Max Sun, I'm going to mean that.
-
the Max Sun, I'm going to mean that. the Max Sun, I'm going to mean that. Okay, the duel. Okay, the duel. Okay, the duel. All right, we can see that the time the All right, we can see that the time the All right, we can see that the time the time it took to generate that video is time it took to generate that video is time it took to generate that video is exactly pretty much exactly the same on exactly pretty much exactly the same on exactly pretty much exactly the same on both of those, which is 226.7 and 224.6 both of those, which is 226.7 and 224.6 both of those, which is 226.7 and 224.6 six versus 167 seconds that it took to six versus 167 seconds that it took to six versus 167 seconds that it took to generate on the B70. But if it's a tie generate on the B70. But if it's a tie generate on the B70. But if it's a tie on speed, the decision comes down to on speed, the decision comes down to on speed, the decision comes down to what speed chart doesn't actually show what speed chart doesn't actually show what speed chart doesn't actually show like power. Now imagine for a second you like power. Now imagine for a second you like power. Now imagine for a second you have to power two of these B60s or one have to power two of these B60s or one have to power two of these B60s or one of those Max Sun cards. Which one do you of those Max Sun cards. Which one do you of those Max Sun cards. Which one do you think is going to take more power? Here think is going to take more power? Here think is going to take more power? Here we go. Same workload and the Max Sun we go. Same workload and the Max Sun we go. Same workload and the Max Sun actually sips a little bit less power. actually sips a little bit less power. actually sips a little bit less power. about 129 watts versus 134 with the dual about 129 watts versus 134 with the dual about 129 watts versus 134 with the dual B60s. Oh my god, I said it again. With B60s. Oh my god, I said it again. With B60s. Oh my god, I said it again. With two B60 GPU cards. So, it's pretty two B60 GPU cards. So, it's pretty two B60 GPU cards. So, it's pretty modest difference, but it's pretty modest difference, but it's pretty modest difference, but it's pretty consistent. You will be using just a consistent. You will be using just a consistent. You will be using just a little bit less power with one card. It little bit less power with one card. It little bit less power with one card. It may not even make that much of a may not even make that much of a may not even make that much of a difference for you. But here's the real difference for you. But here's the real difference for you. But here's the real reason you'd want to go with a Max Sun reason you'd want to go with a Max Sun reason you'd want to go with a Max Sun card and why you'd want to actually card and why you'd want to actually card and why you'd want to actually spend more on buying one of those than spend more on buying one of those than spend more on buying one of those than two B60s. Knowledge.
-
two B60s. Knowledge. two B60s. Knowledge. >> Huh? Uh, no, that's wrong video. Um, >> Huh? Uh, no, that's wrong video. Um, >> Huh? Uh, no, that's wrong video. Um, density. Okay, that was a stupid joke. density. Okay, that was a stupid joke. density. Okay, that was a stupid joke. Two GPUs in one slot means you can fill Two GPUs in one slot means you can fill Two GPUs in one slot means you can fill up a four slot machine like this one and up a four slot machine like this one and up a four slot machine like this one and get even more VRAM total. For example, get even more VRAM total. For example, get even more VRAM total. For example, recently I made a video with four B60 recently I made a video with four B60 recently I made a video with four B60 GPUs in this chassis on this GPUs in this chassis on this GPUs in this chassis on this motherboard. And because each one of motherboard. And because each one of motherboard. And because each one of these is a 2unit GPU and that these is a 2unit GPU and that these is a 2unit GPU and that motherboard has seven slots, you can motherboard has seven slots, you can motherboard has seven slots, you can only physically fit four of these in only physically fit four of these in only physically fit four of these in there. uh unless you do uh ret timers there. uh unless you do uh ret timers there. uh unless you do uh ret timers and riser cables and yeah, you can do and riser cables and yeah, you can do and riser cables and yeah, you can do that. But I mean inside the box, let's that. But I mean inside the box, let's that. But I mean inside the box, let's set that as our limit. Four of those set that as our limit. Four of those set that as our limit. Four of those will be 96 GB of VRAM. Four of these, will be 96 GB of VRAM. Four of these, will be 96 GB of VRAM. Four of these, which are pretty much the same size, which are pretty much the same size, which are pretty much the same size, these are also two space GPUs. Four of these are also two space GPUs. Four of these are also two space GPUs. Four of these will give you 128 VRAM. The Maxon, these will give you 128 VRAM. The Maxon, these will give you 128 VRAM. The Maxon, if you take four of those, it'll give if you take four of those, it'll give if you take four of those, it'll give you 192. So that's the argument. It's you 192. So that's the argument. It's you 192. So that's the argument. It's not about the speed. It's about fitting not about the speed. It's about fitting not about the speed. It's about fitting those workloads that the cheaper cards those workloads that the cheaper cards those workloads that the cheaper cards can't touch. And at a price nowhere near can't touch. And at a price nowhere near can't touch. And at a price nowhere near Nvidia's. 1750 for this Mac Sun card Nvidia's. 1750 for this Mac Sun card Nvidia's. 1750 for this Mac Sun card time 4. $7,000 for 192 GB of VRAM of GPU time 4. $7,000 for 192 GB of VRAM of GPU time 4. $7,000 for 192 GB of VRAM of GPU power. $7,000 versus $24,000 if I were power. $7,000 versus $24,000 if I were power. $7,000 versus $24,000 if I were to use Nvidia RTX Pro 5000s for the same to use Nvidia RTX Pro 5000s for the same to use Nvidia RTX Pro 5000s for the same amount of VRAM. Now, Nvidia, of course, amount of VRAM. Now, Nvidia, of course, amount of VRAM. Now, Nvidia, of course, as I've said before, has their CUDA as I've said before, has their CUDA as I've said before, has their CUDA stack, and they have GDDR7, which is the stack, and they have GDDR7, which is the stack, and they have GDDR7, which is the newer generation of memory. So, yeah, newer generation of memory. So, yeah, newer generation of memory. So, yeah, it's going to be faster. They also have it's going to be faster. They also have it's going to be faster. They also have different memory bandwidths. Nvidia is
-
different memory bandwidths. Nvidia is different memory bandwidths. Nvidia is much faster in memory bandwidth. So, much faster in memory bandwidth. So, much faster in memory bandwidth. So, it's going to be faster in token it's going to be faster in token it's going to be faster in token generation. And I've actually done the generation. And I've actually done the generation. And I've actually done the comparisons to these uh cards before in comparisons to these uh cards before in comparisons to these uh cards before in the previous videos. You can go check the previous videos. You can go check the previous videos. You can go check those out. I'll link them down below. those out. I'll link them down below. those out. I'll link them down below. But, I only have one of these, and I've But, I only have one of these, and I've But, I only have one of these, and I've had to bifurcate that PCI slot. and had to bifurcate that PCI slot. and had to bifurcate that PCI slot. and consumer motherboards, not all of them consumer motherboards, not all of them consumer motherboards, not all of them will support bifurcation. So, you have will support bifurcation. So, you have will support bifurcation. So, you have to be careful. Just a warning for you. to be careful. Just a warning for you. to be careful. Just a warning for you. And I don't know if even in this pro And I don't know if even in this pro And I don't know if even in this pro motherboard if every single lane will motherboard if every single lane will motherboard if every single lane will bifurcate. And I can put four of these bifurcate. And I can put four of these bifurcate. And I can put four of these in there. So, that's what I'm going to in there. So, that's what I'm going to in there. So, that's what I'm going to call an expert. Hi, Wendle. The famous call an expert. Hi, Wendle. The famous call an expert. Hi, Wendle. The famous office. office. office. >> You just need to buy a giant house with >> You just need to buy a giant house with >> You just need to buy a giant house with a workshop. a workshop. a workshop. >> Yes. Yes. And I need to make a lot more >> Yes. Yes. And I need to make a lot more >> Yes. Yes. And I need to make a lot more money to buy the giant house in money to buy the giant house in money to buy the giant house in Washington DC. Washington DC. Washington DC. Oh, that's Wendle from Level One Text. Oh, that's Wendle from Level One Text. Oh, that's Wendle from Level One Text. And I called him because he has the And I called him because he has the And I called him because he has the exact same machine in a video not too exact same machine in a video not too exact same machine in a video not too long ago. But did you have this inside long ago. But did you have this inside long ago. But did you have this inside one of them? one of them? one of them? >> No, I do not have any Macs on anything. >> No, I do not have any Macs on anything. >> No, I do not have any Macs on anything. >> Those are pretty cool. So my question to >> Those are pretty cool. So my question to >> Those are pretty cool. So my question to you, we have this motherboard. We have you, we have this motherboard. We have you, we have this motherboard. We have seven slots in there, but when you have seven slots in there, but when you have seven slots in there, but when you have two slot cards in there, you can put two slot cards in there, you can put two slot cards in there, you can put four. Do you think this board will four. Do you think this board will four. Do you think this board will support that?
-
support that? support that? >> It'd be easy to test because you can >> It'd be easy to test because you can >> It'd be easy to test because you can take a one card and just move it. I take a one card and just move it. I take a one card and just move it. I could do that, but I think asking a could do that, but I think asking a could do that, but I think asking a Wendle would be way easier than me doing Wendle would be way easier than me doing Wendle would be way easier than me doing that and rebooting this thing multiple that and rebooting this thing multiple that and rebooting this thing multiple times. times. times. >> I think you still got all the lanes, so >> I think you still got all the lanes, so >> I think you still got all the lanes, so double check the manual and the manual double check the manual and the manual double check the manual and the manual will tell you for sure. But bifurcating will tell you for sure. But bifurcating will tell you for sure. But bifurcating a slot, even down to like X4, X4, X4, a slot, even down to like X4, X4, X4, a slot, even down to like X4, X4, X4, X4, generally pretty supported on all X4, generally pretty supported on all X4, generally pretty supported on all all boards these days. all boards these days. all boards these days. >> Well, that's good to know. And it >> Well, that's good to know. And it >> Well, that's good to know. And it depends on the CPU, too, right? The only depends on the CPU, too, right? The only depends on the CPU, too, right? The only thing that's been squirly on those Intel thing that's been squirly on those Intel thing that's been squirly on those Intel Zeons has been MMU separation cuz Zeons has been MMU separation cuz Zeons has been MMU separation cuz sometimes the immu group one to MU group sometimes the immu group one to MU group sometimes the immu group one to MU group two, you don't get the full bandwidth, two, you don't get the full bandwidth, two, you don't get the full bandwidth, but Zeon 6, I think they've solved a lot but Zeon 6, I think they've solved a lot but Zeon 6, I think they've solved a lot of those issues. of those issues. of those issues. >> Nice. >> Nice. >> Nice. All right, Wendle, take care. Thank you. All right, Wendle, take care. Thank you. All right, Wendle, take care. Thank you. >> Take care. >> Take care. >> Take care. >> As a side note, Wendell told me about >> As a side note, Wendell told me about >> As a side note, Wendell told me about something really cool that I can't talk something really cool that I can't talk something really cool that I can't talk about yet, but he'll have on his channel about yet, but he'll have on his channel about yet, but he'll have on his channel pretty soon. So, make sure you subscribe pretty soon. So, make sure you subscribe pretty soon. So, make sure you subscribe to Level One Tech channel so you don't to Level One Tech channel so you don't to Level One Tech channel so you don't miss it. I'll leave a link to it down miss it. I'll leave a link to it down miss it. I'll leave a link to it down below. So, if I were you, mostly below. So, if I were you, mostly below. So, if I were you, mostly chatting, small models, one strong card chatting, small models, one strong card chatting, small models, one strong card like a B70 is simpler and snappier. You like a B70 is simpler and snappier. You like a B70 is simpler and snappier. You can put this into any motherboard pretty can put this into any motherboard pretty can put this into any motherboard pretty much. You don't need a pro motherboard.
-
much. You don't need a pro motherboard. much. You don't need a pro motherboard. But if you need 27 billion parameters But if you need 27 billion parameters But if you need 27 billion parameters and up, big context, lots of users on a and up, big context, lots of users on a and up, big context, lots of users on a budget, then you might need to double up budget, then you might need to double up budget, then you might need to double up on a B70 or double up on a B60 or if you on a B70 or double up on a B60 or if you on a B70 or double up on a B60 or if you have limited space, the Max on card. have limited space, the Max on card. have limited space, the Max on card. This dual one card system is a smart way This dual one card system is a smart way This dual one card system is a smart way to get there. if your board can do it. to get there. if your board can do it. to get there. if your board can do it. That's pretty much the whole point of That's pretty much the whole point of That's pretty much the whole point of this card. And credit to Maxon. Two GPUs this card. And credit to Maxon. Two GPUs this card. And credit to Maxon. Two GPUs on one board, one slot, and zero on one board, one slot, and zero on one board, one slot, and zero penalty. That's really clever and at the penalty. That's really clever and at the penalty. That's really clever and at the price that nobody's doing it. But let me price that nobody's doing it. But let me price that nobody's doing it. But let me dream for a second. Now, can you do it dream for a second. Now, can you do it dream for a second. Now, can you do it with a dual B70? Each B70 is 32 gigs. with a dual B70? Each B70 is 32 gigs. with a dual B70? Each B70 is 32 gigs. So, a dual B70, 64 gigs on one board. So, a dual B70, 64 gigs on one board. So, a dual B70, 64 gigs on one board. Then fill four slots and that's 256 gigs Then fill four slots and that's 256 gigs Then fill four slots and that's 256 gigs of VRAM, eight GPUs in one machine. of VRAM, eight GPUs in one machine. of VRAM, eight GPUs in one machine. Still nowhere near Nvidia pricing. Still nowhere near Nvidia pricing. Still nowhere near Nvidia pricing. Maxon, if you're listening, build that Maxon, if you're listening, build that Maxon, if you're listening, build that one. I want to see it. So, two cheap one. I want to see it. So, two cheap one. I want to see it. So, two cheap cards or one dual card? Which one would cards or one dual card? Which one would cards or one dual card? Which one would you buy? Let me know in the comments you buy? Let me know in the comments you buy? Let me know in the comments down below. I read all your comments. down below. I read all your comments. down below. I read all your comments. Watch the B60 video over here and the Watch the B60 video over here and the Watch the B60 video over here and the B70 video over here. Thanks for watching B70 video over here. Thanks for watching B70 video over here. Thanks for watching and I'll see you next time.
Summary
The main theme is the cost-effectiveness of running large language models locally, focusing on VRAM capacity. Key subjects discussed are Nvidia GPUs versus Intel's B-series GPUs for VRAM and the performance implications of multi-GPU setups. The practical takeaway is exploring cheaper alternatives like Intel GPUs and specialized dual-GPU cards from manufacturers like Maxon for high VRAM needs.