3 New PCs, One Giant AI Model… This Shouldn’t Work
Read full transcript 14 segments
-
Right now, this little machine is Right now, this little machine is running a 70 billion parameter AI model, running a 70 billion parameter AI model, running a 70 billion parameter AI model, and so is this one and this one, all and so is this one and this one, all and so is this one and this one, all three together. But, here's the part three together. But, here's the part three together. But, here's the part that doesn't add up. Not one of them is that doesn't add up. Not one of them is that doesn't add up. Not one of them is big enough to hold this model, not even big enough to hold this model, not even big enough to hold this model, not even close. By the way, this is Intel's close. By the way, this is Intel's close. By the way, this is Intel's newest silicon, a brand new chip, a newest silicon, a brand new chip, a newest silicon, a brand new chip, a brand new GPU, a dedicated AI chip brand new GPU, a dedicated AI chip brand new GPU, a dedicated AI chip sitting on top of that, the NPU. And I sitting on top of that, the NPU. And I sitting on top of that, the NPU. And I got three of them to try something got three of them to try something got three of them to try something crazy. So, how is it running a model crazy. So, how is it running a model crazy. So, how is it running a model that none of them can actually fit? that none of them can actually fit? that none of them can actually fit? >> [music] >> [music] >> [music] >> Okay, so the dream here is simple. Mini >> Okay, so the dream here is simple. Mini >> Okay, so the dream here is simple. Mini PCs are getting pretty ridiculous. And PCs are getting pretty ridiculous. And PCs are getting pretty ridiculous. And if you're a developer, this is pretty if you're a developer, this is pretty if you're a developer, this is pretty convenient stuff here. Not only will it convenient stuff here. Not only will it convenient stuff here. Not only will it run your IDEs and code smoothly now, run your IDEs and code smoothly now, run your IDEs and code smoothly now, these used to be kind of a joke just a these used to be kind of a joke just a these used to be kind of a joke just a couple years ago. Now, they're serious. couple years ago. Now, they're serious. couple years ago. Now, they're serious. These things can also run AI locally, These things can also run AI locally, These things can also run AI locally, your own models, your own data, no API your own models, your own data, no API your own models, your own data, no API bill showing up at the end of the month. bill showing up at the end of the month. bill showing up at the end of the month. I have it next to my Mac mini to show I have it next to my Mac mini to show I have it next to my Mac mini to show you the relative size of this new one. you the relative size of this new one. you the relative size of this new one. It's shorter, it's narrower, it's a It's shorter, it's narrower, it's a It's shorter, it's narrower, it's a little bit longer than a Mac mini, but little bit longer than a Mac mini, but little bit longer than a Mac mini, but it is pretty small. Question is, can you it is pretty small. Question is, can you it is pretty small. Question is, can you take a few of these and bolt them take a few of these and bolt them take a few of these and bolt them together into one machine that runs together into one machine that runs together into one machine that runs stuff that you couldn't run on just one?
-
stuff that you couldn't run on just one? stuff that you couldn't run on just one? The part that people forget about a The part that people forget about a The part that people forget about a clean local setup is what happens when clean local setup is what happens when clean local setup is what happens when you need storage somewhere else. For you need storage somewhere else. For you need storage somewhere else. For anyone dealing with servers, data sets, anyone dealing with servers, data sets, anyone dealing with servers, data sets, checkpoints, or large project files, checkpoints, or large project files, checkpoints, or large project files, storage gets complicated fast. You want storage gets complicated fast. You want storage gets complicated fast. You want redundancy and remote access, but you redundancy and remote access, but you redundancy and remote access, but you don't want to give up privacy just to don't want to give up privacy just to don't want to give up privacy just to get it. Most services make uploading get it. Most services make uploading get it. Most services make uploading easy, but they're not really built easy, but they're not really built easy, but they're not really built around the idea that your data should around the idea that your data should around the idea that your data should stay private. That is where Internxt stay private. That is where Internxt stay private. That is where Internxt comes in. It's a privacy first cloud comes in. It's a privacy first cloud comes in. It's a privacy first cloud storage platform built around keeping storage platform built around keeping storage platform built around keeping your data under control. Your files are your data under control. Your files are your data under control. Your files are protected end-to-end zero-knowledge protected end-to-end zero-knowledge protected end-to-end zero-knowledge encryption, so only you can access them. encryption, so only you can access them. encryption, so only you can access them. Not even Internxt can access them. It Not even Internxt can access them. It Not even Internxt can access them. It also uses post-quantum cryptography to also uses post-quantum cryptography to also uses post-quantum cryptography to help you protect against current and help you protect against current and help you protect against current and future threats. And it fits real future threats. And it fits real future threats. And it fits real workflows, too, with web access, desktop workflows, too, with web access, desktop workflows, too, with web access, desktop and mobile apps, plus CLI, WebDAV, and mobile apps, plus CLI, WebDAV, and mobile apps, plus CLI, WebDAV, Rclone, and NAS support. Yeah, I can Rclone, and NAS support. Yeah, I can Rclone, and NAS support. Yeah, I can sync my NAS right to it. I like this sync my NAS right to it. I like this sync my NAS right to it. I like this more as a companion to a local setup, more as a companion to a local setup, more as a companion to a local setup, not a replacement for it. On top of not a replacement for it. On top of not a replacement for it. On top of storage, you also get privacy tools, storage, you also get privacy tools, storage, you also get privacy tools, file versioning for supported file file versioning for supported file file versioning for supported file types, and the lifetime ultimate plan types, and the lifetime ultimate plan types, and the lifetime ultimate plan gives you 5 TB gives you 5 TB gives you 5 TB for a one-time payment, no subscription.
-
for a one-time payment, no subscription. for a one-time payment, no subscription. So, the picture is simple, private cloud So, the picture is simple, private cloud So, the picture is simple, private cloud storage that actually fits the way storage that actually fits the way storage that actually fits the way technical people already work. This is technical people already work. This is technical people already work. This is not just vague security language. not just vague security language. not just vague security language. Internxt is open source and Internxt is open source and Internxt is open source and independently audited by Securium. It's independently audited by Securium. It's independently audited by Securium. It's also GDPR compliant and ISO 27001 also GDPR compliant and ISO 27001 also GDPR compliant and ISO 27001 certified. Use my link in the certified. Use my link in the certified. Use my link in the description or go to description or go to description or go to internxt.com/alexxiskin internxt.com/alexxiskin internxt.com/alexxiskin and use code alexxiskin to get 87% off and use code alexxiskin to get 87% off and use code alexxiskin to get 87% off the lifetime 5 TB plan. So, in case the lifetime 5 TB plan. So, in case the lifetime 5 TB plan. So, in case you're not familiar, ASUS just released you're not familiar, ASUS just released you're not familiar, ASUS just released these, and these are the NUC 16 Pro. these, and these are the NUC 16 Pro. these, and these are the NUC 16 Pro. This particular design actually varies This particular design actually varies This particular design actually varies quite a bit. You get Wi-Fi 7 and quite a bit. You get Wi-Fi 7 and quite a bit. You get Wi-Fi 7 and Bluetooth 6 in these, and they carry the Bluetooth 6 in these, and they carry the Bluetooth 6 in these, and they carry the Pro name because they have redundancy Pro name because they have redundancy Pro name because they have redundancy for pretty much everything. Two for pretty much everything. Two for pretty much everything. Two Thunderbolts, two HDMIs, four USB-As, Thunderbolts, two HDMIs, four USB-As, Thunderbolts, two HDMIs, four USB-As, dual Ethernet ports, upgradeable RAM, dual Ethernet ports, upgradeable RAM, dual Ethernet ports, upgradeable RAM, but they can come with Intel Core Ultra but they can come with Intel Core Ultra but they can come with Intel Core Ultra 5 325 or 559, Core Ultra 7 356, and 5 325 or 559, Core Ultra 7 356, and 5 325 or 559, Core Ultra 7 356, and these are Core Ultra X7 358H. So, this these are Core Ultra X7 358H. So, this these are Core Ultra X7 358H. So, this is like the high-end ones. Here's my is like the high-end ones. Here's my is like the high-end ones. Here's my favorite part, though. If you want to favorite part, though. If you want to favorite part, though. If you want to get inside, get inside, get inside, it's just a little pull tab like that.
-
it's just a little pull tab like that. it's just a little pull tab like that. More mini PCs should be doing this kind More mini PCs should be doing this kind More mini PCs should be doing this kind of thing. Look how easy it is to get in of thing. Look how easy it is to get in of thing. Look how easy it is to get in and get out. Done. and get out. Done. and get out. Done. I was inside of that just now. They can I was inside of that just now. They can I was inside of that just now. They can actually get up to the Intel Core Ultra actually get up to the Intel Core Ultra actually get up to the Intel Core Ultra X9, which I don't have, but I'd imagine X9, which I don't have, but I'd imagine X9, which I don't have, but I'd imagine those would be pretty crazy. They also those would be pretty crazy. They also those would be pretty crazy. They also have the new Arc B390 GPUs inside and a have the new Arc B390 GPUs inside and a have the new Arc B390 GPUs inside and a separate MPU that's the dedicated AI separate MPU that's the dedicated AI separate MPU that's the dedicated AI engine. Now, my version also comes with engine. Now, my version also comes with engine. Now, my version also comes with 64 gigs of memory. This one on Newegg 64 gigs of memory. This one on Newegg 64 gigs of memory. This one on Newegg has 32, and it's already at $1,700. So, has 32, and it's already at $1,700. So, has 32, and it's already at $1,700. So, you can imagine how much these cost. you can imagine how much these cost. you can imagine how much these cost. Check the [music] current pricing cuz Check the [music] current pricing cuz Check the [music] current pricing cuz that's always changing. Now, even though that's always changing. Now, even though that's always changing. Now, even though there's There's no question that these there's There's no question that these there's There's no question that these would make incredible little dev would make incredible little dev would make incredible little dev machines. I'm going to be doing machines. I'm going to be doing machines. I'm going to be doing something you probably shouldn't be something you probably shouldn't be something you probably shouldn't be doing with them, making an AI cluster doing with them, making an AI cluster doing with them, making an AI cluster out of them. And that dual Thunderbolt out of them. And that dual Thunderbolt out of them. And that dual Thunderbolt that I was talking about, keep an eye on that I was talking about, keep an eye on that I was talking about, keep an eye on that because it comes back later in a that because it comes back later in a that because it comes back later in a big way. So, if you've been watching big way. So, if you've been watching big way. So, if you've been watching this channel at all recently, you know this channel at all recently, you know this channel at all recently, you know that I have a few machines here that that I have a few machines here that that I have a few machines here that I've been clustering including these I've been clustering including these I've been clustering including these Dell GB10s, which are basically the same Dell GB10s, which are basically the same Dell GB10s, which are basically the same thing as the DGX Spark, and the Mac thing as the DGX Spark, and the Mac thing as the DGX Spark, and the Mac [music] Studios. And all those machines [music] Studios. And all those machines [music] Studios. And all those machines can be clustered using RDMA, which can be clustered using RDMA, which can be clustered using RDMA, which basically means you can go back and basically means you can go back and basically means you can go back and watch the videos. But basically that watch the videos. But basically that watch the videos. But basically that means that the more machines you add, means that the more machines you add, means that the more machines you add, the faster it gets, generally speaking.
-
the faster it gets, generally speaking. the faster it gets, generally speaking. This is basically for things like large This is basically for things like large This is basically for things like large language models, LLMs. When you're language models, LLMs. When you're language models, LLMs. When you're generating text, generating images, if generating text, generating images, if generating text, generating images, if you decrease the network latency between you decrease the network latency between you decrease the network latency between the machines, you're going to get faster the machines, you're going to get faster the machines, you're going to get faster generation. And when you're doing text generation. And when you're doing text generation. And when you're doing text generation or image generation or generation or image generation or generation or image generation or whatnot, there's always two stages. One whatnot, there's always two stages. One whatnot, there's always two stages. One is prompt processing. [music] If you're is prompt processing. [music] If you're is prompt processing. [music] If you're working with code, for example, and working with code, for example, and working with code, for example, and you've got your IDE talking to an LLM you've got your IDE talking to an LLM you've got your IDE talking to an LLM for code generation or debugging or for code generation or debugging or for code generation or debugging or whatnot, you're going to be sending all whatnot, you're going to be sending all whatnot, you're going to be sending all that code as context to the LLM. That's that code as context to the LLM. That's that code as context to the LLM. That's where prompt processing [music] where prompt processing [music] where prompt processing [music] comes in handy. So, prompt processing comes in handy. So, prompt processing comes in handy. So, prompt processing has to be fast. I won't get into the has to be fast. I won't get into the has to be fast. I won't get into the details of that setup in this video details of that setup in this video details of that setup in this video because I have a bunch of other videos because I have a bunch of other videos because I have a bunch of other videos showing that. Now I just want to test showing that. Now I just want to test showing that. Now I just want to test how fast these things do that. And the how fast these things do that. And the how fast these things do that. And the first thing I wanted to know on a brand first thing I wanted to know on a brand first thing I wanted to know on a brand new chip, can that new Arc GPU actually new chip, can that new Arc GPU actually new chip, can that new Arc GPU actually speed up an LLM? So, I ran the same speed up an LLM? So, I ran the same speed up an LLM? So, I ran the same model on the CPU and then on the GPU. model on the CPU and then on the GPU. model on the CPU and then on the GPU. Turning the GPU on essentially doubles Turning the GPU on essentially doubles Turning the GPU on essentially doubles the speed that it reads your prompt. the speed that it reads your prompt. the speed that it reads your prompt. Just over 1,000 tokens per second up to Just over 1,000 tokens per second up to Just over 1,000 tokens per second up to 2,200 tokens per second. [music] 2,200 tokens per second. [music] 2,200 tokens per second. [music] Boom. Nice. The second part of LLM's Boom. Nice. The second part of LLM's Boom. Nice. The second part of LLM's inference is token generation. After inference is token generation. After inference is token generation. After it's done its prompt processing, it it's done its prompt processing, it it's done its prompt processing, it switches over to the generating part.
-
switches over to the generating part. switches over to the generating part. And that, right here my friends, is And that, right here my friends, is And that, right here my friends, is about 46 tokens per second on the CPU about 46 tokens per second on the CPU about 46 tokens per second on the CPU and about 46 tokens per second on the and about 46 tokens per second on the and about 46 tokens per second on the GPU. It didn't move. And this is the GPU. It didn't move. And this is the GPU. It didn't move. And this is the thing I want you to remember because it thing I want you to remember because it thing I want you to remember because it comes back to haunt us the rest of this comes back to haunt us the rest of this comes back to haunt us the rest of this video. This is the memory wall. Token video. This is the memory wall. Token video. This is the memory wall. Token generation relies on memory bandwidth. generation relies on memory bandwidth. generation relies on memory bandwidth. Prompt processing relies on that brand Prompt processing relies on that brand Prompt processing relies on that brand new hot GPU chip that's really fast. new hot GPU chip that's really fast. new hot GPU chip that's really fast. Both sides are very important. And on Both sides are very important. And on Both sides are very important. And on these particular chips, the Meteor Lake these particular chips, the Meteor Lake these particular chips, the Meteor Lake chips, the GPU shares the same memory as chips, the GPU shares the same memory as chips, the GPU shares the same memory as the CPU and generating text is limited the CPU and generating text is limited the CPU and generating text is limited by that memory speed. So, the GPU's by that memory speed. So, the GPU's by that memory speed. So, the GPU's extra muscle does nothing for it. For extra muscle does nothing for it. For extra muscle does nothing for it. For the generation tokens per second, memory the generation tokens per second, memory the generation tokens per second, memory bandwidth is the boss. So, the GPU bandwidth is the boss. So, the GPU bandwidth is the boss. So, the GPU helps, but it's not magic and there's helps, but it's not magic and there's helps, but it's not magic and there's still that third engine that I hadn't still that third engine that I hadn't still that third engine that I hadn't touched, the NPU. If the GPU's capped by touched, the NPU. If the GPU's capped by touched, the NPU. If the GPU's capped by memory, maybe a dedicated AI chip is the memory, maybe a dedicated AI chip is the memory, maybe a dedicated AI chip is the real story, right? So, let's get it into real story, right? So, let's get it into real story, right? So, let's get it into the game. First problem, the tool that the game. First problem, the tool that the game. First problem, the tool that we can use on these machines, the most we can use on these machines, the most we can use on these machines, the most popular tool for doing inference is popular tool for doing inference is popular tool for doing inference is Llama.cpp. It can't even talk to the Llama.cpp. It can't even talk to the Llama.cpp. It can't even talk to the NPU. So, I switch over to Intel's own NPU. So, I switch over to Intel's own NPU. So, I switch over to Intel's own OpenVINO. This is Intel's own software.
-
OpenVINO. This is Intel's own software. OpenVINO. This is Intel's own software. And small model on the NPU actually And small model on the NPU actually And small model on the NPU actually works. Cool. By the way, if you're not works. Cool. By the way, if you're not works. Cool. By the way, if you're not familiar with OpenVINO, it's supposed to familiar with OpenVINO, it's supposed to familiar with OpenVINO, it's supposed to be able to use the GPU, the CPU, and the be able to use the GPU, the CPU, and the be able to use the GPU, the CPU, and the NPU and basically automatically decide NPU and basically automatically decide NPU and basically automatically decide for you which is best. And OpenVINO has for you which is best. And OpenVINO has for you which is best. And OpenVINO has their own toolkit and models on Hugging their own toolkit and models on Hugging their own toolkit and models on Hugging Face. And these are specifically Face. And these are specifically Face. And these are specifically OpenVINO models with the OV extension so OpenVINO models with the OV extension so OpenVINO models with the OV extension so you can know what they are cuz Hugging you can know what they are cuz Hugging you can know what they are cuz Hugging Face model naming conventions are so Face model naming conventions are so Face model naming conventions are so conventiony. conventiony. conventiony. Right. And then I try a bigger model and Right. And then I try a bigger model and Right. And then I try a bigger model and it fails. Intel's own pre-built models it fails. Intel's own pre-built models it fails. Intel's own pre-built models will not run on Intel's own NPU. So, I will not run on Intel's own NPU. So, I will not run on Intel's own NPU. So, I had to go and rebuild the model myself had to go and rebuild the model myself had to go and rebuild the model myself just to get it to load. That's the just to get it to load. That's the just to get it to load. That's the bleeding edge for you. But we know these bleeding edge for you. But we know these bleeding edge for you. But we know these days what with these chips. They come days what with these chips. They come days what with these chips. They come out, they're great, but if you want to out, they're great, but if you want to out, they're great, but if you want to do any kind of bleeding edge AI stuff, do any kind of bleeding edge AI stuff, do any kind of bleeding edge AI stuff, the software has to catch up. However, the software has to catch up. However, the software has to catch up. However, once I got it running, here's all three once I got it running, here's all three once I got it running, here's all three engines across a few models. Smallish engines across a few models. Smallish engines across a few models. Smallish models, GPU is fastest everywhere, but models, GPU is fastest everywhere, but models, GPU is fastest everywhere, but the NPU beats the CPU. It's just not a the NPU beats the CPU. It's just not a the NPU beats the CPU. It's just not a speed demon, and it's a bit slow to spin speed demon, and it's a bit slow to spin speed demon, and it's a bit slow to spin up. Now, hold on. I just used two up. Now, hold on. I just used two up. Now, hold on. I just used two completely different pieces of software.
-
completely different pieces of software. completely different pieces of software. Llama.cpp for the GPU and OpenVINO for Llama.cpp for the GPU and OpenVINO for Llama.cpp for the GPU and OpenVINO for the NPU. But, I did mention that the NPU. But, I did mention that the NPU. But, I did mention that OpenVINO also works on the GPU. So, on OpenVINO also works on the GPU. So, on OpenVINO also works on the GPU. So, on the exact same GPU, which one is the exact same GPU, which one is the exact same GPU, which one is actually faster, OpenVINO or Llama.cpp? actually faster, OpenVINO or Llama.cpp? actually faster, OpenVINO or Llama.cpp? Well, I ran that, too. Same model, same Well, I ran that, too. Same model, same Well, I ran that, too. Same model, same chip, and I did not expect this. The chip, and I did not expect this. The chip, and I did not expect this. The free open-source tool beat Intel's own free open-source tool beat Intel's own free open-source tool beat Intel's own software by about 2 and 1/2 times. 34 software by about 2 and 1/2 times. 34 software by about 2 and 1/2 times. 34 tokens per second with Vulcan. Vulcan is tokens per second with Vulcan. Vulcan is tokens per second with Vulcan. Vulcan is just the API that Llama.cpp was using in just the API that Llama.cpp was using in just the API that Llama.cpp was using in this case versus Intel's 14 tokens per this case versus Intel's 14 tokens per this case versus Intel's 14 tokens per second. On Intel's own GPU, come on. So, second. On Intel's own GPU, come on. So, second. On Intel's own GPU, come on. So, file that away because when we get to file that away because when we get to file that away because when we get to the cluster, we're going to be sticking the cluster, we're going to be sticking the cluster, we're going to be sticking with Llama.cpp here. It's faster, and with Llama.cpp here. It's faster, and with Llama.cpp here. It's faster, and it's really the only way to get it's really the only way to get it's really the only way to get clustering going anyway on this clustering going anyway on this clustering going anyway on this particular setup. But, there's one more particular setup. But, there's one more particular setup. But, there's one more thing that I wanted to settle before thing that I wanted to settle before thing that I wanted to settle before leaving one single box. Speed was never leaving one single box. Speed was never leaving one single box. Speed was never NPU's whole pitch anyway. Efficiency is. NPU's whole pitch anyway. Efficiency is. NPU's whole pitch anyway. Efficiency is. Everyone says their NPU is the most Everyone says their NPU is the most Everyone says their NPU is the most efficient one, and everybody's got an efficient one, and everybody's got an efficient one, and everybody's got an NPU now. Okay, let's actually measure NPU now. Okay, let's actually measure NPU now. Okay, let's actually measure it. Real watts while it's generating.
-
it. Real watts while it's generating. it. Real watts while it's generating. And the answer is both things are true. And the answer is both things are true. And the answer is both things are true. The NPU actually draws the least amount The NPU actually draws the least amount The NPU actually draws the least amount of power, 17 watts versus 24 for the GPU of power, 17 watts versus 24 for the GPU of power, 17 watts versus 24 for the GPU versus almost 30 for the CPU. And the versus almost 30 for the CPU. And the versus almost 30 for the CPU. And the power draw? The NPU has the lowest draw. power draw? The NPU has the lowest draw. power draw? The NPU has the lowest draw. It's the coolest and the quietest. That It's the coolest and the quietest. That It's the coolest and the quietest. That part is real. However, the GPU is part is real. However, the GPU is part is real. However, the GPU is actually more efficient per token. actually more efficient per token. actually more efficient per token. That's that red line right there. That's that red line right there. That's that red line right there. Because it's twice as fast. It finishes Because it's twice as fast. It finishes Because it's twice as fast. It finishes the work sooner and uses less total the work sooner and uses less total the work sooner and uses less total energy per token. And usually you're energy per token. And usually you're energy per token. And usually you're going to be doing these kinds of bursty going to be doing these kinds of bursty going to be doing these kinds of bursty operations on a machine like this, operations on a machine like this, operations on a machine like this, unless you're running agents 24/7. So, unless you're running agents 24/7. So, unless you're running agents 24/7. So, the NPU sips the least power, but the the NPU sips the least power, but the the NPU sips the least power, but the GPU gets the most AI done per joule. The GPU gets the most AI done per joule. The GPU gets the most AI done per joule. The CPU CPU CPU it loses on both, but that's kind of it loses on both, but that's kind of it loses on both, but that's kind of always been the case. However, the CPU always been the case. However, the CPU always been the case. However, the CPU is making a comeback and I'll talk more is making a comeback and I'll talk more is making a comeback and I'll talk more about that in future videos. So, the about that in future videos. So, the about that in future videos. So, the whole lesson from one box, memory whole lesson from one box, memory whole lesson from one box, memory bandwidth runs the show. bandwidth runs the show. bandwidth runs the show. >> [music] >> [music] >> [music] >> So, I squeezed everything I could out of >> So, I squeezed everything I could out of >> So, I squeezed everything I could out of one machine. The only way to go bigger one machine. The only way to go bigger one machine. The only way to go bigger is more machines. I brought in the other is more machines. I brought in the other is more machines. I brought in the other two Knox, wired them all together, and two Knox, wired them all together, and two Knox, wired them all together, and split one model across all three. More split one model across all three. More split one model across all three. More machines, more speed, right? No.
-
machines, more speed, right? No. machines, more speed, right? No. >> [laughter] >> [laughter] >> [laughter] >> I was kind of hopeful, but >> I was kind of hopeful, but >> I was kind of hopeful, but now I'm a little disappointed. But, now I'm a little disappointed. But, now I'm a little disappointed. But, we've seen this story before with my we've seen this story before with my we've seen this story before with my Framework video cluster situation before Framework video cluster situation before Framework video cluster situation before I got RDMA working on that, and my I got RDMA working on that, and my I got RDMA working on that, and my MinisForum cluster video, it got slower. MinisForum cluster video, it got slower. MinisForum cluster video, it got slower. The same model on one machine was about The same model on one machine was about The same model on one machine was about 35 tokens a second, and across all 35 tokens a second, and across all 35 tokens a second, and across all three, 17. It got essentially cut in three, 17. It got essentially cut in three, 17. It got essentially cut in half. This is Qwen 3.6 35 billion, by half. This is Qwen 3.6 35 billion, by half. This is Qwen 3.6 35 billion, by the way, which easily fits into one the way, which easily fits into one the way, which easily fits into one machine. Once you see why this happens machine. Once you see why this happens machine. Once you see why this happens though, it all makes sense, even if it though, it all makes sense, even if it though, it all makes sense, even if it stinks. This kind of clustering splits stinks. This kind of clustering splits stinks. This kind of clustering splits the model into chunks, and every single the model into chunks, and every single the model into chunks, and every single token has to hop from one machine to token has to hop from one machine to token has to hop from one machine to another machine, to the third over the another machine, to the third over the another machine, to the third over the network. network. network. Over these Ethernet wires going to the Over these Ethernet wires going to the Over these Ethernet wires going to the switch, and then back. You're not adding switch, and then back. You're not adding switch, and then back. You're not adding power, you're adding traffic. So, we power, you're adding traffic. So, we power, you're adding traffic. So, we have the same memory wall as before, but have the same memory wall as before, but have the same memory wall as before, but now we've just added the network tax on now we've just added the network tax on now we've just added the network tax on top of it. So, if clustering doesn't top of it. So, if clustering doesn't top of it. So, if clustering doesn't make a model faster, what's it actually make a model faster, what's it actually make a model faster, what's it actually for? With the caveat that I've mentioned for? With the caveat that I've mentioned for? With the caveat that I've mentioned before, clustering can make before, clustering can make before, clustering can make inference faster if you're using the inference faster if you're using the inference faster if you're using the right technology, which is RDMA, which right technology, which is RDMA, which right technology, which is RDMA, which I'm showing in other videos, and I'll I'm showing in other videos, and I'll I'm showing in other videos, and I'll link those down below. But, the answer link those down below. But, the answer link those down below. But, the answer is the whole reason I did it. It's not is the whole reason I did it. It's not is the whole reason I did it. It's not for speed, it's for size. Let's find the for speed, it's for size. Let's find the for speed, it's for size. Let's find the model that no single one of these model that no single one of these model that no single one of these machines could ever run on its own.
-
machines could ever run on its own. machines could ever run on its own. Ah, yes, the good old Llama 3.370B. Ah, yes, the good old Llama 3.370B. Ah, yes, the good old Llama 3.370B. This is a dense model, it's large, it's This is a dense model, it's large, it's This is a dense model, it's large, it's been around for a little while, but all been around for a little while, but all been around for a little while, but all the models have been either mixture of the models have been either mixture of the models have been either mixture of experts models, which means they have a experts models, which means they have a experts models, which means they have a certain number of parameters, like 200 certain number of parameters, like 200 certain number of parameters, like 200 something, but they only have a small something, but they only have a small something, but they only have a small number of active parameters. So, they've number of active parameters. So, they've number of active parameters. So, they've been bigger in parameter count than the been bigger in parameter count than the been bigger in parameter count than the 70B, or much lower in parameter count 70B, or much lower in parameter count 70B, or much lower in parameter count than the 70B, but they're not really than the 70B, but they're not really than the 70B, but they're not really dense, they're mixture of experts. This dense, they're mixture of experts. This dense, they're mixture of experts. This one is dense, and the one I ran is 75 one is dense, and the one I ran is 75 one is dense, and the one I ran is 75 GB. GB. GB. A single machine has 64 gigs of memory. A single machine has 64 gigs of memory. A single machine has 64 gigs of memory. It physically cannot load this, period. It physically cannot load this, period. It physically cannot load this, period. But, three of them together, that's 192 But, three of them together, that's 192 But, three of them together, that's 192 GB in the pool. So, I split the model GB in the pool. So, I split the model GB in the pool. So, I split the model across all three, and it works. It across all three, and it works. It across all three, and it works. It actually works, folks. Three little actually works, folks. Three little actually works, folks. Three little machines, not one of them big enough on machines, not one of them big enough on machines, not one of them big enough on its own to run the 70 billion parameter its own to run the 70 billion parameter its own to run the 70 billion parameter model. It's slow, okay? 1.4 tokens a model. It's slow, okay? 1.4 tokens a model. It's slow, okay? 1.4 tokens a second, but that's not the point. The second, but that's not the point. The second, but that's not the point. The point is, it ran. And the fun part is point is, it ran. And the fun part is point is, it ran. And the fun part is that, that, that, I asked it, "What is remarkable about I asked it, "What is remarkable about I asked it, "What is remarkable about three small computers running one big AI three small computers running one big AI three small computers running one big AI model?" And it described its own model?" And it described its own model?" And it described its own situation right back to me. Okay, so it situation right back to me. Okay, so it situation right back to me. Okay, so it works, but it's slow, and every one of works, but it's slow, and every one of works, but it's slow, and every one of those tokens is hopping between the those tokens is hopping between the those tokens is hopping between the machines over plain 2.5 gig ethernet.
-
machines over plain 2.5 gig ethernet. machines over plain 2.5 gig ethernet. The fix seems obvious, right? Make the The fix seems obvious, right? Make the The fix seems obvious, right? Make the network way faster. Each one of these network way faster. Each one of these network way faster. Each one of these machines has two Thunderbolt ports on machines has two Thunderbolt ports on machines has two Thunderbolt ports on it. Remember that second port I told you it. Remember that second port I told you it. Remember that second port I told you about? 20 gigabit, like eight times the about? 20 gigabit, like eight times the about? 20 gigabit, like eight times the bandwidth. So, I wire all three together bandwidth. So, I wire all three together bandwidth. So, I wire all three together in a triangle. Every machine straight to in a triangle. Every machine straight to in a triangle. Every machine straight to every other one. This has to fix it, and every other one. This has to fix it, and every other one. This has to fix it, and it did it did it did basically nothing. Loading the model got basically nothing. Loading the model got basically nothing. Loading the model got a little faster on the small one, but a little faster on the small one, but a little faster on the small one, but the big 70 billion parameter barely the big 70 billion parameter barely the big 70 billion parameter barely budged. And generating tokens? budged. And generating tokens? budged. And generating tokens? On a 70 billion parameter, it was On a 70 billion parameter, it was On a 70 billion parameter, it was identical. 1.43 before, 1.43 after. The identical. 1.43 before, 1.43 after. The identical. 1.43 before, 1.43 after. The smaller model actually got worse, and it smaller model actually got worse, and it smaller model actually got worse, and it kept crashing. Now, here's why. It's the kept crashing. Now, here's why. It's the kept crashing. Now, here's why. It's the same wall from that very beginning. The same wall from that very beginning. The same wall from that very beginning. The bottleneck was never the cable or the bottleneck was never the cable or the bottleneck was never the cable or the bandwidth speed. It's memory speed and bandwidth speed. It's memory speed and bandwidth speed. It's memory speed and thousands of tiny little messages flying thousands of tiny little messages flying thousands of tiny little messages flying back and forth every second. A wider back and forth every second. A wider back and forth every second. A wider highway doesn't fix the traffic jam that highway doesn't fix the traffic jam that highway doesn't fix the traffic jam that isn't on the highway. Well played, isn't on the highway. Well played, isn't on the highway. Well played, physics. Splitting one model across physics. Splitting one model across physics. Splitting one model across machines is a dead end for speed. But, machines is a dead end for speed. But, machines is a dead end for speed. But, there is a completely different way to there is a completely different way to there is a completely different way to use these three as a cluster. And I've use these three as a cluster. And I've use these three as a cluster. And I've shown kind of a demo of this in a shown kind of a demo of this in a shown kind of a demo of this in a previous video. We don't split the model previous video. We don't split the model previous video. We don't split the model at all. We put a full copy on every at all. We put a full copy on every at all. We put a full copy on every machine and just send each request to a machine and just send each request to a machine and just send each request to a different one. This, of course, will different one. This, of course, will different one. This, of course, will only work for models that will fit only work for models that will fit only work for models that will fit inside the RAM of one machine. By 70 inside the RAM of one machine. By 70 inside the RAM of one machine. By 70 billion parameters, aha, this one. This billion parameters, aha, this one. This billion parameters, aha, this one. This one scales. One machine handled about one scales. One machine handled about one scales. One machine handled about 196 tokens a second under load. Three 196 tokens a second under load. Three 196 tokens a second under load. Three machines, almost 500 tokens a second.
-
machines, almost 500 tokens a second. machines, almost 500 tokens a second. Two and a half times faster. Finally. Two and a half times faster. Finally. Two and a half times faster. Finally. So, here's the rule, and it's the whole So, here's the rule, and it's the whole So, here's the rule, and it's the whole thing in one line. You cluster one way thing in one line. You cluster one way thing in one line. You cluster one way for size, you split the model, and a for size, you split the model, and a for size, you split the model, and a completely different way for speed and completely different way for speed and completely different way for speed and for throughput, copy the model and for throughput, copy the model and for throughput, copy the model and spread the work. Don't mix them up. So, spread the work. Don't mix them up. So, spread the work. Don't mix them up. So, step back and look what actually step back and look what actually step back and look what actually happened here. Same lesson at two happened here. Same lesson at two happened here. Same lesson at two different scales. There's three AI different scales. There's three AI different scales. There's three AI engines in one box or three boxes in one engines in one box or three boxes in one engines in one box or three boxes in one cluster. Now, I think that this could be cluster. Now, I think that this could be cluster. Now, I think that this could be solved at a software level to some solved at a software level to some solved at a software level to some degree. Uh, I'm not 100% sure about degree. Uh, I'm not 100% sure about degree. Uh, I'm not 100% sure about that, but Apple did [music] it. Just that, but Apple did [music] it. Just that, but Apple did [music] it. Just magically updated the software and magically updated the software and magically updated the software and suddenly RDMA works. Do not look at suddenly RDMA works. Do not look at suddenly RDMA works. Do not look at Petella in his videos. By the way, his Petella in his videos. By the way, his Petella in his videos. By the way, his channel is somewhere down below. channel is somewhere down below. channel is somewhere down below. Also, he got it working with his Strix Also, he got it working with his Strix Also, he got it working with his Strix Halo boxes. I did the same thing Halo boxes. I did the same thing Halo boxes. I did the same thing following his example. So, maybe there following his example. So, maybe there following his example. So, maybe there is a way to unlock that in software is a way to unlock that in software is a way to unlock that in software here. Is the newest Intel stuff ready here. Is the newest Intel stuff ready here. Is the newest Intel stuff ready with the software stack? The hardware is with the software stack? The hardware is with the software stack? The hardware is real and it's impressive, but the real and it's impressive, but the real and it's impressive, but the bleeding edge means rough edges. Intel's bleeding edge means rough edges. Intel's bleeding edge means rough edges. Intel's own software wasn't ready for Intel's own software wasn't ready for Intel's own software wasn't ready for Intel's own chip. I'm rooting for them. They own chip. I'm rooting for them. They own chip. I'm rooting for them. They made some really nice leaps recently, made some really nice leaps recently, made some really nice leaps recently, especially with their discrete especially with their discrete especially with their discrete professional GPUs and I've made a video professional GPUs and I've made a video professional GPUs and I've made a video about the B60, B70. You can check all about the B60, B70. You can check all about the B60, B70. You can check all those out. So, if I were you, for model those out. So, if I were you, for model those out. So, if I were you, for model that fits on one machine, the cluster is that fits on one machine, the cluster is that fits on one machine, the cluster is kind of pointless. One box wins every kind of pointless. One box wins every kind of pointless. One box wins every time. A cluster like this would be time. A cluster like this would be time. A cluster like this would be better served as a Proxmox server and better served as a Proxmox server and better served as a Proxmox server and virtual machines, for example. This
-
virtual machines, for example. This virtual machines, for example. This whole setup only earns its keep and its whole setup only earns its keep and its whole setup only earns its keep and its price tag when the model is too big for price tag when the model is too big for price tag when the model is too big for any single machine or you've got an any single machine or you've got an any single machine or you've got an office of people and you want to serve office of people and you want to serve office of people and you want to serve them. I'm curious, if you were to run a them. I'm curious, if you were to run a them. I'm curious, if you were to run a cluster like this with many PCs that are cluster like this with many PCs that are cluster like this with many PCs that are APUs, for example, like this one, would APUs, for example, like this one, would APUs, for example, like this one, would you run it for AI or would you run it you run it for AI or would you run it you run it for AI or would you run it for for >> [music] >> Proxmox or virtual machines? What would >> Proxmox or virtual machines? What would >> Proxmox or virtual machines? What would you do? Let me know in the comments down you do? Let me know in the comments down you do? Let me know in the comments down below. Very curious to read that and I below. Very curious to read that and I below. Very curious to read that and I read all your comments. Do check out read all your comments. Do check out read all your comments. Do check out that cluster setup right over here and that cluster setup right over here and that cluster setup right over here and the framework cluster setup over here. the framework cluster setup over here. the framework cluster setup over here. Thanks for watching and I'll see you Thanks for watching and I'll see you Thanks for watching and I'll see you next [music] time.
Summary
The transcript explores the surprising capability of mini PCs, equipped with Intel's newest silicon, to run large AI models that exceed their individual storage limits through clever integration. This advancement makes local AI development and data privacy more accessible, with Internxt offering a solution for secure, end-to-end encrypted cloud storage that complements local setups rather than replacing them. The key takeaway is that powerful local AI processing and robust data privacy are becoming increasingly feasible and essential for developers.