3 New PCs, One Giant AI Model… This Shouldn’t Work
Read full transcript 2 segments
-
Right now, this little machine is Right now, this little machine is running a 70 billion parameter AI model, running a 70 billion parameter AI model, running a 70 billion parameter AI model, and so is this one, and this one. All and so is this one, and this one. All and so is this one, and this one. All three together. But, here's the part three together. But, here's the part three together. But, here's the part that doesn't add up. Not one of them is that doesn't add up. Not one of them is that doesn't add up. Not one of them is big enough to hold this model. Not even big enough to hold this model. Not even big enough to hold this model. Not even close. So, in case you're not familiar, close. So, in case you're not familiar, close. So, in case you're not familiar, ASUS just released these These are the ASUS just released these These are the ASUS just released these These are the NUC 16 Pro, and these are Core Ultra X7 NUC 16 Pro, and these are Core Ultra X7 NUC 16 Pro, and these are Core Ultra X7 358H. They also have the new Arc B390 358H. They also have the new Arc B390 358H. They also have the new Arc B390 GPUs inside, up to 50 tops for the MPU. GPUs inside, up to 50 tops for the MPU. GPUs inside, up to 50 tops for the MPU. Now, my version also comes with 64 gigs Now, my version also comes with 64 gigs Now, my version also comes with 64 gigs of memory. And I got three of them to of memory. And I got three of them to of memory. And I got three of them to try something crazy that you probably try something crazy that you probably try something crazy that you probably shouldn't be doing with them. shouldn't be doing with them. shouldn't be doing with them. >> [music] >> More machines, more speed, right? No. >> More machines, more speed, right? No. I was kind of hopeful. The good old I was kind of hopeful. The good old I was kind of hopeful. The good old Llama 3.3 70B. This is a dense model. Llama 3.3 70B. This is a dense model. Llama 3.3 70B. This is a dense model. It's large, and it works. 1.4 tokens a It's large, and it works. 1.4 tokens a It's large, and it works. 1.4 tokens a second. Now, I'm a little disappointed, second. Now, I'm a little disappointed, second. Now, I'm a little disappointed, but there is a completely different way but there is a completely different way but there is a completely different way to use these three as a cluster. We to use these three as a cluster. We to use these three as a cluster. We don't split the model at all. We put a don't split the model at all. We put a don't split the model at all. We put a full [music] copy on every machine, and full [music] copy on every machine, and full [music] copy on every machine, and just send each request to a different just send each request to a different just send each request to a different one. Bye, 70 billion parameters. Aha!
-
one. Bye, 70 billion parameters. Aha! one. Bye, 70 billion parameters. Aha! This one. This one scales. One machine This one. This one scales. One machine This one. This one scales. One machine handled about 196 tokens a second under handled about 196 tokens a second under handled about 196 tokens a second under load. Three machines, almost 500 tokens load. Three machines, almost 500 tokens load. Three machines, almost 500 tokens a second. Two and a half times faster. a second. Two and a half times faster. a second. Two and a half times faster. Finally.
Summary
The main theme is running a large 70 billion parameter AI model on consumer hardware, specifically the ASUS NUC 16 Pro with Core Ultra X7 processors and Arc GPUs. While initially attempting to split the model proved disappointing, the practical takeaway is that replicating the full model on each device and directing requests to different machines significantly speeds up performance, achieving approximately two and a half times the original speed.