← Back
Aaron Zisk June 16, 2026 1m

Three months wrong about why my 4-node AMD cluster was slow

Read full transcript 2 segments
  1. More memory available to the graphics More memory available to the graphics card than a 5090 or even the RTX Pro card than a 5090 or even the RTX Pro card than a 5090 or even the RTX Pro 6000. 128 GB of unified memory on one 6000. 128 GB of unified memory on one 6000. 128 GB of unified memory on one chip. Still to this day the most chip. Still to this day the most chip. Still to this day the most powerful X86 APU you can buy. And four powerful X86 APU you can buy. And four powerful X86 APU you can buy. And four of them on one desk should basically be of them on one desk should basically be of them on one desk should basically be an AI supercomputer in your living room. an AI supercomputer in your living room. an AI supercomputer in your living room. Well, Well, Well, >> [laughter] >> [laughter] >> [laughter] >> it took me over 3 months to get these >> it took me over 3 months to get these >> it took me over 3 months to get these things actually to work together in a things actually to work together in a things actually to work together in a cluster as fast as they finally did. And cluster as fast as they finally did. And cluster as fast as they finally did. And most of that was just me being wrong most of that was just me being wrong most of that was just me being wrong about what the problem even was in the about what the problem even was in the about what the problem even was in the first place. Let me back up here. The first place. Let me back up here. The first place. Let me back up here. The chip is the AMD Ryzen AI Max Plus 395, chip is the AMD Ryzen AI Max Plus 395, chip is the AMD Ryzen AI Max Plus 395, aka Strix Halo. If you watch this aka Strix Halo. If you watch this aka Strix Halo. If you watch this channel, you've seen me try this chip on channel, you've seen me try this chip on channel, you've seen me try this chip on mini PCs and laptops from GMK Tech, mini PCs and laptops from GMK Tech, mini PCs and laptops from GMK Tech, Beelink, Framework, Asus. And this one Beelink, Framework, Asus. And this one Beelink, Framework, Asus. And this one is the Minisforum MS-S1 Max. It might be is the Minisforum MS-S1 Max. It might be is the Minisforum MS-S1 Max. It might be my favorite. And before this video is my favorite. And before this video is my favorite. And before this video is over, I'm going to try to beat over, I'm going to try to beat over, I'm going to try to beat Minisforum's own marketing video on a Minisforum's own marketing video on a Minisforum's own marketing video on a cluster of these things. Then I tried to cluster of these things. Then I tried to cluster of these things. Then I tried to actually cluster it. Finally, I got all actually cluster it. Finally, I got all actually cluster it. Finally, I got all four mesh together with a 4 ms ping.

  2. four mesh together with a 4 ms ping. four mesh together with a 4 ms ping. DeepSeek-R1. And you can hear it. You DeepSeek-R1. And you can hear it. You DeepSeek-R1. And you can hear it. You can hear it. After a clean reboot, all can hear it. After a clean reboot, all can hear it. After a clean reboot, all four nodes are running. Memory four nodes are running. Memory four nodes are running. Memory utilization is at 99% on all four utilization is at 99% on all four utilization is at 99% on all four machines. And it's running. It's machines. And it's running. It's machines. And it's running. It's actually printing everything out for me. actually printing everything out for me. actually printing everything out for me. Now, this is actually 37 billion active Now, this is actually 37 billion active Now, this is actually 37 billion active parameters in DeepSeek-R1. So, this parameters in DeepSeek-R1. So, this parameters in DeepSeek-R1. So, this takes some serious compute time. Now, takes some serious compute time. Now, takes some serious compute time. Now, you tell me, is this looking like it's you tell me, is this looking like it's you tell me, is this looking like it's faster than Minisforum's demo? We're faster than Minisforum's demo? We're faster than Minisforum's demo? We're using about 180 to 187 W on every using about 180 to 187 W on every using about 180 to 187 W on every machine. machine. machine. 4,058 tokens generated in 10 minutes. 4,058 tokens generated in 10 minutes. 4,058 tokens generated in 10 minutes. And guess what? Minisforum's own product And guess what? Minisforum's own product And guess what? Minisforum's own product video, same chip, same four nodes, video, same chip, same four nodes, video, same chip, same four nodes, DeepSeek-R1 quantized to 4 bits. Their DeepSeek-R1 quantized to 4 bits. Their DeepSeek-R1 quantized to 4 bits. Their numbers, 5.94 tokens per second. Mine, numbers, 5.94 tokens per second. Mine, numbers, 5.94 tokens per second. Mine, with VLN tensor parallelism, 6.23 tokens with VLN tensor parallelism, 6.23 tokens with VLN tensor parallelism, 6.23 tokens per second. Yes.

Summary

The main topic is the powerful performance of a cluster of AMD Ryzen AI Max Plus 395 (Strix Halo) chips for AI tasks. Key subjects include the 128GB unified memory, X86 APU capabilities, and running the DeepSeek-R1 model. The practical takeaway is that a DIY cluster of these chips can outperform vendor demonstrations for AI inference, achieving impressive token generation speeds.

View original episode ↗