← Back
Aaron Zisk June 30, 2026 14m

Your OS Changes Everything for Local AI

Read full transcript 11 segments
  1. This is the GEEKOM A9 Max. Does that This is the GEEKOM A9 Max. Does that sound familiar? sound familiar? sound familiar? Uh yeah, it's the same name from last Uh yeah, it's the same name from last Uh yeah, it's the same name from last year's model. But it's got the AMD Ryzen year's model. But it's got the AMD Ryzen year's model. But it's got the AMD Ryzen AI 9 HX 470 this time. It's an updated AI 9 HX 470 this time. It's an updated AI 9 HX 470 this time. It's an updated Strix Point APU. 12 Zen 5 cores, Radeon Strix Point APU. 12 Zen 5 cores, Radeon Strix Point APU. 12 Zen 5 cores, Radeon 890M IGPU, and 50 TOPS NPU. For the AI 890M IGPU, and 50 TOPS NPU. For the AI 890M IGPU, and 50 TOPS NPU. For the AI crowd specifically, which operating crowd specifically, which operating crowd specifically, which operating system should you run on it? You can do system should you run on it? You can do system should you run on it? You can do Windows out of the box. You can do WSL, Windows out of the box. You can do WSL, Windows out of the box. You can do WSL, which is Windows Subsystem for Linux. It which is Windows Subsystem for Linux. It which is Windows Subsystem for Linux. It runs inside Windows but in a virtual runs inside Windows but in a virtual runs inside Windows but in a virtual machine if you want the whole Linux machine if you want the whole Linux machine if you want the whole Linux toolchain. Or you can install Linux on toolchain. Or you can install Linux on toolchain. Or you can install Linux on bare metal because some people say Linux bare metal because some people say Linux bare metal because some people say Linux is faster for AI. But let's find out is faster for AI. But let's find out is faster for AI. But let's find out because I ended up running a bunch of because I ended up running a bunch of because I ended up running a bunch of benchmarks here on all three OSs. Same benchmarks here on all three OSs. Same benchmarks here on all three OSs. Same hardware, same model, same prompts, and hardware, same model, same prompts, and hardware, same model, same prompts, and the winners might actually surprise you, the winners might actually surprise you, the winners might actually surprise you, especially for long context prefill. especially for long context prefill. especially for long context prefill. Also, as we go through this video, some Also, as we go through this video, some Also, as we go through this video, some of the numbers came in about half of of the numbers came in about half of of the numbers came in about half of what the machine chip should be doing. what the machine chip should be doing. what the machine chip should be doing. So, we'll come back to that. So, these So, we'll come back to that. So, these So, we'll come back to that. So, these days I'm always flipping between models. days I'm always flipping between models. days I'm always flipping between models. GPT for research, Claude for coding, GPT for research, Claude for coding, GPT for research, Claude for coding, Nano Banana for image generation, VEO Nano Banana for image generation, VEO Nano Banana for image generation, VEO Kling and Runway for video. Six tabs, Kling and Runway for video. Six tabs, Kling and Runway for video. Six tabs, six bills, and counting. Enter Chat LLM six bills, and counting. Enter Chat LLM six bills, and counting. Enter Chat LLM Teams. One dashboard houses every top Teams. One dashboard houses every top Teams. One dashboard houses every top LLM and route LLM picks the right one.

  2. LLM and route LLM picks the right one. LLM and route LLM picks the right one. GPT Mini for ultra fast answers, Claude GPT Mini for ultra fast answers, Claude GPT Mini for ultra fast answers, Claude Sonnet for coding, Gemini Pro for Sonnet for coding, Gemini Pro for Sonnet for coding, Gemini Pro for massive context. They recently added massive context. They recently added massive context. They recently added Gemini 3 and GPT 5.1 the moment they Gemini 3 and GPT 5.1 the moment they Gemini 3 and GPT 5.1 the moment they dropped. Create professional dropped. Create professional dropped. Create professional presentations with graphs, charts, and presentations with graphs, charts, and presentations with graphs, charts, and deep research detail content. Need deep research detail content. Need deep research detail content. Need human-sounding copy? Humanize rewrites human-sounding copy? Humanize rewrites human-sounding copy? Humanize rewrites text to defeat AI detectors. Need text to defeat AI detectors. Need text to defeat AI detectors. Need visuals? Pick Frontier or open-source visuals? Pick Frontier or open-source visuals? Pick Frontier or open-source models. Nano Banana, Midjourney, Flux models. Nano Banana, Midjourney, Flux models. Nano Banana, Midjourney, Flux for images, Magnific for upscaling, plus for images, Magnific for upscaling, plus for images, Magnific for upscaling, plus VEO Wan and Sora for video. All built VEO Wan and Sora for video. All built VEO Wan and Sora for video. All built in. You also get Abacus AI deep agent to in. You also get Abacus AI deep agent to in. You also get Abacus AI deep agent to pretty much do anything. Build pretty much do anything. Build pretty much do anything. Build full-stack apps, websites, reports with full-stack apps, websites, reports with full-stack apps, websites, reports with just text prompts, and deploy them on just text prompts, and deploy them on just text prompts, and deploy them on the spot. They have Abacus AI Desktop, the spot. They have Abacus AI Desktop, the spot. They have Abacus AI Desktop, which is the brand new coding editor and which is the brand new coding editor and which is the brand new coding editor and assistant that lets you vibe code and assistant that lets you vibe code and assistant that lets you vibe code and build production ready apps. And the build production ready apps. And the build production ready apps. And the kicker? It's just $10 a month, less than kicker? It's just $10 a month, less than kicker? It's just $10 a month, less than one premium model. Head over to one premium model. Head over to one premium model. Head over to chat.llm.abacus.ai chat.llm.abacus.ai chat.llm.abacus.ai or click the link below to level up with or click the link below to level up with or click the link below to level up with chat.llm teams. Just some quick specs on chat.llm teams. Just some quick specs on chat.llm teams. Just some quick specs on this particular one, HX470 like I this particular one, HX470 like I this particular one, HX470 like I mentioned, Strix Point. It's got 32 gigs mentioned, Strix Point. It's got 32 gigs mentioned, Strix Point. It's got 32 gigs of DDR5 5600 memory and mine comes with of DDR5 5600 memory and mine comes with of DDR5 5600 memory and mine comes with 2 terabytes SSD. And when it comes to 2 terabytes SSD. And when it comes to 2 terabytes SSD. And when it comes to ports, this thing is loaded. Dual USB 4, ports, this thing is loaded. Dual USB 4, ports, this thing is loaded. Dual USB 4, dual HDMI 2.1, dual 2.5 gig ethernet, dual HDMI 2.1, dual 2.5 gig ethernet, dual HDMI 2.1, dual 2.5 gig ethernet, five USB A ports, an audio jack, and an five USB A ports, an audio jack, and an five USB A ports, an audio jack, and an SD card reader on side. This actually SD card reader on side. This actually SD card reader on side. This actually pretty rare to see on a mini PC this pretty rare to see on a mini PC this pretty rare to see on a mini PC this size. Today it's LLM test using a bunch size. Today it's LLM test using a bunch size. Today it's LLM test using a bunch of models and I picked these so I have a of models and I picked these so I have a of models and I picked these so I have a reference point cuz I recently compared reference point cuz I recently compared reference point cuz I recently compared the Apple's M4 Pro with the Certan from

  3. the Apple's M4 Pro with the Certan from the Apple's M4 Pro with the Certan from Beelink, which I tested a couple weeks Beelink, which I tested a couple weeks Beelink, which I tested a couple weeks ago. That one has the same exact chip. ago. That one has the same exact chip. ago. That one has the same exact chip. First up, the OS you're already probably First up, the OS you're already probably First up, the OS you're already probably running. I start where 90% of viewers will start, I start where 90% of viewers will start, Windows and Ollama. Now, by default, Windows and Ollama. Now, by default, Windows and Ollama. Now, by default, Ollama on Windows is going to use 100% Ollama on Windows is going to use 100% Ollama on Windows is going to use 100% CPU. The GPU is idle. Ollama is just CPU. The GPU is idle. Ollama is just CPU. The GPU is idle. Ollama is just ignoring the AMD GPU entirely here. Even ignoring the AMD GPU entirely here. Even ignoring the AMD GPU entirely here. Even when Vulcan is set to true, for some when Vulcan is set to true, for some when Vulcan is set to true, for some reason Ollama is quietly failing here reason Ollama is quietly failing here reason Ollama is quietly failing here back to the CPU. So, we're going to swap back to the CPU. So, we're going to swap back to the CPU. So, we're going to swap over to llama.cpp and use Vulcan over to llama.cpp and use Vulcan over to llama.cpp and use Vulcan directly. Vulcan is the API that talks directly. Vulcan is the API that talks directly. Vulcan is the API that talks to GPU directly. And there it is, loaded to GPU directly. And there it is, loaded to GPU directly. And there it is, loaded Vulcan backend. you all start writing Vulcan backend. you all start writing Vulcan backend. you all start writing your angry comments that I'm using old your angry comments that I'm using old your angry comments that I'm using old models, first of all, it doesn't really models, first of all, it doesn't really models, first of all, it doesn't really matter. This is kind of a relative matter. This is kind of a relative matter. This is kind of a relative comparison. And second of all, here's comparison. And second of all, here's comparison. And second of all, here's Gemma 4 12B, okay? Chill out. There's Gemma 4 12B, okay? Chill out. There's Gemma 4 12B, okay? Chill out. There's always going to be new models, people, always going to be new models, people, always going to be new models, people, every single week. And the machines and every single week. And the machines and every single week. And the machines and CPUs and GPUs don't get upgrades every CPUs and GPUs don't get upgrades every CPUs and GPUs don't get upgrades every single week. So, take these model single week. So, take these model single week. So, take these model comparisons in a relative sense. There comparisons in a relative sense. There comparisons in a relative sense. There it is, running on the GPU through it is, running on the GPU through it is, running on the GPU through llama.cpp. No CPU offload on this one at llama.cpp. No CPU offload on this one at llama.cpp. No CPU offload on this one at all. We're hitting about 100% all. We're hitting about 100% all. We're hitting about 100% utilization on the GPU. So, about 10% utilization on the GPU. So, about 10% utilization on the GPU. So, about 10% gain over Ollama CPU. By the way, I have gain over Ollama CPU. By the way, I have gain over Ollama CPU. By the way, I have the charts over here. That's why I'll be the charts over here. That's why I'll be the charts over here. That's why I'll be looking over there on to the side. Now, looking over there on to the side. Now, looking over there on to the side. Now, as the models get larger, the gain as the models get larger, the gain as the models get larger, the gain shrinks because they all get slower and shrinks because they all get slower and shrinks because they all get slower and slower. So, we have 14% difference at slower. So, we have 14% difference at slower. So, we have 14% difference at 1.5 billion parameters down to 10% at 14 1.5 billion parameters down to 10% at 14 1.5 billion parameters down to 10% at 14 billion. Not really what you would billion. Not really what you would billion. Not really what you would expect going from CPU to iGPU here, but expect going from CPU to iGPU here, but expect going from CPU to iGPU here, but there's a third path on Windows there's a third path on Windows there's a third path on Windows specifically, and that's AMD's lemonade specifically, and that's AMD's lemonade specifically, and that's AMD's lemonade server. And you have hybrid recipes that

  4. server. And you have hybrid recipes that server. And you have hybrid recipes that use the NPU for pre-fill and the GPU for use the NPU for pre-fill and the GPU for use the NPU for pre-fill and the GPU for decode. Those are the two stages of decode. Those are the two stages of decode. Those are the two stages of inference. I talked more about that in inference. I talked more about that in inference. I talked more about that in other videos. Decode numbers from other videos. Decode numbers from other videos. Decode numbers from lemonade come in tied with Llama.cpp lemonade come in tied with Llama.cpp lemonade come in tied with Llama.cpp Vulcan, about a percent apart at 3 Vulcan, about a percent apart at 3 Vulcan, about a percent apart at 3 billion parameters. And the NPU shows up billion parameters. And the NPU shows up billion parameters. And the NPU shows up the moment you push long prompts, which the moment you push long prompts, which the moment you push long prompts, which we'll get into more when we're doing we'll get into more when we're doing we'll get into more when we're doing Linux. So, on Windows, Llama.cpp Vulcan Linux. So, on Windows, Llama.cpp Vulcan Linux. So, on Windows, Llama.cpp Vulcan beats default Ollama, and lemonade ties beats default Ollama, and lemonade ties beats default Ollama, and lemonade ties Vulcan on decode. So, if you want to Vulcan on decode. So, if you want to Vulcan on decode. So, if you want to utilize the iGPU on Windows, just don't utilize the iGPU on Windows, just don't utilize the iGPU on Windows, just don't go through Ollama. There's probably go through Ollama. There's probably go through Ollama. There's probably tweaks and bugs that you can do, but by tweaks and bugs that you can do, but by tweaks and bugs that you can do, but by default, it looks like it's not working default, it looks like it's not working default, it looks like it's not working very well. Go straight to Llama.cpp. very well. Go straight to Llama.cpp. very well. Go straight to Llama.cpp. Though, Though, Though, even the best Windows path is landing even the best Windows path is landing even the best Windows path is landing lower than I expected from this chip. lower than I expected from this chip. lower than I expected from this chip. Maybe it's Windows. Maybe it's Windows. Maybe it's Windows. Let's [music] try WSL, the Linux Let's [music] try WSL, the Linux Let's [music] try WSL, the Linux experience on a Windows machine. You get experience on a Windows machine. You get experience on a Windows machine. You get to it just by installing WSL first. I to it just by installing WSL first. I to it just by installing WSL first. I have other videos showing that, and then have other videos showing that, and then have other videos showing that, and then Ubuntu. Boom. Look at that. Now we're in Ubuntu. Boom. Look at that. Now we're in Ubuntu. Boom. Look at that. Now we're in Ubuntu. Oh my gosh. I installed the same Ubuntu. Oh my gosh. I installed the same Ubuntu. Oh my gosh. I installed the same exact Ollama here exact Ollama here exact Ollama here with the same exact models. I actually with the same exact models. I actually with the same exact models. I actually like WSL, and I do a lot of development like WSL, and I do a lot of development like WSL, and I do a lot of development work inside of WSL on Windows, just work inside of WSL on Windows, just work inside of WSL on Windows, just because of the Linux ergonomics. It because of the Linux ergonomics. It because of the Linux ergonomics. It makes it convenient. And there it is.

  5. makes it convenient. And there it is. makes it convenient. And there it is. It's actually doing the thing, and we're It's actually doing the thing, and we're It's actually doing the thing, and we're still executing everything on the CPU. still executing everything on the CPU. still executing everything on the CPU. So, there's no advantage here for So, there's no advantage here for So, there's no advantage here for Ollama, at least. How about Llama.cpp? Ollama, at least. How about Llama.cpp? Ollama, at least. How about Llama.cpp? Before I do that, let's run Vulcan info Before I do that, let's run Vulcan info Before I do that, let's run Vulcan info summary. And the driver is showing that summary. And the driver is showing that summary. And the driver is showing that it's a software rasterizer only. The it's a software rasterizer only. The it's a software rasterizer only. The iGPU is basically invisible. But when iGPU is basically invisible. But when iGPU is basically invisible. But when you install WSL, it tells you that it you install WSL, it tells you that it you install WSL, it tells you that it can see the GPU. So, what do we do? Some can see the GPU. So, what do we do? Some can see the GPU. So, what do we do? Some of the internet folks, you know, the of the internet folks, you know, the of the internet folks, you know, the internet folks say that WSL can't use internet folks say that WSL can't use internet folks say that WSL can't use AMD iGPUs. But that's not actually true. AMD iGPUs. But that's not actually true. AMD iGPUs. But that's not actually true. Look at what's installed. We have Asahi, Look at what's installed. We have Asahi, Look at what's installed. We have Asahi, Intel, Novell, Radeon, but we don't have Intel, Novell, Radeon, but we don't have Intel, Novell, Radeon, but we don't have dznICD.json dznICD.json dznICD.json here, which is the file we need to have here, which is the file we need to have here, which is the file we need to have that support. It's Mesa Dozen driver. that support. It's Mesa Dozen driver. that support. It's Mesa Dozen driver. Dozen? Dozen? I don't know how to Dozen? Dozen? I don't know how to Dozen? Dozen? I don't know how to pronounce that. So, that driver is not pronounce that. So, that driver is not pronounce that. So, that driver is not in the stock Ubuntu 26.04 in the stock Ubuntu 26.04 in the stock Ubuntu 26.04 package that ships with, you know, the package that ships with, you know, the package that ships with, you know, the installation here. So, first we need to installation here. So, first we need to installation here. So, first we need to add the proper repository, the package add the proper repository, the package add the proper repository, the package there. And then, once that's done, we there. And then, once that's done, we there. And then, once that's done, we install Mesa Vulcan drivers and Vulcan install Mesa Vulcan drivers and Vulcan install Mesa Vulcan drivers and Vulcan tools. Boom. And now we have it, right tools. Boom. And now we have it, right tools. Boom. And now we have it, right there. dznICD.json.

  6. there. dznICD.json. there. dznICD.json. So, now we have a GPU zero is Microsoft So, now we have a GPU zero is Microsoft So, now we have a GPU zero is Microsoft Direct3D 12 regenerating and WSL sees Direct3D 12 regenerating and WSL sees Direct3D 12 regenerating and WSL sees the iGPU now. There's a GPU. We're at the iGPU now. There's a GPU. We're at the iGPU now. There's a GPU. We're at 83%. Hmm, a little bit lower utilization 83%. Hmm, a little bit lower utilization 83%. Hmm, a little bit lower utilization here, but still here, but still here, but still it's all happening on the GPU. Well, I it's all happening on the GPU. Well, I it's all happening on the GPU. Well, I think mostly at least. There is a little think mostly at least. There is a little think mostly at least. There is a little bit of CPU usage there. CPU usage. But bit of CPU usage there. CPU usage. But bit of CPU usage there. CPU usage. But llama.cpp Vulcan picks up automatically llama.cpp Vulcan picks up automatically llama.cpp Vulcan picks up automatically and produces real Vulcan numbers here. and produces real Vulcan numbers here. and produces real Vulcan numbers here. Not just CPU fallback numbers in Not just CPU fallback numbers in Not just CPU fallback numbers in disguise. Compared to going straight disguise. Compared to going straight disguise. Compared to going straight into Windows native, WSL gives up into Windows native, WSL gives up into Windows native, WSL gives up roughly about a sixth of the throughput. roughly about a sixth of the throughput. roughly about a sixth of the throughput. That may be an acceptable price for you That may be an acceptable price for you That may be an acceptable price for you to pay if you want that Linux tool chain to pay if you want that Linux tool chain to pay if you want that Linux tool chain inside Windows. But neither path is inside Windows. But neither path is inside Windows. But neither path is getting where this chip should be getting where this chip should be getting where this chip should be landing. Weird. landing. Weird. landing. Weird. >> [music] >> [music] >> [music] >> Now, bare metal Linux. The path that >> Now, bare metal Linux. The path that >> Now, bare metal Linux. The path that nobody actually wants to take, but nobody actually wants to take, but nobody actually wants to take, but everyone tells you is faster. Let's see. everyone tells you is faster. Let's see. everyone tells you is faster. Let's see. Here I'm going to do Ollama run. So, I Here I'm going to do Ollama run. So, I Here I'm going to do Ollama run. So, I got the same exact models, the same got the same exact models, the same got the same exact models, the same exact stacks installed here on Linux as exact stacks installed here on Linux as exact stacks installed here on Linux as well. And I'm running this 14 billion well. And I'm running this 14 billion well. And I'm running this 14 billion parameter model, and it's writing out parameter model, and it's writing out parameter model, and it's writing out the story. And look at that. On Ollama, the story. And look at that. On Ollama, the story. And look at that. On Ollama, we're using 99% of the GPU. So, I ran we're using 99% of the GPU. So, I ran we're using 99% of the GPU. So, I ran all three Linux backends, Ollama, all three Linux backends, Ollama, all three Linux backends, Ollama, llama.cpp with Vulcan with radv, and llama.cpp with Vulcan with radv, and llama.cpp with Vulcan with radv, and llama.cpp with ROCm. ROCm is AMD's own llama.cpp with ROCm. ROCm is AMD's own llama.cpp with ROCm. ROCm is AMD's own thing. radv is a third-party thing. And thing. radv is a third-party thing. And thing. radv is a third-party thing. And decodes land within a few percent of decodes land within a few percent of decodes land within a few percent of each other. What? Very close. Basically, each other. What? Very close. Basically, each other. What? Very close. Basically, doesn't matter what backend you're doesn't matter what backend you're doesn't matter what backend you're using. They are all very close. And they using. They are all very close. And they using. They are all very close. And they all use the GPU. But, here's where Linux

  7. all use the GPU. But, here's where Linux all use the GPU. But, here's where Linux does pull ahead. And I hinted about this does pull ahead. And I hinted about this does pull ahead. And I hinted about this before. This is long prompt prefill at before. This is long prompt prefill at before. This is long prompt prefill at the 14 billion parameter model. This is the 14 billion parameter model. This is the 14 billion parameter model. This is a kind of workflow that matters for rag, a kind of workflow that matters for rag, a kind of workflow that matters for rag, agents, coding assistance. You saw me agents, coding assistance. You saw me agents, coding assistance. You saw me typing write a story before. Well, this typing write a story before. Well, this typing write a story before. Well, this is the opposite of that. This is like a is the opposite of that. This is like a is the opposite of that. This is like a really long prompt, and that takes time really long prompt, and that takes time really long prompt, and that takes time to process. It's like pasting in your to process. It's like pasting in your to process. It's like pasting in your entire code repo and having the LLM entire code repo and having the LLM entire code repo and having the LLM evaluate that to give you answers. And evaluate that to give you answers. And evaluate that to give you answers. And Linux with radv is roughly three times Linux with radv is roughly three times Linux with radv is roughly three times faster on Linux than Windows. So, this faster on Linux than Windows. So, this faster on Linux than Windows. So, this actually proves it. Not 3%, three times. actually proves it. Not 3%, three times. actually proves it. Not 3%, three times. And that's an open-source driver beating And that's an open-source driver beating And that's an open-source driver beating AMD's own proprietary driver on AMD's AMD's own proprietary driver on AMD's AMD's own proprietary driver on AMD's own hardware, too. So, the OS verdict so own hardware, too. So, the OS verdict so own hardware, too. So, the OS verdict so far, long prefill, Linux radv wins by far, long prefill, Linux radv wins by far, long prefill, Linux radv wins by three times. WSL is viable if you three times. WSL is viable if you three times. WSL is viable if you install kisak-mesa. By the way, I never install kisak-mesa. By the way, I never install kisak-mesa. By the way, I never knew what kisak-mesa is before doing knew what kisak-mesa is before doing knew what kisak-mesa is before doing this video and researching it. So, this video and researching it. So, this video and researching it. So, I'm not an expert in WSL. I'm not an expert in WSL. I'm not an expert in WSL. I was just asked to do this for this I was just asked to do this for this I was just asked to do this for this video. So, that's it. Windows is fine video. So, that's it. Windows is fine video. So, that's it. Windows is fine for chat if you skip Ollama. There's for chat if you skip Ollama. There's for chat if you skip Ollama. There's also LM Studio, which I didn't cover also LM Studio, which I didn't cover also LM Studio, which I didn't cover here, but it's a really nice tool that I here, but it's a really nice tool that I here, but it's a really nice tool that I prefer using when I'm on a graphical prefer using when I'm on a graphical prefer using when I'm on a graphical interface, like Windows or macOS. But, interface, like Windows or macOS. But, interface, like Windows or macOS. But, besides prefill, there's decode. Every besides prefill, there's decode. Every besides prefill, there's decode. Every OS, every backend, pretty much the same.

  8. OS, every backend, pretty much the same. OS, every backend, pretty much the same. That's not an OS problem. That is a That's not an OS problem. That is a That's not an OS problem. That is a different problem. Oh, it's not the chip because I just Oh, it's not the chip because I just tested the same chip in the Strix 10 a tested the same chip in the Strix 10 a tested the same chip in the Strix 10 a couple weeks ago. couple weeks ago. couple weeks ago. And it was fine. The GPU is at 100%. And it was fine. The GPU is at 100%. And it was fine. The GPU is at 100%. That's doing the pre-fill part. That's That's doing the pre-fill part. That's That's doing the pre-fill part. That's the calculations. The decode is much the calculations. The decode is much the calculations. The decode is much slower though. And that's a clue right slower though. And that's a clue right slower though. And that's a clue right there. That's a clue because decode there. That's a clue because decode there. That's a clue because decode happens in memory. Memory bandwidth is happens in memory. Memory bandwidth is happens in memory. Memory bandwidth is what determines slower or faster decode what determines slower or faster decode what determines slower or faster decode speed. So, I cracked it open. Two slots, speed. So, I cracked it open. Two slots, speed. So, I cracked it open. Two slots, one stick. The chip is designed to feed one stick. The chip is designed to feed one stick. The chip is designed to feed from both channels in parallel. With one from both channels in parallel. With one from both channels in parallel. With one stick, it runs at a single channel stick, it runs at a single channel stick, it runs at a single channel bandwidth. So, roughly half the bandwidth. So, roughly half the bandwidth. So, roughly half the throughput the silicon expects. That's throughput the silicon expects. That's throughput the silicon expects. That's it. This is the whole reason every back it. This is the whole reason every back it. This is the whole reason every back end on every OS hit the same wall. You end on every OS hit the same wall. You end on every OS hit the same wall. You can't compute your way around a memory can't compute your way around a memory can't compute your way around a memory bandwidth bottleneck. [music] One stick, bandwidth bottleneck. [music] One stick, bandwidth bottleneck. [music] One stick, and this is crucial for AI workloads. and this is crucial for AI workloads. and this is crucial for AI workloads. Uh Uh Uh you want to have as much memory you want to have as much memory you want to have as much memory bandwidth as possible. So, you want two bandwidth as possible. So, you want two bandwidth as possible. So, you want two chips in there. All right, we got 64 chips in there. All right, we got 64 chips in there. All right, we got 64 total system memory now, and we have two total system memory now, and we have two total system memory now, and we have two channels. This should make things a channels. This should make things a channels. This should make things a little bit speedier.

  9. little bit speedier. little bit speedier. Every single decode workload roughly Every single decode workload roughly Every single decode workload roughly doubled. I didn't make you sit through doubled. I didn't make you sit through doubled. I didn't make you sit through the whole me doing it again [music] the whole me doing it again [music] the whole me doing it again [music] thing, but here's the results. Windows thing, but here's the results. Windows thing, but here's the results. Windows Vulcan 2.13 Vulcan 2.13 Vulcan 2.13 times faster. Linux Radv over two times times faster. Linux Radv over two times times faster. Linux Radv over two times faster. Ollama with ROCm 1.86 times faster. Ollama with ROCm 1.86 times faster. Ollama with ROCm 1.86 times faster. And it keeps going. [music] You faster. And it keeps going. [music] You faster. And it keeps going. [music] You get the idea. And that's the textbook get the idea. And that's the textbook get the idea. And that's the textbook signature of a memory bandwidth bound signature of a memory bandwidth bound signature of a memory bandwidth bound workload. The compute on the GPU, workload. The compute on the GPU, workload. The compute on the GPU, [music] that was never the bottleneck. [music] that was never the bottleneck. [music] that was never the bottleneck. The RAM bus was. And here is the answer The RAM bus was. And here is the answer The RAM bus was. And here is the answer to the question I open with. Which OS to the question I open with. Which OS to the question I open with. Which OS wins? These numbers, nine on the 14 wins? These numbers, nine on the 14 wins? These numbers, nine on the 14 billion, 17 on the 7 billion, 37 on the billion, 17 on the 7 billion, 37 on the billion, 17 on the 7 billion, 37 on the 3 billion. This is what the chip is 3 billion. This is what the chip is 3 billion. This is what the chip is actually capable of. And you might have actually capable of. And you might have actually capable of. And you might have caught my video about the SER 10 machine caught my video about the SER 10 machine caught my video about the SER 10 machine from Beelink, which is a similar class from Beelink, which is a similar class from Beelink, which is a similar class machine. And the numbers are very close machine. And the numbers are very close machine. And the numbers are very close here. As you can see the single channel here. As you can see the single channel here. As you can see the single channel is on this chart, too. The GEEKOM is is on this chart, too. The GEEKOM is is on this chart, too. The GEEKOM is slightly ahead. I mean, it's within the slightly ahead. I mean, it's within the slightly ahead. I mean, it's within the margin of error here. But this is kind margin of error here. But this is kind margin of error here. But this is kind of a sanity checkpoint chart [music] of a sanity checkpoint chart [music] of a sanity checkpoint chart [music] that I used and it shows that the that I used and it shows that the that I used and it shows that the Beelink machine actually comes with dual Beelink machine actually comes with dual Beelink machine actually comes with dual channel memory already pre-installed.

  10. channel memory already pre-installed. channel memory already pre-installed. The Giga does not. And if I were to The Giga does not. And if I were to The Giga does not. And if I were to mention anything to them, if they're mention anything to them, if they're mention anything to them, if they're watching this video, I'd say watching this video, I'd say watching this video, I'd say please put in two chips in there. And please put in two chips in there. And please put in two chips in there. And just to make sure none of this is a Qwen just to make sure none of this is a Qwen just to make sure none of this is a Qwen 2.5 thing cuz I know some of you are 2.5 thing cuz I know some of you are 2.5 thing cuz I know some of you are going to complain, "Oh, it's an old going to complain, "Oh, it's an old going to complain, "Oh, it's an old model." I did Gemma 4 12 billion landed model." I did Gemma 4 12 billion landed model." I did Gemma 4 12 billion landed last week. Also, it's a totally last week. Also, it's a totally last week. Also, it's a totally different model architecture. So, yeah, different model architecture. So, yeah, different model architecture. So, yeah, it's it's a valid test. This is decode it's it's a valid test. This is decode it's it's a valid test. This is decode and you can see that Windows and Linux and you can see that Windows and Linux and you can see that Windows and Linux are pretty much tied up. WSL a little are pretty much tied up. WSL a little are pretty much tied up. WSL a little bit less. But this is just looking at a bit less. But this is just looking at a bit less. But this is just looking at a single point of data. WSL gives you a single point of data. WSL gives you a single point of data. WSL gives you a lot more. I don't want to stay here and lot more. I don't want to stay here and lot more. I don't want to stay here and defend things for people that are not defend things for people that are not defend things for people that are not going to listen anyway, but with WSL you going to listen anyway, but with WSL you going to listen anyway, but with WSL you get the Linux and the Windows, get the Linux and the Windows, get the Linux and the Windows, everything. So, you're paying a little everything. So, you're paying a little everything. So, you're paying a little bit of a price for that. All right, if bit of a price for that. All right, if bit of a price for that. All right, if you're picking something for your own you're picking something for your own you're picking something for your own tool chain, for your own workflows, that tool chain, for your own workflows, that tool chain, for your own workflows, that prompt processing, the prefill stage, prompt processing, the prefill stage, prompt processing, the prefill stage, Linux is a big winner on that one in my Linux is a big winner on that one in my Linux is a big winner on that one in my tests. And then another thing is the tests. And then another thing is the tests. And then another thing is the RAM. If you're buying any AMD mini PC, RAM. If you're buying any AMD mini PC, RAM. If you're buying any AMD mini PC, if it says single channel or 1X SODIMM, if it says single channel or 1X SODIMM, if it says single channel or 1X SODIMM, budget around 80 bucks for a matched budget around 80 bucks for a matched budget around 80 bucks for a matched pair, or let's translate that to today's pair, or let's translate that to today's pair, or let's translate that to today's money, 300 bucks. The Giga A9 Max with money, 300 bucks. The Giga A9 Max with money, 300 bucks. The Giga A9 Max with that swap is exactly the box that AMD that swap is exactly the box that AMD that swap is exactly the box that AMD spec sheet promises. Now, keep in mind spec sheet promises. Now, keep in mind spec sheet promises. Now, keep in mind when you're searching for this machine, when you're searching for this machine, when you're searching for this machine, you might have come across the A9 Max you might have come across the A9 Max you might have come across the A9 Max from last year, the one with a 370 chip from last year, the one with a 370 chip from last year, the one with a 370 chip instead of the 470 chip. They're both instead of the 470 chip. They're both instead of the 470 chip. They're both going to be available for sale and from going to be available for sale and from going to be available for sale and from personal experience, I've had the A9 Max

  11. personal experience, I've had the A9 Max personal experience, I've had the A9 Max on my desk in the other office for over on my desk in the other office for over on my desk in the other office for over a year now and I've been using it for my a year now and I've been using it for my a year now and I've been using it for my Windows development stuff. It's been a Windows development stuff. It's been a Windows development stuff. It's been a wonderful machine and physically, it's wonderful machine and physically, it's wonderful machine and physically, it's pretty much identical [music] pretty much identical [music] pretty much identical [music] to the new A9 Max. They even have the to the new A9 Max. They even have the to the new A9 Max. They even have the same exact name. Just that chip is same exact name. Just that chip is same exact name. Just that chip is different. That's it. So, you can save different. That's it. So, you can save different. That's it. So, you can save some real money here by grabbing some real money here by grabbing some real money here by grabbing yourself the older A9 Max with a 370 yourself the older A9 Max with a 370 yourself the older A9 Max with a 370 chip. Just confirm which one you're chip. Just confirm which one you're chip. Just confirm which one you're buying first. Pay attention to that very buying first. Pay attention to that very buying first. Pay attention to that very carefully. You'll be giving up about 9 carefully. You'll be giving up about 9 carefully. You'll be giving up about 9 or 10% of performance in AI tasks and or 10% of performance in AI tasks and or 10% of performance in AI tasks and maybe other tasks as well because the maybe other tasks as well because the maybe other tasks as well because the 470 is the newer chip, but also consider 470 is the newer chip, but also consider 470 is the newer chip, but also consider the price difference between the two and the price difference between the two and the price difference between the two and the fact that with the 370 you do get the fact that with the 370 you do get the fact that with the 370 you do get the two dims in there for your memory. the two dims in there for your memory. the two dims in there for your memory. Uh if you want to watch my review of the Uh if you want to watch my review of the Uh if you want to watch my review of the dev tests on the this year 10 from last dev tests on the this year 10 from last dev tests on the this year 10 from last week [music] right here and the previous week [music] right here and the previous week [music] right here and the previous version of the A9 Max right over here. version of the A9 Max right over here. version of the A9 Max right over here. [music] Thanks for watching and I'll see [music] Thanks for watching and I'll see [music] Thanks for watching and I'll see you next time.

Summary

The main theme is the comparison of the GEEKOM A9 Max with its new AMD Ryzen AI 9 HX 470 APU, benchmarking performance across Windows, WSL, and bare-metal Linux for AI tasks. The conclusion highlights the introduction of Chat LLM Teams, a unified platform for various AI models and content creation tools, presenting a practical solution for consolidating AI workflows. The takeaway is that while the hardware is impressive, the true innovation lies in the streamlined management and accessibility of diverse AI capabilities through integrated software platforms.

View original episode ↗