← Back
Aaron Zisk June 16, 2026 1m

AMD's Strix Successor Just Caught the M4 Pro

Read full transcript 2 segments
  1. I've spent a bit of time testing this I've spent a bit of time testing this brand new mini PC, the SIR 10 from brand new mini PC, the SIR 10 from brand new mini PC, the SIR 10 from Beink. [music] Beink. [music] Beink. [music] So, this little box has three chips on So, this little box has three chips on So, this little box has three chips on it and they all can do AI. There's a 12 it and they all can do AI. There's a 12 it and they all can do AI. There's a 12 core CPU, Radeon 890MIG GPU, and there's core CPU, Radeon 890MIG GPU, and there's core CPU, Radeon 890MIG GPU, and there's a 55 [music] tops XDNA 2 NPU. AMD says a 55 [music] tops XDNA 2 NPU. AMD says a 55 [music] tops XDNA 2 NPU. AMD says all three together do 86 tops of AI all three together do 86 tops of AI all three together do 86 tops of AI work. Tops. Yeah, we all know the work. Tops. Yeah, we all know the work. Tops. Yeah, we all know the marketing stuff. So, I ran a couple of marketing stuff. So, I ran a couple of marketing stuff. So, I ran a couple of them as a test. Llama [music] 3.23B, 23B them as a test. Llama [music] 3.23B, 23B them as a test. Llama [music] 3.23B, 23B we went from 27 tokens per second to 37 we went from 27 tokens per second to 37 we went from 27 tokens per second to 37 1/2. Quen 2.5 1.5B went from 47.5 1/2. Quen 2.5 1.5B went from 47.5 1/2. Quen 2.5 1.5B went from 47.5 to 68.3 [music] tokens per second. So to 68.3 [music] tokens per second. So to 68.3 [music] tokens per second. So there's some differences there, but there's some differences there, but there's some differences there, but there's still one chip in this box we there's still one chip in this box we there's still one chip in this box we haven't even touched, and that's this haven't even touched, and that's this haven't even touched, and that's this NPU. So first try, I used Lemonade NPU. So first try, I used Lemonade NPU. So first try, I used Lemonade Server. They got these hybrid recipes Server. They got these hybrid recipes Server. They got these hybrid recipes where the MPU does the prefill and the where the MPU does the prefill and the where the MPU does the prefill and the iGPU does the decode. Those are the two iGPU does the decode. Those are the two iGPU does the decode. Those are the two stages of inference. It even comes with stages of inference. It even comes with stages of inference. It even comes with this little handy app. Let's do a high this little handy app. Let's do a high this little handy app. Let's do a high here. 14.3 on this one. And I ran a here. 14.3 on this one. And I ran a here. 14.3 on this one. And I ran a longer prompt. And that's the actual longer prompt. And that's the actual longer prompt. And that's the actual prefill workload on Quen 7B. The CPU got prefill workload on Quen 7B. The CPU got prefill workload on Quen 7B. The CPU got 255 tokens per second there. The Vulcan 255 tokens per second there. The Vulcan 255 tokens per second there. The Vulcan iGPU 240. And the NPU with the hybrid iGPU 240. And the NPU with the hybrid iGPU 240. And the NPU with the hybrid mode 631 tokens per second. That's 2 and mode 631 tokens per second. That's 2 and mode 631 tokens per second. That's 2 and 1/2 times the iGPU on the same chip. So 1/2 times the iGPU on the same chip. So 1/2 times the iGPU on the same chip. So now we know what each chip wants to do.

  2. now we know what each chip wants to do. now we know what each chip wants to do. And CPU is the fall back. IGPU does And CPU is the fall back. IGPU does And CPU is the fall back. IGPU does streaming chat. NPU does long prompts. streaming chat. NPU does long prompts. streaming chat. NPU does long prompts. So, quick recommendations here. You So, quick recommendations here. You So, quick recommendations here. You should buy the S 10 if you care about should buy the S 10 if you care about should buy the S 10 if you care about the MPU for rag and long context AI the MPU for rag and long context AI the MPU for rag and long context AI work. If you want upgradeable RAM, or if work. If you want upgradeable RAM, or if work. If you want upgradeable RAM, or if you want 10 gig Ethernet port, you you want 10 gig Ethernet port, you you want 10 gig Ethernet port, you should probably skip it if you already should probably skip it if you already should probably skip it if you already have the S 9 and your main workload is have the S 9 and your main workload is have the S 9 and your main workload is the iGPU LLM path. Thanks for watching the iGPU LLM path. Thanks for watching the iGPU LLM path. Thanks for watching and I'll see you next

Summary

This analysis focuses on the Beink SIR 10 mini PC, highlighting its three AI-capable chips: CPU, GPU, and NPU, and their performance with AI models like Llama and Quiken. The practical takeaway is that the NPU excels at long prompts and RAG workloads, making the SIR 10 a worthwhile purchase for those prioritizing this functionality, while the iGPU is better for streaming chat.

View original episode ↗