AMD Built the DGX Spark Rival I Predicted… But There's a Catch
Read full transcript 14 segments
-
Maybe other vendors will also catch up Maybe other vendors will also catch up like AMD and create their own tiny like AMD and create their own tiny like AMD and create their own tiny little AMD box. Well, they did. This is little AMD box. Well, they did. This is little AMD box. Well, they did. This is AMD's answer to the DJX Spark. Brand AMD's answer to the DJX Spark. Brand AMD's answer to the DJX Spark. Brand new, just came out, and I've got a few new, just came out, and I've got a few new, just came out, and I've got a few things to say about it. But first, a things to say about it. But first, a things to say about it. But first, a quick word from the sponsor of this quick word from the sponsor of this quick word from the sponsor of this video, Merlin AI. It's an all-in-one AI video, Merlin AI. It's an all-in-one AI video, Merlin AI. It's an all-in-one AI tool, and they gave my audience a big tool, and they gave my audience a big tool, and they gave my audience a big discount. I keep multiple AI tools discount. I keep multiple AI tools discount. I keep multiple AI tools around because each one is good at around because each one is good at around because each one is good at something, but it gets really expensive, something, but it gets really expensive, something, but it gets really expensive, and bouncing between tabs breaks my and bouncing between tabs breaks my and bouncing between tabs breaks my focus. Merlin AI puts Chad GPT, Claw, focus. Merlin AI puts Chad GPT, Claw, focus. Merlin AI puts Chad GPT, Claw, Gemini, and more in one place so I can Gemini, and more in one place so I can Gemini, and more in one place so I can pick the best one for the moment, pick the best one for the moment, pick the best one for the moment, whether I'm coding, researching, or whether I'm coding, researching, or whether I'm coding, researching, or writing for a video. Watch this. I click writing for a video. Watch this. I click writing for a video. Watch this. I click the Merlin AI extension, chat with the the Merlin AI extension, chat with the the Merlin AI extension, chat with the web page to summarize what I'm reading, web page to summarize what I'm reading, web page to summarize what I'm reading, and pull out the important parts. And I and pull out the important parts. And I and pull out the important parts. And I even have my choice of models right at even have my choice of models right at even have my choice of models right at my fingertips. If I need something my fingertips. If I need something my fingertips. If I need something deeper, I turn on deep research, and it deeper, I turn on deep research, and it deeper, I turn on deep research, and it builds a clean, structured report from builds a clean, structured report from builds a clean, structured report from multiple sources. And it also has quick multiple sources. And it also has quick multiple sources. And it also has quick modes like web, academic, and Reddit modes like web, academic, and Reddit modes like web, academic, and Reddit search. If you pay separately, Chad GPT search. If you pay separately, Chad GPT search. If you pay separately, Chad GPT is $20. Claude is $20, Gemini is $20, is $20. Claude is $20, Gemini is $20, is $20. Claude is $20, Gemini is $20, and that adds up fast. Merlin AI is and that adds up fast. Merlin AI is and that adds up fast. Merlin AI is cheaper because they buy AI API access cheaper because they buy AI API access cheaper because they buy AI API access in [music] bulk. APIs cost less than the in [music] bulk. APIs cost less than the in [music] bulk. APIs cost less than the $20 plans. And most people don't even $20 plans. And most people don't even $20 plans. And most people don't even use $20 worth of API in a month. And use $20 worth of API in a month. And use $20 worth of API in a month. And here's the discount. I click pricing, here's the discount. I click pricing, here's the discount. I click pricing, continue, it takes me to Stripe, I enter continue, it takes me to Stripe, I enter continue, it takes me to Stripe, I enter the promo code, and the total drops to the promo code, and the total drops to the promo code, and the total drops to $60 for the year. That's basically five $60 for the year. That's basically five $60 for the year. That's basically five bucks a month. I don't know how long bucks a month. I don't know how long bucks a month. I don't know how long this deal will be available, so grab it this deal will be available, so grab it this deal will be available, so grab it soon. The link is in the description. If soon. The link is in the description. If soon. The link is in the description. If you watch this channel for a while, you you watch this channel for a while, you you watch this channel for a while, you already know the chip that's inside of
-
already know the chip that's inside of already know the chip that's inside of this thing pretty well. All these came this thing pretty well. All these came this thing pretty well. All these came out in 2025. The Beink GTR9, the GMK out in 2025. The Beink GTR9, the GMK out in 2025. The Beink GTR9, the GMK Techch Evo X2, the Framework Desktop, Techch Evo X2, the Framework Desktop, Techch Evo X2, the Framework Desktop, Minis Forum, MS1 Minis Forum, MS1 Minis Forum, MS1 Max, and of course, ASUS has it in a Max, and of course, ASUS has it in a Max, and of course, ASUS has it in a laptop form as well. These are general laptop form as well. These are general laptop form as well. These are general purpose Stricks Halo boxes. Stricks Halo purpose Stricks Halo boxes. Stricks Halo purpose Stricks Halo boxes. Stricks Halo is the name of the chip that's inside of is the name of the chip that's inside of is the name of the chip that's inside of these. The official name is Ryzen AI Max these. The official name is Ryzen AI Max these. The official name is Ryzen AI Max Plus 395. So AMD took that exact chip Plus 395. So AMD took that exact chip Plus 395. So AMD took that exact chip and put it in this tiny little dedicated and put it in this tiny little dedicated and put it in this tiny little dedicated box and charged 4 grand for it. Yeah, box and charged 4 grand for it. Yeah, box and charged 4 grand for it. Yeah, that's the same price that Nvidia that's the same price that Nvidia that's the same price that Nvidia charged for the DJX Spark when it came charged for the DJX Spark when it came charged for the DJX Spark when it came out. And my first reaction was, what? out. And my first reaction was, what? out. And my first reaction was, what? 4,000 bucks for a chip I can already 4,000 bucks for a chip I can already 4,000 bucks for a chip I can already have cheaper. All right, look. What? have cheaper. All right, look. What? have cheaper. All right, look. What? Okay. Uh, surely the framework 128 2 TB. Okay. Uh, surely the framework 128 2 TB. Okay. Uh, surely the framework 128 2 TB. Oh, that's without any expansion. Well, Oh, that's without any expansion. Well, Oh, that's without any expansion. Well, the minutes form came down a little bit. the minutes form came down a little bit. the minutes form came down a little bit. This is the retail right now. And the This is the retail right now. And the This is the retail right now. And the GMK Techch Evo X2 128 also came down a GMK Techch Evo X2 128 also came down a GMK Techch Evo X2 128 also came down a bit from 4,000. So the whole market kind bit from 4,000. So the whole market kind bit from 4,000. So the whole market kind of quietly drifted up to 4 grand while of quietly drifted up to 4 grand while of quietly drifted up to 4 grand while nobody was looking during the past year. nobody was looking during the past year. nobody was looking during the past year. In other words, AMD is not gouging you In other words, AMD is not gouging you In other words, AMD is not gouging you here. Uh they're just right where here. Uh they're just right where here. Uh they're just right where everybody else is. And by the way, this everybody else is. And by the way, this everybody else is. And by the way, this is now officially the smallest machine is now officially the smallest machine is now officially the smallest machine that I've got. Smaller than the that I've got. Smaller than the that I've got. Smaller than the framework, much smaller than all these framework, much smaller than all these framework, much smaller than all these other streaks Halo machines, and even a other streaks Halo machines, and even a other streaks Halo machines, and even a touch smaller than the Spark. It packs touch smaller than the Spark. It packs touch smaller than the Spark. It packs 128 gigs of unified memory, which means 128 gigs of unified memory, which means 128 gigs of unified memory, which means memory that can be used by the GPU into memory that can be used by the GPU into memory that can be used by the GPU into that. So, the question today is not the that. So, the question today is not the that. So, the question today is not the chip. We already know the chip, it's chip. We already know the chip, it's chip. We already know the chip, it's everything around it. Did AMD actually everything around it. Did AMD actually everything around it. Did AMD actually close the gap with the Spark, not on close the gap with the Spark, not on close the gap with the Spark, not on silicon, but on the stuff that matters
-
silicon, but on the stuff that matters silicon, but on the stuff that matters when you sit down and use it. That's when you sit down and use it. That's when you sit down and use it. That's what I wanted to find out. Now, with what I wanted to find out. Now, with what I wanted to find out. Now, with this Ryzen AI Halo machine, AMD is this Ryzen AI Halo machine, AMD is this Ryzen AI Halo machine, AMD is targeting AI deaths, just like the DJX targeting AI deaths, just like the DJX targeting AI deaths, just like the DJX Spark was. And to me, there are three Spark was. And to me, there are three Spark was. And to me, there are three groups of people looking at this thing groups of people looking at this thing groups of people looking at this thing and wondering if they should fork over and wondering if they should fork over and wondering if they should fork over the 4 grand for it. [music] Group one, the 4 grand for it. [music] Group one, the 4 grand for it. [music] Group one, developers, my people. And I'm going to developers, my people. And I'm going to developers, my people. And I'm going to make this real quick because there make this real quick because there make this real quick because there really isn't much of a story here. For really isn't much of a story here. For really isn't much of a story here. For regular dev work, building code, running regular dev work, building code, running regular dev work, building code, running your stack, all the CPUbound stuff, the your stack, all the CPUbound stuff, the your stack, all the CPUbound stuff, the Spark and the M4 Pro are basically all Spark and the M4 Pro are basically all Spark and the M4 Pro are basically all very close. To get a sense of this, I very close. To get a sense of this, I very close. To get a sense of this, I ran my Mandle Broad test, which is a ran my Mandle Broad test, which is a ran my Mandle Broad test, which is a Python interpretive code, which pegs all Python interpretive code, which pegs all Python interpretive code, which pegs all the cores all at once. One and two and the cores all at once. One and two and the cores all at once. One and two and three. All right, we got some noise three. All right, we got some noise three. All right, we got some noise going on here. And you can see every going on here. And you can see every going on here. And you can see every single core of every single machine is single core of every single machine is single core of every single machine is being used. And this is what it looks being used. And this is what it looks being used. And this is what it looks like while running. 164 watts on the like while running. 164 watts on the like while running. 164 watts on the Halo, 164 on the Spark. Very close. And Halo, 164 on the Spark. Very close. And Halo, 164 on the Spark. Very close. And about 80 to 90 on the Mac. Since this is about 80 to 90 on the Mac. Since this is about 80 to 90 on the Mac. Since this is a multi-core test, they all can do code a multi-core test, they all can do code a multi-core test, they all can do code compilation pretty similarly. We got compilation pretty similarly. We got compilation pretty similarly. We got 17.3 on the Mac, 15.4 on the Spark, and 17.3 on the Mac, 15.4 on the Spark, and 17.3 on the Mac, 15.4 on the Spark, and 18.4 on the Halo. Of course, there's 18.4 on the Halo. Of course, there's 18.4 on the Halo. Of course, there's nuances here and there, but I don't nuances here and there, but I don't nuances here and there, but I don't think that this is the primary target think that this is the primary target think that this is the primary target audience for this machine, even though audience for this machine, even though audience for this machine, even though it can be used for that purpose. But it can be used for that purpose. But it can be used for that purpose. But should you buy a $4,000 machine just to should you buy a $4,000 machine just to should you buy a $4,000 machine just to compile code? Come on, you know the compile code? Come on, you know the compile code? Come on, you know the answer to that. There are way cheaper answer to that. There are way cheaper answer to that. There are way cheaper alternatives for that if that's all you alternatives for that if that's all you alternatives for that if that's all you need, this isn't the one for that.
-
need, this isn't the one for that. need, this isn't the one for that. That's not why you're here anyway. That's not why you're here anyway. That's not why you're here anyway. You're here for the AI stuff. So, that You're here for the AI stuff. So, that You're here for the AI stuff. So, that brings me to the next group. brings me to the next group. brings me to the next group. All right, group two, the tinkerers. And All right, group two, the tinkerers. And All right, group two, the tinkerers. And this is where it gets good because this this is where it gets good because this this is where it gets good because this box actually comes with all the software box actually comes with all the software box actually comes with all the software pre-installed. LM Studio Olama both pre-installed. LM Studio Olama both pre-installed. LM Studio Olama both running with the Llama CPP under the running with the Llama CPP under the running with the Llama CPP under the hood. So Llama CPP is built on it. hood. So Llama CPP is built on it. hood. So Llama CPP is built on it. Lemonade is included. Comfy UI is Lemonade is included. Comfy UI is Lemonade is included. Comfy UI is included. So I thought I'd do a little included. So I thought I'd do a little included. So I thought I'd do a little performance measurements and I've got performance measurements and I've got performance measurements and I've got Llama Beni to be the client. By the way, Llama Beni to be the client. By the way, Llama Beni to be the client. By the way, if you're not familiar, check out Llama if you're not familiar, check out Llama if you're not familiar, check out Llama Beni. This is different than Llama Beni. This is different than Llama Beni. This is different than Llama Bench. Llama Bench is part of Llama CPP. Bench. Llama Bench is part of Llama CPP. Bench. Llama Bench is part of Llama CPP. It gets built with it. However, LA bench It gets built with it. However, LA bench It gets built with it. However, LA bench can be used with other inference engines can be used with other inference engines can be used with other inference engines like VLM and SG Lang and a bunch of like VLM and SG Lang and a bunch of like VLM and SG Lang and a bunch of other ones. I'm running Gemma 412B on other ones. I'm running Gemma 412B on other ones. I'm running Gemma 412B on all these machines. You can see the GPU all these machines. You can see the GPU all these machines. You can see the GPU output right here. Halo Spark Mac. And output right here. Halo Spark Mac. And output right here. Halo Spark Mac. And boom. And there they go. They do prefill boom. And there they go. They do prefill boom. And there they go. They do prefill first, then immediately switch to first, then immediately switch to first, then immediately switch to decode. And this is a race, folks. So, decode. And this is a race, folks. So, decode. And this is a race, folks. So, here we go. You can see the GPU usage here we go. You can see the GPU usage here we go. You can see the GPU usage pretty high on all these. 97% here, 95% pretty high on all these. 97% here, 95% pretty high on all these. 97% here, 95% on the Spark, and almost 100% on the on the Spark, and almost 100% on the on the Spark, and almost 100% on the Mac. The Mac is actually done now.
-
Mac. The Mac is actually done now. Mac. The Mac is actually done now. Slightly different technology. On the Slightly different technology. On the Slightly different technology. On the Mac, we're using MLX. On the two other Mac, we're using MLX. On the two other Mac, we're using MLX. On the two other machines, we're using Llama CPP. That's machines, we're using Llama CPP. That's machines, we're using Llama CPP. That's the GGUF guff models. And I've shown in the GGUF guff models. And I've shown in the GGUF guff models. And I've shown in previous videos that MLX is a little bit previous videos that MLX is a little bit previous videos that MLX is a little bit faster than the GGUF models on the Mac. faster than the GGUF models on the Mac. faster than the GGUF models on the Mac. And we definitely see that in the live And we definitely see that in the live And we definitely see that in the live chart here as it's running. Oh, I can chart here as it's running. Oh, I can chart here as it's running. Oh, I can hear the fan. That's the new AMD Halo. hear the fan. That's the new AMD Halo. hear the fan. That's the new AMD Halo. It's got a little bit of a fan going on It's got a little bit of a fan going on It's got a little bit of a fan going on right now. And this is what the right now. And this is what the right now. And this is what the temperatures look like. By the way, the temperatures look like. By the way, the temperatures look like. By the way, the Halo Box has ports venting on the top of Halo Box has ports venting on the top of Halo Box has ports venting on the top of the machine, which is interesting, which the machine, which is interesting, which the machine, which is interesting, which is going to be limiting. You can't stack is going to be limiting. You can't stack is going to be limiting. You can't stack them like you can with the Sparks. them like you can with the Sparks. them like you can with the Sparks. Sparks go from back to front. Although, Sparks go from back to front. Although, Sparks go from back to front. Although, stacking sparks is probably not great stacking sparks is probably not great stacking sparks is probably not great either cuz it's probably the hottest either cuz it's probably the hottest either cuz it's probably the hottest machine here. So, you'd want a little machine here. So, you'd want a little machine here. So, you'd want a little bit of space in between them. But these bit of space in between them. But these bit of space in between them. But these definitely you don't want anything definitely you don't want anything definitely you don't want anything around it because there's vents on all around it because there's vents on all around it because there's vents on all the sides of the halo. What do we got? the sides of the halo. What do we got? the sides of the halo. What do we got? Okay, Llama gives us two numbers that Okay, Llama gives us two numbers that Okay, Llama gives us two numbers that matter. Token generation is how fast it matter. Token generation is how fast it matter. Token generation is how fast it actually spits out and generates the actually spits out and generates the actually spits out and generates the text. That one is memory bandwidth text. That one is memory bandwidth text. That one is memory bandwidth bound. So it has to do with memory. Both bound. So it has to do with memory. Both bound. So it has to do with memory. Both the DJX Spark and the M4 Pro have a the DJX Spark and the M4 Pro have a the DJX Spark and the M4 Pro have a memory bandwidth of 273 GB per second.
-
memory bandwidth of 273 GB per second. memory bandwidth of 273 GB per second. And the Halo has 256 GB per second, so And the Halo has 256 GB per second, so And the Halo has 256 GB per second, so slightly less. And this actually slightly less. And this actually slightly less. And this actually reflects in the token generation speed. reflects in the token generation speed. reflects in the token generation speed. So we get 26.4 on the DGX Spark, 24.6 So we get 26.4 on the DGX Spark, 24.6 So we get 26.4 on the DGX Spark, 24.6 tokens per second on the Ryzen Halo Box. tokens per second on the Ryzen Halo Box. tokens per second on the Ryzen Halo Box. just a little bit higher on the DJX just a little bit higher on the DJX just a little bit higher on the DJX Spark, but the uh M4 Pro actually beats Spark, but the uh M4 Pro actually beats Spark, but the uh M4 Pro actually beats them both. 33.8 tokens per second. them both. 33.8 tokens per second. them both. 33.8 tokens per second. They're all pretty close. The other They're all pretty close. The other They're all pretty close. The other number you should look at is prefill or number you should look at is prefill or number you should look at is prefill or prompt processing. That's the compute prompt processing. That's the compute prompt processing. That's the compute heavy part. That's what the GPU does. heavy part. That's what the GPU does. heavy part. That's what the GPU does. So, as far as generating tokens, the So, as far as generating tokens, the So, as far as generating tokens, the Halo actually keeps up with the Spark Halo actually keeps up with the Spark Halo actually keeps up with the Spark pretty nicely. That's the number you can pretty nicely. That's the number you can pretty nicely. That's the number you can live with if you're chatting with it and live with if you're chatting with it and live with if you're chatting with it and you're a tinkerer. By the way, I ran you're a tinkerer. By the way, I ran you're a tinkerer. By the way, I ran this previously. That's why I have this previously. That's why I have this previously. That's why I have slightly different numbers on the chart slightly different numbers on the chart slightly different numbers on the chart here. And in that case, the Halo here. And in that case, the Halo here. And in that case, the Halo actually beat the Spark. So, they're actually beat the Spark. So, they're actually beat the Spark. So, they're very close. Now, prefill. This is the very close. Now, prefill. This is the very close. Now, prefill. This is the one where the Spark's CUDA pedigree one where the Spark's CUDA pedigree one where the Spark's CUDA pedigree shows up. And I'm going to have to be shows up. And I'm going to have to be shows up. And I'm going to have to be straight about it with you here. I've straight about it with you here. I've straight about it with you here. I've shown this before on the channel. The shown this before on the channel. The shown this before on the channel. The Halo does about 650. The Spark 2000, Halo does about 650. The Spark 2000, Halo does about 650. The Spark 2000, let's call it three times faster. let's call it three times faster. let's call it three times faster. However, this is actually almost double However, this is actually almost double However, this is actually almost double the speed that I got on the exact same the speed that I got on the exact same the speed that I got on the exact same chip back last year. So, even though chip back last year. So, even though chip back last year. So, even though it's the same chip, the advances in it's the same chip, the advances in it's the same chip, the advances in software since then have actually software since then have actually software since then have actually enabled faster prefill on this chip.
-
enabled faster prefill on this chip. enabled faster prefill on this chip. Software matters a lot, and AMD has Software matters a lot, and AMD has Software matters a lot, and AMD has really been working hard on this. Now, really been working hard on this. Now, really been working hard on this. Now, for most of what group two people will for most of what group two people will for most of what group two people will do, the pre-fill gap just doesn't reach do, the pre-fill gap just doesn't reach do, the pre-fill gap just doesn't reach you at all. If you're chatting, if you at all. If you're chatting, if you at all. If you're chatting, if you're streaming responses or if you're you're streaming responses or if you're you're streaming responses or if you're running an agent in your editor, you running an agent in your editor, you running an agent in your editor, you basically will never feel it. Where you basically will never feel it. Where you basically will never feel it. Where you will feel it is the compute heavy stuff will feel it is the compute heavy stuff will feel it is the compute heavy stuff like giant context dumps or pure image like giant context dumps or pure image like giant context dumps or pure image and video work. I did a stable diffusion and video work. I did a stable diffusion and video work. I did a stable diffusion test and here the DJX Spark was doing test and here the DJX Spark was doing test and here the DJX Spark was doing about 2.9 iterations a second in the about 2.9 iterations a second in the about 2.9 iterations a second in the Halo 1.3 and I had a hard time getting Halo 1.3 and I had a hard time getting Halo 1.3 and I had a hard time getting it to work on the Mac at all. Uh there it to work on the Mac at all. Uh there it to work on the Mac at all. Uh there is a way. It's just uh I'm still working is a way. It's just uh I'm still working is a way. It's just uh I'm still working on it and didn't get it working. Bad on it and didn't get it working. Bad on it and didn't get it working. Bad Alex. Bad Alex. But I didn't expect too Alex. Bad Alex. But I didn't expect too Alex. Bad Alex. But I didn't expect too much from the Mac cuz it's really good much from the Mac cuz it's really good much from the Mac cuz it's really good at generating tokens. It doesn't have at generating tokens. It doesn't have at generating tokens. It doesn't have that very compute heavy and comput that very compute heavy and comput that very compute heavy and comput intensive GPU in there. So it'll do the intensive GPU in there. So it'll do the intensive GPU in there. So it'll do the job. But let's keep going, shall we? job. But let's keep going, shall we? job. But let's keep going, shall we? Next time, get ready for the video Next time, get ready for the video Next time, get ready for the video better. better. better. >> Thanks. >> Thanks. >> Thanks. >> Nice preparation, Alex. >> Nice preparation, Alex. >> Nice preparation, Alex. >> Okay. I also did WAN 2.2 video >> Okay. I also did WAN 2.2 video >> Okay. I also did WAN 2.2 video generation. This one was uh kind of generation. This one was uh kind of generation. This one was uh kind of nasty. [laughter] A 5second clip on the nasty. [laughter] A 5second clip on the nasty. [laughter] A 5second clip on the Spark takes about 5 1/2 minutes and 75 Spark takes about 5 1/2 minutes and 75 Spark takes about 5 1/2 minutes and 75 minutes on the Halo. Yeah. Um I was minutes on the Halo. Yeah. Um I was minutes on the Halo. Yeah. Um I was expecting a little bit faster on the expecting a little bit faster on the expecting a little bit faster on the Halo, but TBD TBD. Oh, uh just a note, Halo, but TBD TBD. Oh, uh just a note, Halo, but TBD TBD. Oh, uh just a note, this is Windows also. Um, there is an this is Windows also. Um, there is an this is Windows also. Um, there is an catch to that and I showed this in a catch to that and I showed this in a catch to that and I showed this in a recent video about specifically AMD recent video about specifically AMD recent video about specifically AMD chips and how they behave with Linux chips and how they behave with Linux chips and how they behave with Linux versus Windows in prefill scenarios.
-
versus Windows in prefill scenarios. versus Windows in prefill scenarios. That's another video. I'll link to that That's another video. I'll link to that That's another video. I'll link to that down below. But hold on, hold on. I got down below. But hold on, hold on. I got down below. But hold on, hold on. I got something coming up. You just wait. something coming up. You just wait. something coming up. You just wait. Anyway, none of this should surprise you Anyway, none of this should surprise you Anyway, none of this should surprise you really. The Spark is a CUDA machine top really. The Spark is a CUDA machine top really. The Spark is a CUDA machine top to bottom. That's what it's for. The to bottom. That's what it's for. The to bottom. That's what it's for. The Halo was never trying to win a computer Halo was never trying to win a computer Halo was never trying to win a computer drag race here. It's a different kind of drag race here. It's a different kind of drag race here. It's a different kind of box. So, let's keep going. Now, for the box. So, let's keep going. Now, for the box. So, let's keep going. Now, for the twist. AMD put out their own benchmarks. twist. AMD put out their own benchmarks. twist. AMD put out their own benchmarks. It's right here. I'm looking over here It's right here. I'm looking over here It's right here. I'm looking over here cuz that's where my laptop is. If you cuz that's where my laptop is. If you cuz that's where my laptop is. If you just go to the homepage of the Ryzen AI just go to the homepage of the Ryzen AI just go to the homepage of the Ryzen AI Halo, you'll see everything about it. By Halo, you'll see everything about it. By Halo, you'll see everything about it. By the way, available at MicroEnter. You the way, available at MicroEnter. You the way, available at MicroEnter. You can pre-order it right now. The ultimate can pre-order it right now. The ultimate can pre-order it right now. The ultimate AI dev platform. And these are exactly AI dev platform. And these are exactly AI dev platform. And these are exactly the machines they're comparing it the machines they're comparing it the machines they're comparing it against. Hm. According to AMD, the Halo against. Hm. According to AMD, the Halo against. Hm. According to AMD, the Halo doesn't just keep up with the Spark, it doesn't just keep up with the Spark, it doesn't just keep up with the Spark, it beats it in certain situations. Look at beats it in certain situations. Look at beats it in certain situations. Look at this token generation head-to-head with this token generation head-to-head with this token generation head-to-head with Spark on the GLM Flash 30B model up to Spark on the GLM Flash 30B model up to Spark on the GLM Flash 30B model up to 14% faster. Quen 3.51 122B 12% faster. 14% faster. Quen 3.51 122B 12% faster. 14% faster. Quen 3.51 122B 12% faster. GPTOSS 12B 7% faster and Quen 3.635B GPTOSS 12B 7% faster and Quen 3.635B GPTOSS 12B 7% faster and Quen 3.635B 4% faster. Across the board, Halo wins 4% faster. Across the board, Halo wins 4% faster. Across the board, Halo wins in these particular models. And if we go in these particular models. And if we go in these particular models. And if we go slightly up here, they're doing image slightly up here, they're doing image slightly up here, they're doing image generation against the Apple M4 Pro generation against the Apple M4 Pro generation against the Apple M4 Pro three, four, even up to seven times three, four, even up to seven times three, four, even up to seven times faster. So, case closed, right? AMD faster. So, case closed, right? AMD faster. So, case closed, right? AMD wins. Now, hold on. Let me just say up wins. Now, hold on. Let me just say up wins. Now, hold on. Let me just say up front that I'm glad AMD actually put front that I'm glad AMD actually put front that I'm glad AMD actually put real numbers out there at all. Most real numbers out there at all. Most real numbers out there at all. Most companies just give you like a chart and companies just give you like a chart and companies just give you like a chart and a vibe and a render. But every benchmark a vibe and a render. But every benchmark a vibe and a render. But every benchmark slide ever made is a company showing you
-
slide ever made is a company showing you slide ever made is a company showing you their best angle. So, the best we can do their best angle. So, the best we can do their best angle. So, the best we can do is just read it carefully and fill in is just read it carefully and fill in is just read it carefully and fill in the parts that didn't make the slide the parts that didn't make the slide the parts that didn't make the slide ourselves. Now, first of all, against ourselves. Now, first of all, against ourselves. Now, first of all, against the Spark, AMD shows only token the Spark, AMD shows only token the Spark, AMD shows only token generation here because that's the generation here because that's the generation here because that's the number that Halo is really strong at. number that Halo is really strong at. number that Halo is really strong at. So, I get it. But they skip prefill So, I get it. But they skip prefill So, I get it. But they skip prefill here. That's the number where the spark here. That's the number where the spark here. That's the number where the spark is three times ahead in certain cases on is three times ahead in certain cases on is three times ahead in certain cases on certain models. So, here you're seeing certain models. So, here you're seeing certain models. So, here you're seeing the half that looks best. And that tells the half that looks best. And that tells the half that looks best. And that tells me that group two, the tinkerers, is me that group two, the tinkerers, is me that group two, the tinkerers, is actually what this is for because actually what this is for because actually what this is for because prefill is more useful in professional prefill is more useful in professional prefill is more useful in professional situations. Second, the image generation situations. Second, the image generation situations. Second, the image generation stuff. This is not against the Spark at stuff. This is not against the Spark at stuff. This is not against the Spark at all. This is against the Apple M4 Pro, all. This is against the Apple M4 Pro, all. This is against the Apple M4 Pro, which makes sense once you remember that which makes sense once you remember that which makes sense once you remember that the Spark actually wins image generation the Spark actually wins image generation the Spark actually wins image generation 2:1 and video by 16 to1 in my case. I'll 2:1 and video by 16 to1 in my case. I'll 2:1 and video by 16 to1 in my case. I'll have to test Linux later. So, here have to test Linux later. So, here have to test Linux later. So, here you're seeing a real win for the Halo you're seeing a real win for the Halo you're seeing a real win for the Halo just against different opponents than just against different opponents than just against different opponents than the headline implies. Third, if you the headline implies. Third, if you the headline implies. Third, if you notice here, they're using mixture of notice here, they're using mixture of notice here, they're using mixture of experts models here. And that happens to experts models here. And that happens to experts models here. And that happens to work really well with 128 gig machines work really well with 128 gig machines work really well with 128 gig machines cuz you can fit pretty big models like cuz you can fit pretty big models like cuz you can fit pretty big models like GPTOSS 12B and Quen 3.5122B. But the GPTOSS 12B and Quen 3.5122B. But the GPTOSS 12B and Quen 3.5122B. But the number of active parameters is that number of active parameters is that number of active parameters is that second number. I don't know how many second number. I don't know how many second number. I don't know how many active parameters 12B has, but Quen active parameters 12B has, but Quen active parameters 12B has, but Quen 3.5122B has active 10B A 10B and Quen 3.5122B has active 10B A 10B and Quen 3.5122B has active 10B A 10B and Quen 3.635B active 3B. So the 3 billion 3.635B active 3B. So the 3 billion 3.635B active 3B. So the 3 billion parameters is actually active in that parameters is actually active in that parameters is actually active in that Quen model. Finally, the price. Now, Quen model. Finally, the price. Now, Quen model. Finally, the price. Now, there's a footnote on this page at the there's a footnote on this page at the there's a footnote on this page at the very bottom here. They tell you about
-
very bottom here. They tell you about very bottom here. They tell you about the test that they performed. And they the test that they performed. And they the test that they performed. And they also mentioned that the retail price for also mentioned that the retail price for also mentioned that the retail price for Nvidia DJX Spark is $4,699. Nvidia DJX Spark is $4,699. Nvidia DJX Spark is $4,699. And for a few months, it was it came out And for a few months, it was it came out And for a few months, it was it came out at $4,000 and for a few months it was at $4,000 and for a few months it was at $4,000 and for a few months it was $4,700. Then it went back down to $4,000 $4,700. Then it went back down to $4,000 $4,700. Then it went back down to $4,000 and now it's back up to 4,500. Why does and now it's back up to 4,500. Why does and now it's back up to 4,500. Why does that matter? It doesn't matter in the that matter? It doesn't matter in the that matter? It doesn't matter in the grand scheme of things, but in some of grand scheme of things, but in some of grand scheme of things, but in some of their results, they're actually showing their results, they're actually showing their results, they're actually showing more tokens per dollar line. It's going more tokens per dollar line. It's going more tokens per dollar line. It's going to be kind of hard to pin down that to be kind of hard to pin down that to be kind of hard to pin down that number at this point. Sometimes it's number at this point. Sometimes it's number at this point. Sometimes it's true and sometimes it isn't. So, where true and sometimes it isn't. So, where true and sometimes it isn't. So, where does that leave group two? Well, in a does that leave group two? Well, in a does that leave group two? Well, in a really good spot. On raw speeds, it's a really good spot. On raw speeds, it's a really good spot. On raw speeds, it's a trade. The Spark takes prefill and trade. The Spark takes prefill and trade. The Spark takes prefill and compute while the Halo ties on token compute while the Halo ties on token compute while the Halo ties on token generation and brings the memory. So for generation and brings the memory. So for generation and brings the memory. So for things that you do all day like things that you do all day like things that you do all day like chatting, Olama, LM Studio, running chatting, Olama, LM Studio, running chatting, Olama, LM Studio, running agents in VS Code, this thing will be agents in VS Code, this thing will be agents in VS Code, this thing will be just great. And there's one more thing just great. And there's one more thing just great. And there's one more thing and that's the best thing and this is and that's the best thing and this is and that's the best thing and this is kind of an exciting time for AMD and kind of an exciting time for AMD and kind of an exciting time for AMD and these machines because every model and these machines because every model and these machines because every model and every tool that I threw at this box just every tool that I threw at this box just every tool that I threw at this box just worked without me fighting it. This was worked without me fighting it. This was worked without me fighting it. This was not the case last year. So just to not the case last year. So just to not the case last year. So just to validate the claims on the site, I'm validate the claims on the site, I'm validate the claims on the site, I'm running Quen 3.6 35B. Actually, the running Quen 3.6 35B. Actually, the running Quen 3.6 35B. Actually, the Spark is a little bit faster here, too.
-
Spark is a little bit faster here, too. Spark is a little bit faster here, too. And this pretty much matches exactly to And this pretty much matches exactly to And this pretty much matches exactly to the memory bandwidth. So yeah, but look the memory bandwidth. So yeah, but look the memory bandwidth. So yeah, but look at prompt processing here. 1,4 tokens at prompt processing here. 1,4 tokens at prompt processing here. 1,4 tokens per second on the Halo. 1,783 per second on the Halo. 1,783 per second on the Halo. 1,783 tokens per second on the Spark. Halo is tokens per second on the Spark. Halo is tokens per second on the Spark. Halo is more than half at this point. Now, it more than half at this point. Now, it more than half at this point. Now, it may seem like the Spark is better. Uh, may seem like the Spark is better. Uh, may seem like the Spark is better. Uh, but there is a couple things that this but there is a couple things that this but there is a couple things that this box has that the Spark just doesn't. box has that the Spark just doesn't. box has that the Spark just doesn't. First, the DJX Spark is an ARMbased First, the DJX Spark is an ARMbased First, the DJX Spark is an ARMbased machine, Nvidia Grace CPU, ARM. So, it machine, Nvidia Grace CPU, ARM. So, it machine, Nvidia Grace CPU, ARM. So, it runs Linux for ARM. This thing, it's runs Linux for ARM. This thing, it's runs Linux for ARM. This thing, it's boring old x86. And I mean that as the boring old x86. And I mean that as the boring old x86. And I mean that as the highest compliment because x86 means highest compliment because x86 means highest compliment because x86 means everything just works. My whole tool everything just works. My whole tool everything just works. My whole tool chain just runs with no caveats. I can chain just runs with no caveats. I can chain just runs with no caveats. I can install Visual Studio and just go with install Visual Studio and just go with install Visual Studio and just go with no emulation, no ARM version of no emulation, no ARM version of no emulation, no ARM version of anything. For a developer, that's anything. For a developer, that's anything. For a developer, that's massive because contrary to what you'll massive because contrary to what you'll massive because contrary to what you'll see all the loud people say in the see all the loud people say in the see all the loud people say in the comments, we want Linux. Actually, most comments, we want Linux. Actually, most comments, we want Linux. Actually, most of the world still runs on Windows, but of the world still runs on Windows, but of the world still runs on Windows, but I think with AI, things are slowly I think with AI, things are slowly I think with AI, things are slowly shifting towards Linux, and there's good shifting towards Linux, and there's good shifting towards Linux, and there's good reasons for that. There's one other reasons for that. There's one other reasons for that. There's one other thing to know about x86.
-
thing to know about x86. thing to know about x86. It runs a little warmer than ARM, even It runs a little warmer than ARM, even It runs a little warmer than ARM, even though the power usage at full tilt is though the power usage at full tilt is though the power usage at full tilt is actually about the same between these actually about the same between these actually about the same between these two boxes. And the Mac is just it's just two boxes. And the Mac is just it's just two boxes. And the Mac is just it's just the Mac. It sips the power. Now, I did the Mac. It sips the power. Now, I did the Mac. It sips the power. Now, I did run this thing continuously for a while, run this thing continuously for a while, run this thing continuously for a while, and this bottom metal part got pretty and this bottom metal part got pretty and this bottom metal part got pretty toasty. Not a dealbreaker, but another toasty. Not a dealbreaker, but another toasty. Not a dealbreaker, but another reason why you wouldn't want to stack reason why you wouldn't want to stack reason why you wouldn't want to stack these boxes. Also, x86, the price you these boxes. Also, x86, the price you these boxes. Also, x86, the price you pay for running everything under the pay for running everything under the pay for running everything under the sun. And that's why it also comes in two sun. And that's why it also comes in two sun. And that's why it also comes in two ways, a Windows version and a Linux ways, a Windows version and a Linux ways, a Windows version and a Linux version. You get to pick your own. The version. You get to pick your own. The version. You get to pick your own. The Spark is Linux only. Take it or leave Spark is Linux only. Take it or leave Spark is Linux only. Take it or leave it. By the way, if you buy the Windows it. By the way, if you buy the Windows it. By the way, if you buy the Windows version, you get the Windows key for the version, you get the Windows key for the version, you get the Windows key for the same price and then you can install same price and then you can install same price and then you can install Linux on it. One more thing, there is an Linux on it. One more thing, there is an Linux on it. One more thing, there is an NPU inside here up to 50 tops. The Spark NPU inside here up to 50 tops. The Spark NPU inside here up to 50 tops. The Spark doesn't have that and it's not either doesn't have that and it's not either doesn't have that and it's not either or. You can actually use the NPU or. You can actually use the NPU or. You can actually use the NPU alongside of the GPU and the CPU. Well, alongside of the GPU and the CPU. Well, alongside of the GPU and the CPU. Well, Alex, what do you do with the MPU? Who Alex, what do you do with the MPU? Who Alex, what do you do with the MPU? Who cares? The MPU is very efficient. Now, cares? The MPU is very efficient. Now, cares? The MPU is very efficient. Now, it's true that there's not too much it's true that there's not too much it's true that there's not too much software out there using the MPU right software out there using the MPU right software out there using the MPU right now, but things are being built like now, but things are being built like now, but things are being built like Lemonade Server, which actually comes Lemonade Server, which actually comes Lemonade Server, which actually comes [music] pre-installed here. And this [music] pre-installed here. And this [music] pre-installed here. And this allows you to run basically all the same allows you to run basically all the same allows you to run basically all the same old models, but using a hybrid approach.
-
old models, but using a hybrid approach. old models, but using a hybrid approach. You can use the NPU for the pre-filled You can use the NPU for the pre-filled You can use the NPU for the pre-filled part and the GPU for the decode part. part and the GPU for the decode part. part and the GPU for the decode part. There's a bunch of familiar models. They There's a bunch of familiar models. They There's a bunch of familiar models. They come in CPU versions, hybrid versions, come in CPU versions, hybrid versions, come in CPU versions, hybrid versions, and NPU only versions. Interesting. and NPU only versions. Interesting. and NPU only versions. Interesting. Deepseek NPU version only. Write a Deepseek NPU version only. Write a Deepseek NPU version only. Write a paragraph and go. There it goes. And paragraph and go. There it goes. And paragraph and go. There it goes. And it's fully on the NPU. It's not the it's fully on the NPU. It's not the it's fully on the NPU. It's not the fastest thing in the world, but it's fastest thing in the world, but it's fastest thing in the world, but it's only using 50 Ws of power. So, if we go only using 50 Ws of power. So, if we go only using 50 Ws of power. So, if we go back to my video about the Spark from back to my video about the Spark from back to my video about the Spark from last year, Stricks Halo, really good last year, Stricks Halo, really good last year, Stricks Halo, really good platform. They put out the chip. They platform. They put out the chip. They platform. They put out the chip. They said, "Here you go. The chip is here. said, "Here you go. The chip is here. said, "Here you go. The chip is here. They gave it to people, but software is They gave it to people, but software is They gave it to people, but software is still catching up." Yeah, that's what I still catching up." Yeah, that's what I still catching up." Yeah, that's what I said back then. But this box right here said back then. But this box right here said back then. But this box right here is actually AMD's answer to that is actually AMD's answer to that is actually AMD's answer to that criticism. Maybe maybe not this box, but criticism. Maybe maybe not this box, but criticism. Maybe maybe not this box, but what this box comes with. As soon as I what this box comes with. As soon as I what this box comes with. As soon as I unboxed it and turn it on, the software unboxed it and turn it on, the software unboxed it and turn it on, the software is already there and ready to go, is already there and ready to go, is already there and ready to go, pre-installed and preconfigured. There pre-installed and preconfigured. There pre-installed and preconfigured. There are no drivers to chase, no fighting are no drivers to chase, no fighting are no drivers to chase, no fighting Rockom for an afternoon, and no guessing Rockom for an afternoon, and no guessing Rockom for an afternoon, and no guessing which version of Python is installed. LM which version of Python is installed. LM which version of Python is installed. LM Studio, O Lama, Comfy UI are all ready Studio, O Lama, Comfy UI are all ready Studio, O Lama, Comfy UI are all ready to go. Lemonade server also with Rockom to go. Lemonade server also with Rockom to go. Lemonade server also with Rockom support from day zero. This is not your support from day zero. This is not your support from day zero. This is not your grandma Stricks Halo. Okay, this is this grandma Stricks Halo. Okay, this is this grandma Stricks Halo. Okay, this is this is the new Stricks Halo. And they didn't is the new Stricks Halo. And they didn't is the new Stricks Halo. And they didn't stop there. AMD got these playbooks on stop there. AMD got these playbooks on stop there. AMD got these playbooks on their site now. Actual step-by-step their site now. Actual step-by-step their site now. Actual step-by-step guides getting AI workloads running.
-
guides getting AI workloads running. guides getting AI workloads running. Look, they've got AMD Rising AI Halo. Look, they've got AMD Rising AI Halo. Look, they've got AMD Rising AI Halo. They've got AMD Ryzen AI APUs like They've got AMD Ryzen AI APUs like They've got AMD Ryzen AI APUs like stricts point APUs for example and then stricts point APUs for example and then stricts point APUs for example and then Radeon ones are coming soon as well. But Radeon ones are coming soon as well. But Radeon ones are coming soon as well. But under here under the Ryzen AI Halo box, under here under the Ryzen AI Halo box, under here under the Ryzen AI Halo box, this shows you step by step even this shows you step by step even this shows you step by step even filterable by beginner, intermediate, filterable by beginner, intermediate, filterable by beginner, intermediate, and advanced. Ooh, clustering two Ryzen and advanced. Ooh, clustering two Ryzen and advanced. Ooh, clustering two Ryzen AI halos. Interesting. Oh, um, hang on AI halos. Interesting. Oh, um, hang on AI halos. Interesting. Oh, um, hang on just a second. just a second. just a second. Yeah. Oh, Oh, yeah. Remember, there are three Oh, yeah. Remember, there are three groups of people. Group one developers, groups of people. Group one developers, groups of people. Group one developers, group two tinkers, and then there's group two tinkers, and then there's group two tinkers, and then there's group three. Don't forget about that group three. Don't forget about that group three. Don't forget about that because I couldn't answer them with just because I couldn't answer them with just because I couldn't answer them with just one machine. I needed to have two. That one machine. I needed to have two. That one machine. I needed to have two. That video is coming soon. Make sure you video is coming soon. Make sure you video is coming soon. Make sure you don't miss that. In the meantime, you don't miss that. In the meantime, you don't miss that. In the meantime, you can watch this video right here. Thanks can watch this video right here. Thanks can watch this video right here. Thanks [music] for watching and I'll see you in [music] for watching and I'll see you in [music] for watching and I'll see you in the next one.
Summary
The main topic is AMD's new mini-PC, a competitor to Nvidia's DJX Spark, featuring the Ryzen AI Max Plus 395 chip. The core message highlights the value proposition of the Merlin AI tool, which consolidates multiple AI models for efficient workflow management at a significant discount. The practical takeaway is to leverage Merlin AI for cost-effective and streamlined productivity, especially with the limited-time discount.