The Non-NVIDIA AI Card Everyone’s Ignoring
Read full transcript 11 segments
-
This is a GPU. Well, I'm going to keep This is a GPU. Well, I'm going to keep calling it a GPU, but I have to come calling it a GPU, but I have to come calling it a GPU, but I have to come clean right up front. It's not a GPU. It clean right up front. It's not a GPU. It clean right up front. It's not a GPU. It can't even draw a single pixel. There's can't even draw a single pixel. There's can't even draw a single pixel. There's not one video output anywhere on this not one video output anywhere on this not one video output anywhere on this thing. It does have these two ports on thing. It does have these two ports on thing. It does have these two ports on the back, though. More on that later. the back, though. More on that later. the back, though. More on that later. Anyway, this thing is not from Nvidia. Anyway, this thing is not from Nvidia. Anyway, this thing is not from Nvidia. There's no CUDA. There's no Olama There's no CUDA. There's no Olama There's no CUDA. There's no Olama singleclick installation. So, the singleclick installation. So, the singleclick installation. So, the question is, what can it actually do? question is, what can it actually do? question is, what can it actually do? And to answer that, I needed to do a And to answer that, I needed to do a And to answer that, I needed to do a head-to-head comparison that I never head-to-head comparison that I never head-to-head comparison that I never actually planned on doing. So, here's actually planned on doing. So, here's actually planned on doing. So, here's the thing. For years, if you wanted to the thing. For years, if you wanted to the thing. For years, if you wanted to run a serious model locally, the answer run a serious model locally, the answer run a serious model locally, the answer was always the same. Buy Nvidia. CUDA. was always the same. Buy Nvidia. CUDA. was always the same. Buy Nvidia. CUDA. CUDA owns this space. And if you watch CUDA owns this space. And if you watch CUDA owns this space. And if you watch this channel, you know I've built a this channel, you know I've built a this channel, you know I've built a bunch of those rigs. But the real bunch of those rigs. But the real bunch of those rigs. But the real question I care about, the one that I question I care about, the one that I question I care about, the one that I think you care about, is whether there's think you care about, is whether there's think you care about, is whether there's anything else that isn't Nvidia that can anything else that isn't Nvidia that can anything else that isn't Nvidia that can actually compete for local AI. And I'm actually compete for local AI. And I'm actually compete for local AI. And I'm going to give you the numbers on that, going to give you the numbers on that, going to give you the numbers on that, including the parts that aren't great including the parts that aren't great including the parts that aren't great yet. I don't expect one AI prompt to yet. I don't expect one AI prompt to yet. I don't expect one AI prompt to build an entire project for me. In build an entire project for me. In build an entire project for me. In reality, I'm constantly moving between reality, I'm constantly moving between reality, I'm constantly moving between models depending on the job. GPT for models depending on the job. GPT for models depending on the job. GPT for research, Claude for coding, Gemini for research, Claude for coding, Gemini for research, Claude for coding, Gemini for massive context, Nano Banana, massive context, Nano Banana, massive context, Nano Banana, Midjourney, Flux for images, and then Midjourney, Flux for images, and then Midjourney, Flux for images, and then I've got Seed Dance and Cling for video.
-
I've got Seed Dance and Cling for video. I've got Seed Dance and Cling for video. That's why Chat LM by Abacus AI makes That's why Chat LM by Abacus AI makes That's why Chat LM by Abacus AI makes sense. It brings one support for the sense. It brings one support for the sense. It brings one support for the latest GPT, Claw, Gemini, Gro, Deepseek, latest GPT, Claw, Gemini, Gro, Deepseek, latest GPT, Claw, Gemini, Gro, Deepseek, and more in one place the moment they and more in one place the moment they and more in one place the moment they drop. Pick any model from the interface drop. Pick any model from the interface drop. Pick any model from the interface or let route LLM automatically choose or let route LLM automatically choose or let route LLM automatically choose the best model for each prompt. Create the best model for each prompt. Create the best model for each prompt. Create professional presentations with graphs professional presentations with graphs professional presentations with graphs and charts and deep research detailed and charts and deep research detailed and charts and deep research detailed content. Need human sounding copy? Human content. Need human sounding copy? Human content. Need human sounding copy? Human eyes rewrites text to defeat AI eyes rewrites text to defeat AI eyes rewrites text to defeat AI detectors. Need visuals? Pick frontier detectors. Need visuals? Pick frontier detectors. Need visuals? Pick frontier or open- source models. And when you or open- source models. And when you or open- source models. And when you need more than chat, Abacus AI agent can need more than chat, Abacus AI agent can need more than chat, Abacus AI agent can help build complex apps and websites, help build complex apps and websites, help build complex apps and websites, connect payments, or run 24/7 agents connect payments, or run 24/7 agents connect payments, or run 24/7 agents that keep working through longer tasks. that keep working through longer tasks. that keep working through longer tasks. The best part is app hosting, back-end The best part is app hosting, back-end The best part is app hosting, back-end database, and author support comes with database, and author support comes with database, and author support comes with the subscription. All that starts at the subscription. All that starts at the subscription. All that starts at just $10 a month, way cheaper than just $10 a month, way cheaper than just $10 a month, way cheaper than paying for all those subscriptions paying for all those subscriptions paying for all those subscriptions separately. Check out chatlm.abacus.ai separately. Check out chatlm.abacus.ai separately. Check out chatlm.abacus.ai or click the link below. So, what is or click the link below. So, what is or click the link below. So, what is this thing? This is actually the Tense this thing? This is actually the Tense this thing? This is actually the Tense Torrent wormhole N300. There's two chips Torrent wormhole N300. There's two chips Torrent wormhole N300. There's two chips on one board, 24 gigs of memory, and on one board, 24 gigs of memory, and on one board, 24 gigs of memory, and it's about 1,400 bucks. Is it worth the it's about 1,400 bucks. Is it worth the it's about 1,400 bucks. Is it worth the moola, the dollars, the gold coins?
-
moola, the dollars, the gold coins? moola, the dollars, the gold coins? Well, that actually depends on how you Well, that actually depends on how you Well, that actually depends on how you use it, and I'll get into that. But use it, and I'll get into that. But use it, and I'll get into that. But first, I wanted to do some LLM tests. first, I wanted to do some LLM tests. first, I wanted to do some LLM tests. Now, quick confession. This card came up Now, quick confession. This card came up Now, quick confession. This card came up dead after months of sitting around dead after months of sitting around dead after months of sitting around being powered off without any updates. being powered off without any updates. being powered off without any updates. And it actually took a forced firmware And it actually took a forced firmware And it actually took a forced firmware reflash. Yeah, the kind that can break reflash. Yeah, the kind that can break reflash. Yeah, the kind that can break your card. Plus, uh, a fight with some your card. Plus, uh, a fight with some your card. Plus, uh, a fight with some version mismatches. And finally, I got version mismatches. And finally, I got version mismatches. And finally, I got it to show up. You know, kind of like it to show up. You know, kind of like it to show up. You know, kind of like when you leave Windows for a few months when you leave Windows for a few months when you leave Windows for a few months and then come back and turn it on. Yeah, and then come back and turn it on. Yeah, and then come back and turn it on. Yeah, like that. [music] So, is this thing like that. [music] So, is this thing like that. [music] So, is this thing plugandplay? Well, no, not yet. It's plugandplay? Well, no, not yet. It's plugandplay? Well, no, not yet. It's It's getting there. But once it is up, It's getting there. But once it is up, It's getting there. But once it is up, it's actually pretty good. And I got VLM it's actually pretty good. And I got VLM it's actually pretty good. And I got VLM server running Llama 3.18 billion. I server running Llama 3.18 billion. I server running Llama 3.18 billion. I know what you're thinking. Hold on. I know what you're thinking. Hold on. I know what you're thinking. Hold on. I ran some tests. Now, the very first ran some tests. Now, the very first ran some tests. Now, the very first request came back at 3.6 tokens a request came back at 3.6 tokens a request came back at 3.6 tokens a second. So, I was ready to cancel this second. So, I was ready to cancel this second. So, I was ready to cancel this video right then and there and be video right then and there and be video right then and there and be disappointed. But it turns out that it's disappointed. But it turns out that it's disappointed. But it turns out that it's just a one-time warm-up thing, not the just a one-time warm-up thing, not the just a one-time warm-up thing, not the real speed. So, once I got it warmed up real speed. So, once I got it warmed up real speed. So, once I got it warmed up and ran it a few times, 30 to 31 tokens and ran it a few times, 30 to 31 tokens and ran it a few times, 30 to 31 tokens a second, fully local on a card that was a second, fully local on a card that was a second, fully local on a card that was just a brick a little while ago. Let just a brick a little while ago. Let just a brick a little while ago. Let that be a lesson to you. Keep trying.
-
that be a lesson to you. Keep trying. that be a lesson to you. Keep trying. Well, sometimes. Anyway, 31 tokens per Well, sometimes. Anyway, 31 tokens per Well, sometimes. Anyway, 31 tokens per second is faster than I can read. Of second is faster than I can read. Of second is faster than I can read. Of course, uh, that argument doesn't hold course, uh, that argument doesn't hold course, uh, that argument doesn't hold much water anymore now that we've got much water anymore now that we've got much water anymore now that we've got agents in the mix. But what does hold agents in the mix. But what does hold agents in the mix. But what does hold water is concurrency. So, I threw eight water is concurrency. So, I threw eight water is concurrency. So, I threw eight requests at it at once. And it does 164 requests at it at once. And it does 164 requests at it at once. And it does 164 tokens a second total. That's five times tokens a second total. That's five times tokens a second total. That's five times the throughput. As a little inference the throughput. As a little inference the throughput. As a little inference server, it punches way above the single server, it punches way above the single server, it punches way above the single stream number, which is pretty cool. And stream number, which is pretty cool. And stream number, which is pretty cool. And it actually chews through prompts pretty it actually chews through prompts pretty it actually chews through prompts pretty fast. Uh, 3700 tokens per second, up to fast. Uh, 3700 tokens per second, up to fast. Uh, 3700 tokens per second, up to 4,000 tokens per second. And because 4,000 tokens per second. And because 4,000 tokens per second. And because you're all developers, let's give it a you're all developers, let's give it a you're all developers, let's give it a real job. Write a function to merge real job. Write a function to merge real job. Write a function to merge overlapping intervals. Looks good to me. overlapping intervals. Looks good to me. overlapping intervals. Looks good to me. Now, of course, the quality depends on Now, of course, the quality depends on Now, of course, the quality depends on the model that you're using, but it the model that you're using, but it the model that you're using, but it works. But most of my tests are going to works. But most of my tests are going to works. But most of my tests are going to be using Llama Beni. That's this project be using Llama Beni. That's this project be using Llama Beni. That's this project right here on GitHub, which now I'm a right here on GitHub, which now I'm a right here on GitHub, which now I'm a proud contributor to also. proud contributor to also. proud contributor to also. [clears throat] Good job, Alex. Good [clears throat] Good job, Alex. Good [clears throat] Good job, Alex. Good job. job. job. [music] Now, under the hood, this thing [music] Now, under the hood, this thing [music] Now, under the hood, this thing is a completely different animal. A GPU is a completely different animal. A GPU is a completely different animal. A GPU is a giant sea of tiny cores like in is a giant sea of tiny cores like in is a giant sea of tiny cores like in Nvidia's CUDA cores. They're all running Nvidia's CUDA cores. They're all running Nvidia's CUDA cores. They're all running in lock step. Here on the wormhole, it's in lock step. Here on the wormhole, it's in lock step. Here on the wormhole, it's a grid of bigger cores. They're not even a grid of bigger cores. They're not even a grid of bigger cores. They're not even called CUDA cores. They're 106 cores and called CUDA cores. They're 106 cores and called CUDA cores. They're 106 cores and they're wired by a network right on the they're wired by a network right on the they're wired by a network right on the chip. Inside every one of them are five chip. Inside every one of them are five chip. Inside every one of them are five little Risk 5, five risk 5 CPUs. They're little Risk 5, five risk 5 CPUs. They're little Risk 5, five risk 5 CPUs. They're CPUs. And yeah, risk 5, the open- source CPUs. And yeah, risk 5, the open- source CPUs. And yeah, risk 5, the open- source instruction set. But the CPUs are not instruction set. But the CPUs are not instruction set. But the CPUs are not the ones doing the math. They're like the ones doing the math. They're like the ones doing the math. They're like traffic cops. They're streaming data traffic cops. They're streaming data traffic cops. They're streaming data through a dedicated matrix engine with through a dedicated matrix engine with through a dedicated matrix engine with its own chip memory. Data flows through
-
its own chip memory. Data flows through its own chip memory. Data flows through the grid instead of hammering slow the grid instead of hammering slow the grid instead of hammering slow memory. Remember that because it's going memory. Remember that because it's going memory. Remember that because it's going to matter in a minute. And if you want to matter in a minute. And if you want to matter in a minute. And if you want to program at the metal, to program at the metal, to program at the metal, that's their open- source SDK called that's their open- source SDK called that's their open- source SDK called Metallium. Yeah, you can do that though Metallium. Yeah, you can do that though Metallium. Yeah, you can do that though almost nobody needs to because you can almost nobody needs to because you can almost nobody needs to because you can just use VLM which wraps everything for just use VLM which wraps everything for just use VLM which wraps everything for you nicely and that's what I used. But you nicely and that's what I used. But you nicely and that's what I used. But everything they have here is documented everything they have here is documented everything they have here is documented pretty nicely and open source. And this pretty nicely and open source. And this pretty nicely and open source. And this even explains the uh the tensic scores. even explains the uh the tensic scores. even explains the uh the tensic scores. Here's the TT Metal repo for your Here's the TT Metal repo for your Here's the TT Metal repo for your perusal. Hey, updated 54 minutes ago. perusal. Hey, updated 54 minutes ago. perusal. Hey, updated 54 minutes ago. They're on top of things. But here's the They're on top of things. But here's the They're on top of things. But here's the part that makes it a different bet part that makes it a different bet part that makes it a different bet altogether. Those ports on the back, altogether. Those ports on the back, altogether. Those ports on the back, that's not video out. It's networking. that's not video out. It's networking. that's not video out. It's networking. That's the same kind of ports uh that That's the same kind of ports uh that That's the same kind of ports uh that are on the back of the DGX Spark. Those are on the back of the DGX Spark. Those are on the back of the DGX Spark. Those are QSFP ports. And switches like this, are QSFP ports. And switches like this, are QSFP ports. And switches like this, the ones I use to connect a bunch of DJX the ones I use to connect a bunch of DJX the ones I use to connect a bunch of DJX Sparks together. See, Nvidia has Sparks together. See, Nvidia has Sparks together. See, Nvidia has stripped cardtocard connection out of stripped cardtocard connection out of stripped cardtocard connection out of every consumer GPU. Remember NVLink? every consumer GPU. Remember NVLink? every consumer GPU. Remember NVLink? Well, it died after the 3090. Your 4090, Well, it died after the 3090. Your 4090, Well, it died after the 3090. Your 4090, your 5090, they don't have that. Uh they your 5090, they don't have that. Uh they your 5090, they don't have that. Uh they scale over PCIe and as a result, they scale over PCIe and as a result, they scale over PCIe and as a result, they lose about 30% of performance that way.
-
lose about 30% of performance that way. lose about 30% of performance that way. I I haven't actually tested that. I need I I haven't actually tested that. I need I I haven't actually tested that. I need to get some 3090s in here. I'm sure to get some 3090s in here. I'm sure to get some 3090s in here. I'm sure somebody already tested it. Like this somebody already tested it. Like this somebody already tested it. Like this guy, Digital Spaceport. Nvidia saves the guy, Digital Spaceport. Nvidia saves the guy, Digital Spaceport. Nvidia saves the real fabric for $30,000 data center real fabric for $30,000 data center real fabric for $30,000 data center chips. Tenstor [music] chips. Tenstor [music] chips. Tenstor [music] put it on a $1,400 card. The idea is put it on a $1,400 card. The idea is put it on a $1,400 card. The idea is that you can connect many cards and that you can connect many cards and that you can connect many cards and they'll all work together as one, which they'll all work together as one, which they'll all work together as one, which is going to give you a way faster is going to give you a way faster is going to give you a way faster connection than going through the PCI connection than going through the PCI connection than going through the PCI bus. The catch? Well, with one card, you bus. The catch? Well, with one card, you bus. The catch? Well, with one card, you don't use it. It's a bet that you'll don't use it. It's a bet that you'll don't use it. It's a bet that you'll grow into it eventually once you start grow into it eventually once you start grow into it eventually once you start using their ecosystem. using their ecosystem. using their ecosystem. Now, let's be real about what this card Now, let's be real about what this card Now, let's be real about what this card is and isn't. Remember my 30 tokens per is and isn't. Remember my 30 tokens per is and isn't. Remember my 30 tokens per second. Well, Tentor's own official second. Well, Tentor's own official second. Well, Tentor's own official number for this is about 44. Here's number for this is about 44. Here's number for this is about 44. Here's their perf. MD, which does get updated their perf. MD, which does get updated their perf. MD, which does get updated pretty frequently. And there it is. M300 pretty frequently. And there it is. M300 pretty frequently. And there it is. M300 llama 3.18b 44. I even flipped my CPU llama 3.18b 44. I even flipped my CPU llama 3.18b 44. I even flipped my CPU into full performance mode. Remeasured into full performance mode. Remeasured into full performance mode. Remeasured everything. No difference. So, the only everything. No difference. So, the only everything. No difference. So, the only thing that could account for this is thing that could account for this is thing that could account for this is Tenstor software, not my machine. I Tenstor software, not my machine. I Tenstor software, not my machine. I thought they might be like uh on a later thought they might be like uh on a later thought they might be like uh on a later version than mine on some kind of version than mine on some kind of version than mine on some kind of special version that's internal or special version that's internal or special version that's internal or something. And then I noticed that their something. And then I noticed that their something. And then I noticed that their quantization is a little bit different quantization is a little bit different quantization is a little bit different here. I use a strictly BFP8 here. I use a strictly BFP8 here. I use a strictly BFP8 quantization. Wow, that's a popular quantization. Wow, that's a popular quantization. Wow, that's a popular model. And they used a BFP8 MLP, only model. And they used a BFP8 MLP, only model. And they used a BFP8 MLP, only the 302 decoder layer and then BFP4 MLP the 302 decoder layer and then BFP4 MLP the 302 decoder layer and then BFP4 MLP elsewhere because they used less bit elsewhere because they used less bit elsewhere because they used less bit precision here. That's probably what precision here. That's probably what precision here. That's probably what resulted in it being a little faster.
-
resulted in it being a little faster. resulted in it being a little faster. However, I'm standing by what I said. However, I'm standing by what I said. However, I'm standing by what I said. It's an optimization at the software It's an optimization at the software It's an optimization at the software level at that point. You might also level at that point. You might also level at that point. You might also notice other devices listed here like notice other devices listed here like notice other devices listed here like the N150 and the T3K and the TG. Tenstor the N150 and the T3K and the TG. Tenstor the N150 and the T3K and the TG. Tenstor also has a bunch of other devices that also has a bunch of other devices that also has a bunch of other devices that they have. By the way, this video is not they have. By the way, this video is not they have. By the way, this video is not sponsored by Tentorant. They didn't even sponsored by Tentorant. They didn't even sponsored by Tentorant. They didn't even send me this card. I got it from a send me this card. I got it from a send me this card. I got it from a friend. They have these uh black hole friend. They have these uh black hole friend. They have these uh black hole cards. Those are the uh QSFP cables that cards. Those are the uh QSFP cables that cards. Those are the uh QSFP cables that you would use to connect them. If you've you would use to connect them. If you've you would use to connect them. If you've seen my videos connecting the sparks, seen my videos connecting the sparks, seen my videos connecting the sparks, they kind of look like this chunky boys. they kind of look like this chunky boys. they kind of look like this chunky boys. But here's the wormhole N150. You can But here's the wormhole N150. You can But here's the wormhole N150. You can see that it's using pretty much the same see that it's using pretty much the same see that it's using pretty much the same board, but it only has one chip. And the board, but it only has one chip. And the board, but it only has one chip. And the N300 has both of those chips. As a N300 has both of those chips. As a N300 has both of those chips. As a result, the N150, which is actually not result, the N150, which is actually not result, the N150, which is actually not that much less. It's not half the price that much less. It's not half the price that much less. It's not half the price of the 300, but it only has 12 gigs of of the 300, but it only has 12 gigs of of the 300, but it only has 12 gigs of memory. So, it can't even run some memory. So, it can't even run some memory. So, it can't even run some models of the 300 can. The 150 also has models of the 300 can. The 150 also has models of the 300 can. The 150 also has only 228 gigabytes per second memory only 228 gigabytes per second memory only 228 gigabytes per second memory bandwidth, while the 300 has 576 bandwidth, while the 300 has 576 bandwidth, while the 300 has 576 gigabytes per second, which results in gigabytes per second, which results in gigabytes per second, which results in decode speed kind of like this across decode speed kind of like this across decode speed kind of like this across different models. So 24 gigs is a single different models. So 24 gigs is a single different models. So 24 gigs is a single card ceiling for the 300. And card ceiling for the 300. And card ceiling for the 300. And realistically, you're going to be in the realistically, you're going to be in the realistically, you're going to be in the 7 to 9 billion parameter range if you 7 to 9 billion parameter range if you 7 to 9 billion parameter range if you use it with enough context. And it's use it with enough context. And it's use it with enough context. And it's definitely behind the frontier. The definitely behind the frontier. The definitely behind the frontier. The newest models like Gemma 4, Quen 3.5, newest models like Gemma 4, Quen 3.5, newest models like Gemma 4, Quen 3.5, and 3.6, those newer models are nowhere and 3.6, those newer models are nowhere and 3.6, those newer models are nowhere to be found. They're not ported yet. The to be found. They're not ported yet. The to be found. They're not ported yet. The newest thing on this list is Quen 3. And newest thing on this list is Quen 3. And newest thing on this list is Quen 3. And I was able to run Quen 38 8B. Here are I was able to run Quen 38 8B. Here are I was able to run Quen 38 8B. Here are the numbers. Not the bleeding edge, but the numbers. Not the bleeding edge, but the numbers. Not the bleeding edge, but somewhat capable models, which brings us somewhat capable models, which brings us somewhat capable models, which brings us to the question that actually matters.
-
to the question that actually matters. to the question that actually matters. Is any of this enough to justify buying Is any of this enough to justify buying Is any of this enough to justify buying it over an actual GPU. it over an actual GPU. it over an actual GPU. [music] Now, before we look into whether [music] Now, before we look into whether [music] Now, before we look into whether it's worth buying one of these over a it's worth buying one of these over a it's worth buying one of these over a GPU, there's one thing that actually GPU, there's one thing that actually GPU, there's one thing that actually surprised me. The power. I measured the surprised me. The power. I measured the surprised me. The power. I measured the power while it was running and power while it was running and power while it was running and generating and it still only pulled generating and it still only pulled generating and it still only pulled about 70 watts on average. This was on a about 70 watts on average. This was on a about 70 watts on average. This was on a car that's rated for 300 watts. I think car that's rated for 300 watts. I think car that's rated for 300 watts. I think there's definitely some headroom left there's definitely some headroom left there's definitely some headroom left here and I pushed it over several users. here and I pushed it over several users. here and I pushed it over several users. So, it loses on speed, but maybe it wins So, it loses on speed, but maybe it wins So, it loses on speed, but maybe it wins on efficiency. I mean, this is pretty on efficiency. I mean, this is pretty on efficiency. I mean, this is pretty low power usage and efficiency being low power usage and efficiency being low power usage and efficiency being tokens per watt. So, I decided to test tokens per watt. So, I decided to test tokens per watt. So, I decided to test it against an actual GPU. Same box, same it against an actual GPU. Same box, same it against an actual GPU. Same box, same model. I pulled out the 10 storm card model. I pulled out the 10 storm card model. I pulled out the 10 storm card and dropped in an Nvidia RTX5080. and dropped in an Nvidia RTX5080. and dropped in an Nvidia RTX5080. They're in the same price range. See, They're in the same price range. See, They're in the same price range. See, $999 $999 $999 according to Nvidia. And then based on according to Nvidia. And then based on according to Nvidia. And then based on the OEM, you have MSI at $1,800, ASUS at the OEM, you have MSI at $1,800, ASUS at the OEM, you have MSI at $1,800, ASUS at $1,900. Well, here's one for 1400 from $1,900. Well, here's one for 1400 from $1,900. Well, here's one for 1400 from MSI. I ran the same 8-bit quantization MSI. I ran the same 8-bit quantization MSI. I ran the same 8-bit quantization FP8 on the GPU to match Ken Torrance FP8 on the GPU to match Ken Torrance FP8 on the GPU to match Ken Torrance BFB8. And the raw speed, um, BFB8. And the raw speed, um, BFB8. And the raw speed, um, not close. Nope. 93 tokens a second on not close. Nope. 93 tokens a second on not close. Nope. 93 tokens a second on the Nvidia versus 10 torren 30. Three the Nvidia versus 10 torren 30. Three the Nvidia versus 10 torren 30. Three times faster. That's decode. What about times faster. That's decode. What about times faster. That's decode. What about prefill? That's reading the prompts.
-
prefill? That's reading the prompts. prefill? That's reading the prompts. Yeah, the GPU pulls way ahead. But okay, Yeah, the GPU pulls way ahead. But okay, Yeah, the GPU pulls way ahead. But okay, uh the GPU uses way more power, right? uh the GPU uses way more power, right? uh the GPU uses way more power, right? So, what about performance per watt? So, what about performance per watt? So, what about performance per watt? This is supposed to be where the tentor This is supposed to be where the tentor This is supposed to be where the tentor fights back. fights back. fights back. Um Um Um yeah, no, the GPU wins that too. every yeah, no, the GPU wins that too. every yeah, no, the GPU wins that too. every single batch size I ran. Sure, it pulls single batch size I ran. Sure, it pulls single batch size I ran. Sure, it pulls more power, but it cranks out so many more power, but it cranks out so many more power, but it cranks out so many more tokens that the tokens per watt is more tokens that the tokens per watt is more tokens that the tokens per watt is still higher. And I'm being generous to still higher. And I'm being generous to still higher. And I'm being generous to the test store because it's 70 watts is the test store because it's 70 watts is the test store because it's 70 watts is just the chip, not the whole board. If just the chip, not the whole board. If just the chip, not the whole board. If you hook up networking, these uh QSFP you hook up networking, these uh QSFP you hook up networking, these uh QSFP ports, they take up a lot. So, we ports, they take up a lot. So, we ports, they take up a lot. So, we haven't done that. I haven't tested that haven't done that. I haven't tested that haven't done that. I haven't tested that cuz I only have one board. So, let me cuz I only have one board. So, let me cuz I only have one board. So, let me just say it straight. For running an 8 just say it straight. For running an 8 just say it straight. For running an 8 billion model today, the Nvidia GPU is billion model today, the Nvidia GPU is billion model today, the Nvidia GPU is faster and it's actually more efficient. faster and it's actually more efficient. faster and it's actually more efficient. So, case closed, right? Well, I wasn't So, case closed, right? Well, I wasn't So, case closed, right? Well, I wasn't done yet because I had one more card on done yet because I had one more card on done yet because I had one more card on the bench, the AMD R9700. the bench, the AMD R9700. the bench, the AMD R9700. Pretty new card, pretty new GPU, 32 gigs Pretty new card, pretty new GPU, 32 gigs Pretty new card, pretty new GPU, 32 gigs of memory, a lot of memory, so it can of memory, a lot of memory, so it can of memory, a lot of memory, so it can fit larger models. Dedicated FP8 fit larger models. Dedicated FP8 fit larger models. Dedicated FP8 hardware. On paper, it's more capable hardware. On paper, it's more capable hardware. On paper, it's more capable than the other two because it can run than the other two because it can run than the other two because it can run larger models. And the results came in.
-
larger models. And the results came in. larger models. And the results came in. The decode speed wasn't great. Llama 3.1 The decode speed wasn't great. Llama 3.1 The decode speed wasn't great. Llama 3.1 8 billion actually was right in the 8 billion actually was right in the 8 billion actually was right in the middle between those two. But for Quen middle between those two. But for Quen middle between those two. But for Quen 38B, very close to what the 10 Storing 38B, very close to what the 10 Storing 38B, very close to what the 10 Storing card gave us. However, if we take a look card gave us. However, if we take a look card gave us. However, if we take a look at efficiency, this is the worst of all at efficiency, this is the worst of all at efficiency, this is the worst of all three. Dead last. The GPU pulls nearly three. Dead last. The GPU pulls nearly three. Dead last. The GPU pulls nearly 300 W and the output is pretty modest. 300 W and the output is pretty modest. 300 W and the output is pretty modest. And that is the whole story of this And that is the whole story of this And that is the whole story of this video. The Tenstor, the AMD card, the video. The Tenstor, the AMD card, the video. The Tenstor, the AMD card, the hardware is not the problem. The AMD hardware is not the problem. The AMD hardware is not the problem. The AMD chip is seriously capable silicon. So is chip is seriously capable silicon. So is chip is seriously capable silicon. So is Tenstor. The RTX 5080 isn't winning Tenstor. The RTX 5080 isn't winning Tenstor. The RTX 5080 isn't winning because the chip is magical somehow. because the chip is magical somehow. because the chip is magical somehow. It's winning because it has 15 years of It's winning because it has 15 years of It's winning because it has 15 years of CUDA development behind it. Every one of CUDA development behind it. Every one of CUDA development behind it. Every one of these challengers is only as good as how these challengers is only as good as how these challengers is only as good as how far along the software stack is. The far along the software stack is. The far along the software stack is. The real question is not whose chip is real question is not whose chip is real question is not whose chip is faster. The real question is whose faster. The real question is whose faster. The real question is whose software is going to catch up to Nvidia. software is going to catch up to Nvidia. software is going to catch up to Nvidia. And that's what makes this whole space And that's what makes this whole space And that's what makes this whole space actually interesting and worth watching. actually interesting and worth watching. actually interesting and worth watching. So if I were you, if you just want fast So if I were you, if you just want fast So if I were you, if you just want fast oneclick local AI right now today, get a oneclick local AI right now today, get a oneclick local AI right now today, get a GPU or a Mac. And you can argue with me GPU or a Mac. And you can argue with me GPU or a Mac. And you can argue with me in the comments about Mac or GPU. They in the comments about Mac or GPU. They in the comments about Mac or GPU. They have different pros and cons. I've have different pros and cons. I've have different pros and cons. I've actually made videos about that. The 10 actually made videos about that. The 10 actually made videos about that. The 10 storing cards are for tinkerers, the storing cards are for tinkerers, the storing cards are for tinkerers, the folks who want real non- Nvidia box, who folks who want real non- Nvidia box, who folks who want real non- Nvidia box, who want to learn AI silicon hands-on, and want to learn AI silicon hands-on, and want to learn AI silicon hands-on, and who are betting that the software will who are betting that the software will who are betting that the software will catch up? Because 2 years from now, this catch up? Because 2 years from now, this catch up? Because 2 years from now, this exact comparison could look very exact comparison could look very exact comparison could look very different. So, which ecosystem do you different. So, which ecosystem do you different. So, which ecosystem do you think catches up first? Let me know in think catches up first? Let me know in think catches up first? Let me know in the comments down below. You might also the comments down below. You might also the comments down below. You might also enjoy this video next. Thanks for enjoy this video next. Thanks for enjoy this video next. Thanks for watching, and I'll see you next time.
-
watching, and I'll see you next time. watching, and I'll see you next time. [music]
Summary
The main theme discusses alternatives to Nvidia GPUs for local AI tasks, referencing specific AI models like GPT, Claude, and Gemini, and highlighting the potential of the Tense Torrent wormhole N300. The practical takeaway is that while Nvidia's CUDA has dominated, other hardware and platforms like Chat LM by Abacus AI offer a more flexible and cost-effective solution for running diverse AI models.