This was a data center a year ago… Now it's on my desk
Read full transcript 14 segments
-
It was just a year ago that this was a It was just a year ago that this was a data center. And today it's next to my data center. And today it's next to my data center. And today it's next to my laptop on my desk. And everything you're laptop on my desk. And everything you're laptop on my desk. And everything you're about to see, dozens of AIs all thinking about to see, dozens of AIs all thinking about to see, dozens of AIs all thinking at once, is running on this one machine, at once, is running on this one machine, at once, is running on this one machine, not a data center. We've been waiting a not a data center. We've been waiting a not a data center. We've been waiting a long time for the DJX station. And now long time for the DJX station. And now long time for the DJX station. And now it's finally here. But the interesting it's finally here. But the interesting it's finally here. But the interesting part is not how fast it is, and it is part is not how fast it is, and it is part is not how fast it is, and it is plenty fast, and I'm going to test that plenty fast, and I'm going to test that plenty fast, and I'm going to test that out for you. But first, we're going to out for you. But first, we're going to out for you. But first, we're going to see what is this thing, and what is it see what is this thing, and what is it see what is this thing, and what is it made out of? Besides a lot of copper, made out of? Besides a lot of copper, made out of? Besides a lot of copper, this is the Asus Expert Center Pro this is the Asus Expert Center Pro this is the Asus Expert Center Pro ET900NG ET900NG ET900NG G3. That's a long name, but inside it's G3. That's a long name, but inside it's G3. That's a long name, but inside it's known as the DGX station. That's what known as the DGX station. That's what known as the DGX station. That's what powers it. That's the Nvidia bit, and powers it. That's the Nvidia bit, and powers it. That's the Nvidia bit, and the GB300 is the chip. We first heard the GB300 is the chip. We first heard the GB300 is the chip. We first heard about it alongside the DGX Spark on about it alongside the DGX Spark on about it alongside the DGX Spark on stage when Jensen Juan announced it. stage when Jensen Juan announced it. stage when Jensen Juan announced it. It'll be a beast for a lot of industries It'll be a beast for a lot of industries It'll be a beast for a lot of industries that need heavy compute, but today I'm that need heavy compute, but today I'm that need heavy compute, but today I'm looking at it as an AI developer. So, looking at it as an AI developer. So, looking at it as an AI developer. So, it's all LLMs and agents, and I'm going it's all LLMs and agents, and I'm going it's all LLMs and agents, and I'm going big. GB300 is that chip inside. And big. GB300 is that chip inside. And big. GB300 is that chip inside. And GB300s have been in service like the HGX GB300s have been in service like the HGX GB300s have been in service like the HGX for months now, since last year. I for months now, since last year. I for months now, since last year. I tested one earlier this year on the tested one earlier this year on the tested one earlier this year on the channel, actually. But this is the channel, actually. But this is the channel, actually. But this is the desktop variant, the desktop super chip desktop variant, the desktop super chip desktop variant, the desktop super chip as it's known, built for your desk.
-
as it's known, built for your desk. as it's known, built for your desk. Tada. Now, yes, it has that crazy chip Tada. Now, yes, it has that crazy chip Tada. Now, yes, it has that crazy chip inside, but it also has 748 GB of memory inside, but it also has 748 GB of memory inside, but it also has 748 GB of memory to run your AI. Its memory kind of works to run your AI. Its memory kind of works to run your AI. Its memory kind of works like Apple Silicon's unified memory, but like Apple Silicon's unified memory, but like Apple Silicon's unified memory, but instead of one pool, it fuses two. An instead of one pool, it fuses two. An instead of one pool, it fuses two. An ultra fast 256 gig GPU pool and a big ultra fast 256 gig GPU pool and a big ultra fast 256 gig GPU pool and a big 500 plus gig grace pool into one 748 gig 500 plus gig grace pool into one 748 gig 500 plus gig grace pool into one 748 gig pool. I'm saying pool a lot. I want to pool. I'm saying pool a lot. I want to pool. I'm saying pool a lot. I want to go swimming. It's kind of hot outside. go swimming. It's kind of hot outside. go swimming. It's kind of hot outside. I'm also hungry. I'm also hungry. I'm also hungry. Think of it as unified memory with a Think of it as unified memory with a Think of it as unified memory with a fast lane and a not so fast lane, a fast lane and a not so fast lane, a fast lane and a not so fast lane, a slightly slower lane. I'll get back to slightly slower lane. I'll get back to slightly slower lane. I'll get back to that. There's also one neat thing that that. There's also one neat thing that that. There's also one neat thing that you don't normally find on machines that you don't normally find on machines that you don't normally find on machines that are on your desk are 400 gigabit nyx are on your desk are 400 gigabit nyx are on your desk are 400 gigabit nyx network cards. These are called super network cards. These are called super network cards. These are called super nicks and it's using connectx8. These nicks and it's using connectx8. These nicks and it's using connectx8. These are basically things that you'd find in are basically things that you'd find in are basically things that you'd find in a data center connecting multiple GPUs a data center connecting multiple GPUs a data center connecting multiple GPUs together and it's right here. If you see together and it's right here. If you see together and it's right here. If you see my videos connecting and clustering DGX my videos connecting and clustering DGX my videos connecting and clustering DGX sparks, those have connect 7, but it's sparks, those have connect 7, but it's sparks, those have connect 7, but it's basically the same idea. You'll be able basically the same idea. You'll be able basically the same idea. You'll be able to connect multiple boxes with ultra low to connect multiple boxes with ultra low to connect multiple boxes with ultra low latency and high bandwidth. There's also latency and high bandwidth. There's also latency and high bandwidth. There's also three PCIe slots on there, so you can three PCIe slots on there, so you can three PCIe slots on there, so you can expand it with more GPUs if you want to.
-
expand it with more GPUs if you want to. expand it with more GPUs if you want to. This Loner unit actually shipped with an This Loner unit actually shipped with an This Loner unit actually shipped with an RTX Pro 4000 inside, but you can also RTX Pro 4000 inside, but you can also RTX Pro 4000 inside, but you can also put an RTX Pro 6000 in there. Now, the put an RTX Pro 6000 in there. Now, the put an RTX Pro 6000 in there. Now, the Super Chip alone is about 360 watts idle Super Chip alone is about 360 watts idle Super Chip alone is about 360 watts idle and about 1,400 flat out when it's and about 1,400 flat out when it's and about 1,400 flat out when it's running. That's just a chip. The whole running. That's just a chip. The whole running. That's just a chip. The whole machine at the wall is more. And while machine at the wall is more. And while machine at the wall is more. And while you can plug it into a regular outlet in you can plug it into a regular outlet in you can plug it into a regular outlet in United States, which gives you 115 volts United States, which gives you 115 volts United States, which gives you 115 volts and 15 amps, I opted for a 240 volt with and 15 amps, I opted for a 240 volt with and 15 amps, I opted for a 240 volt with 20 amp outlet. That way I'm not running 20 amp outlet. That way I'm not running 20 amp outlet. That way I'm not running into any issues. All right, let's put it into any issues. All right, let's put it into any issues. All right, let's put it to work and see what it can do. to work and see what it can do. to work and see what it can do. We're running some big models today and We're running some big models today and We're running some big models today and here's a quick look at exactly what here's a quick look at exactly what here's a quick look at exactly what we're running >> GLM. >> GLM. >> Now, these models come in different >> Now, these models come in different >> Now, these models come in different quantizations and Blackwell architecture quantizations and Blackwell architecture quantizations and Blackwell architecture can do NVFP4 quantization, which means can do NVFP4 quantization, which means can do NVFP4 quantization, which means floating point 4, but it's the Nvidia floating point 4, but it's the Nvidia floating point 4, but it's the Nvidia format, and it runs really well and format, and it runs really well and format, and it runs really well and efficiently on this machine. Here's an efficiently on this machine. Here's an efficiently on this machine. Here's an example of the NVFP4. And this is example of the NVFP4. And this is example of the NVFP4. And this is Neatron 3 Super 120 billion parameter Neatron 3 Super 120 billion parameter Neatron 3 Super 120 billion parameter model. It's pretty big. I'm also going model. It's pretty big. I'm also going model. It's pretty big. I'm also going to be doing DeepSk V4 flash quen 3 235 to be doing DeepSk V4 flash quen 3 235 to be doing DeepSk V4 flash quen 3 235 billion parameter model. Also NVFP4. And billion parameter model. Also NVFP4. And billion parameter model. Also NVFP4. And we're going to do this one GLM 5.2. I we're going to do this one GLM 5.2. I we're going to do this one GLM 5.2. I know you're all excited about that one.
-
know you're all excited about that one. know you're all excited about that one. That's the big one. As a side note, most That's the big one. As a side note, most That's the big one. As a side note, most models come in 16 bit and quantization models come in 16 bit and quantization models come in 16 bit and quantization reduces the number of bits. So if it's reduces the number of bits. So if it's reduces the number of bits. So if it's 4bit, then you basically shrunk your 4bit, then you basically shrunk your 4bit, then you basically shrunk your model four times. So GLM 5.2 2 becomes model four times. So GLM 5.2 2 becomes model four times. So GLM 5.2 2 becomes 465 GB on disk instead of being like 2 465 GB on disk instead of being like 2 465 GB on disk instead of being like 2 TB. Not only are they smaller, but they TB. Not only are they smaller, but they TB. Not only are they smaller, but they also are faster. And NVFP4 is Nvidia's also are faster. And NVFP4 is Nvidia's also are faster. And NVFP4 is Nvidia's own 4-bit format that Blackwell chips own 4-bit format that Blackwell chips own 4-bit format that Blackwell chips like this one and the ones on the Pro like this one and the ones on the Pro like this one and the ones on the Pro GPUs, they can run it natively. That's GPUs, they can run it natively. That's GPUs, they can run it natively. That's the whole trick that puts a data center the whole trick that puts a data center the whole trick that puts a data center model on your desk. Now, if you notice, model on your desk. Now, if you notice, model on your desk. Now, if you notice, GLM 5.2, even the NVFP4 format is 465 GLM 5.2, even the NVFP4 format is 465 GLM 5.2, even the NVFP4 format is 465 GB. the DGX station high bandwidth GB. the DGX station high bandwidth GB. the DGX station high bandwidth memory HBM, it's not going to fit the memory HBM, it's not going to fit the memory HBM, it's not going to fit the entire model because HBM is only limited entire model because HBM is only limited entire model because HBM is only limited to 256 GB. That part is going to be to 256 GB. That part is going to be to 256 GB. That part is going to be really fast. But if the whole model is really fast. But if the whole model is really fast. But if the whole model is not an HBM, it's going to have to spill not an HBM, it's going to have to spill not an HBM, it's going to have to spill over to the rest of the memory. That's over to the rest of the memory. That's over to the rest of the memory. That's the system memory, the LPDDR5X. So the system memory, the LPDDR5X. So the system memory, the LPDDR5X. So together, they're considered coherent together, they're considered coherent together, they're considered coherent memory. But what the heck is coherent memory. But what the heck is coherent memory. But what the heck is coherent memory? Is it like memory that's aware memory? Is it like memory that's aware memory? Is it like memory that's aware of what's going on? Is it self-aware of what's going on? Is it self-aware of what's going on? Is it self-aware memory? All right, that was that was memory? All right, that was that was memory? All right, that was that was pretty dumb. Here's what's really pretty dumb. Here's what's really pretty dumb. Here's what's really happening. Look at how fast the HPM is.
-
happening. Look at how fast the HPM is. happening. Look at how fast the HPM is. 7 to 8 terabytes per second. Well, if 7 to 8 terabytes per second. Well, if 7 to 8 terabytes per second. Well, if the model spills onto the slower memory, the model spills onto the slower memory, the model spills onto the slower memory, which is 500 GB per second, what which is 500 GB per second, what which is 500 GB per second, what happens? Does the whole thing crawl? happens? Does the whole thing crawl? happens? Does the whole thing crawl? Well, no. Only the bytes that come from Well, no. Only the bytes that come from Well, no. Only the bytes that come from Grace. But GLM is a mixture of experts Grace. But GLM is a mixture of experts Grace. But GLM is a mixture of experts model. And the offloaded experts are model. And the offloaded experts are model. And the offloaded experts are exactly what each token reads. So, you exactly what each token reads. So, you exactly what each token reads. So, you feel most of the slow lane. That's 24 feel most of the slow lane. That's 24 feel most of the slow lane. That's 24 tokens per second. It's kind of like the tokens per second. It's kind of like the tokens per second. It's kind of like the same trick as running a giant model on a same trick as running a giant model on a same trick as running a giant model on a maxed out Mac Studio, which are Oh, I I maxed out Mac Studio, which are Oh, I I maxed out Mac Studio, which are Oh, I I moved. They're not here anymore. Bigger moved. They're not here anymore. Bigger moved. They're not here anymore. Bigger than the GPU's memory, but runs anyway than the GPU's memory, but runs anyway than the GPU's memory, but runs anyway because it's shared. Plus, this has an because it's shared. Plus, this has an because it's shared. Plus, this has an HBM fast year that Macs don't have. So, HBM fast year that Macs don't have. So, HBM fast year that Macs don't have. So, what's up with all these people buying what's up with all these people buying what's up with all these people buying Mac minis for hosting their agents? Is Mac minis for hosting their agents? Is Mac minis for hosting their agents? Is that even doable, or do I need this box? that even doable, or do I need this box? that even doable, or do I need this box? Well, a Mac Mini will run an agent. Well, a Mac Mini will run an agent. Well, a Mac Mini will run an agent. It'll just be on a much smaller scale. It'll just be on a much smaller scale. It'll just be on a much smaller scale. What the big box buys you is hard things What the big box buys you is hard things What the big box buys you is hard things and big models, and it does it fast. But and big models, and it does it fast. But and big models, and it does it fast. But we're going to talk about agents. Just we're going to talk about agents. Just we're going to talk about agents. Just hold that thought. I don't expect one AI hold that thought. I don't expect one AI hold that thought. I don't expect one AI prompt to build an entire project for prompt to build an entire project for prompt to build an entire project for me. In reality, I'm constantly moving me. In reality, I'm constantly moving me. In reality, I'm constantly moving between models depending on the job. GPT between models depending on the job. GPT between models depending on the job. GPT for research, Claude for coding, Gemini for research, Claude for coding, Gemini for research, Claude for coding, Gemini for massive context, nano banana, for massive context, nano banana, for massive context, nano banana, midjourney, flux for images, and then midjourney, flux for images, and then midjourney, flux for images, and then I've got seed dance and cling for video.
-
I've got seed dance and cling for video. I've got seed dance and cling for video. That's why chat lm by abacusai makes That's why chat lm by abacusai makes That's why chat lm by abacusai makes sense. It brings day one support for the sense. It brings day one support for the sense. It brings day one support for the latest GPT, Claw, Gemini, Grock, latest GPT, Claw, Gemini, Grock, latest GPT, Claw, Gemini, Grock, Deepseek, and more in one place the Deepseek, and more in one place the Deepseek, and more in one place the moment they drop. Pick any model from moment they drop. Pick any model from moment they drop. Pick any model from the interface or let Route LLM the interface or let Route LLM the interface or let Route LLM automatically choose the best model for automatically choose the best model for automatically choose the best model for each prompt. Create professional each prompt. Create professional each prompt. Create professional presentations with graphs and charts and presentations with graphs and charts and presentations with graphs and charts and deep research, detailed content. Need deep research, detailed content. Need deep research, detailed content. Need human sounding copy? Humanize rewrites human sounding copy? Humanize rewrites human sounding copy? Humanize rewrites text to defeat AI detectors. Need text to defeat AI detectors. Need text to defeat AI detectors. Need visuals? Pick Frontier or open- source visuals? Pick Frontier or open- source visuals? Pick Frontier or open- source models. And when you need more than models. And when you need more than models. And when you need more than chat, Abacus AI agent can help build chat, Abacus AI agent can help build chat, Abacus AI agent can help build complex apps and websites, connect complex apps and websites, connect complex apps and websites, connect payments, or run 247 agents that keep payments, or run 247 agents that keep payments, or run 247 agents that keep working through longer tasks. The best working through longer tasks. The best working through longer tasks. The best part is app hosting, back-end database, part is app hosting, back-end database, part is app hosting, back-end database, and off support comes with the and off support comes with the and off support comes with the subscription. All that starts at just subscription. All that starts at just subscription. All that starts at just $10 a month. Way cheaper than paying for $10 a month. Way cheaper than paying for $10 a month. Way cheaper than paying for all those subscriptions separately. all those subscriptions separately. all those subscriptions separately. Check out chat llm.abacus.ai Check out chat llm.abacus.ai Check out chat llm.abacus.ai or click the link below. Now, if you're or click the link below. Now, if you're or click the link below. Now, if you're using this machine running models and using this machine running models and using this machine running models and you're plugging in your code editor, for you're plugging in your code editor, for you're plugging in your code editor, for example, this is going to be something example, this is going to be something example, this is going to be something else. Boom. Okay, there it is. Neatron else. Boom. Okay, there it is. Neatron else. Boom. Okay, there it is. Neatron 120B going at 185 tokens per second 120B going at 185 tokens per second 120B going at 185 tokens per second right there in decode. Wo, we're up to right there in decode. Wo, we're up to right there in decode. Wo, we're up to 230 250. Oh, wow. It just keeps going 230 250. Oh, wow. It just keeps going 230 250. Oh, wow. It just keeps going up. What's going on here? Oh, it's doing up. What's going on here? Oh, it's doing up. What's going on here? Oh, it's doing a concurrency sweep. We didn't get to a concurrency sweep. We didn't get to a concurrency sweep. We didn't get to that part yet. Run single DGX. All that part yet. Run single DGX. All that part yet. Run single DGX. All right. Sorry about that. There we go.
-
right. Sorry about that. There we go. right. Sorry about that. There we go. about 184 tokens per second decode and about 184 tokens per second decode and about 184 tokens per second decode and we're getting some nice prefill there. we're getting some nice prefill there. we're getting some nice prefill there. Prefill is important for agents when Prefill is important for agents when Prefill is important for agents when you're sending in your prompt to process you're sending in your prompt to process you're sending in your prompt to process because your prompt is not just going to because your prompt is not just going to because your prompt is not just going to be hello, how are you or write a story be hello, how are you or write a story be hello, how are you or write a story or hi like I sometimes am guilty of or hi like I sometimes am guilty of or hi like I sometimes am guilty of showing off on this channel, but that's showing off on this channel, but that's showing off on this channel, but that's just chat. Your prompt is going to just chat. Your prompt is going to just chat. Your prompt is going to include all that code and all the files include all that code and all the files include all that code and all the files that the agent is automatically going to that the agent is automatically going to that the agent is automatically going to send in along with all the context of send in along with all the context of send in along with all the context of your conversation thus far and all that your conversation thus far and all that your conversation thus far and all that needs to be processed. For that you need needs to be processed. For that you need needs to be processed. For that you need fast prompt processing also known as fast prompt processing also known as fast prompt processing also known as prefill throughput. Here is the prefill prefill throughput. Here is the prefill prefill throughput. Here is the prefill for Neatron single stream. So like for Neatron single stream. So like for Neatron single stream. So like chatting basically 2200 on Neatron, 2500 chatting basically 2200 on Neatron, 2500 chatting basically 2200 on Neatron, 2500 on Quen, 1,800 on DeepSeek and GLM is a on Quen, 1,800 on DeepSeek and GLM is a on Quen, 1,800 on DeepSeek and GLM is a big chunker. So 243 also that's big chunker. So 243 also that's big chunker. So 243 also that's offloaded to that slower memory that I offloaded to that slower memory that I offloaded to that slower memory that I mentioned. That's why we're getting mentioned. That's why we're getting mentioned. That's why we're getting slower numbers here. But what would slower numbers here. But what would slower numbers here. But what would happen if we were not to use a single happen if we were not to use a single happen if we were not to use a single stream? If we were to use multiple stream? If we were to use multiple stream? If we were to use multiple streams, streams, streams, everything so far was just me, one user, everything so far was just me, one user, everything so far was just me, one user, one chat. But a machine like this isn't one chat. But a machine like this isn't one chat. But a machine like this isn't built for just chatting with with one built for just chatting with with one built for just chatting with with one person. It's built to serve a crowd. All person. It's built to serve a crowd. All person. It's built to serve a crowd. All right, they may be pushing it. Maybe right, they may be pushing it. Maybe right, they may be pushing it. Maybe like a few people in an office like a few people in an office like a few people in an office environment, maybe a few agents.
-
environment, maybe a few agents. environment, maybe a few agents. In fact, maybe a lot of agents. That's In fact, maybe a lot of agents. That's In fact, maybe a lot of agents. That's coming up soon. So, let's talk about coming up soon. So, let's talk about coming up soon. So, let's talk about concurrency. And that's something I concurrency. And that's something I concurrency. And that's something I showed before on the channel many times. showed before on the channel many times. showed before on the channel many times. And people were like, why do why do I And people were like, why do why do I And people were like, why do why do I need 4,000 tokens a second? Hold that need 4,000 tokens a second? Hold that need 4,000 tokens a second? Hold that thought. Concurrency just means how many thought. Concurrency just means how many thought. Concurrency just means how many requests hit the model at the same time. requests hit the model at the same time. requests hit the model at the same time. A whole team using it at once, an app A whole team using it at once, an app A whole team using it at once, an app with a 100 users, or as we'll see, a with a 100 users, or as we'll see, a with a 100 users, or as we'll see, a swarm of agents. And here's the swarm of agents. And here's the swarm of agents. And here's the interesting question. When I pile all interesting question. When I pile all interesting question. When I pile all those requests on at the same time, what those requests on at the same time, what those requests on at the same time, what happens? Does each person slow to a happens? Does each person slow to a happens? Does each person slow to a crawl? Does the machine choke or does crawl? Does the machine choke or does crawl? Does the machine choke or does something else happen? Well, let's see. something else happen? Well, let's see. something else happen? Well, let's see. I'm going to kick this off. And right I'm going to kick this off. And right I'm going to kick this off. And right now, I'm going to start with just me. now, I'm going to start with just me. now, I'm going to start with just me. About 180 tokens a second. We already About 180 tokens a second. We already About 180 tokens a second. We already saw this. Now watch as I add more at the saw this. Now watch as I add more at the saw this. Now watch as I add more at the same time. This is two now. 250 tokens same time. This is two now. 250 tokens same time. This is two now. 250 tokens per second. Four. We're getting up to per second. Four. We're getting up to per second. Four. We're getting up to 300 tokens per second. Now 350. Four. 300 tokens per second. Now 350. Four. 300 tokens per second. Now 350. Four. Concurrency of four. Now we're at 400 Concurrency of four. Now we're at 400 Concurrency of four. Now we're at 400 tokens per second. At concurrency of 8. tokens per second. At concurrency of 8. tokens per second. At concurrency of 8. And we can keep going. This machine will And we can keep going. This machine will And we can keep going. This machine will handle it. Here's concurrency of 16. It handle it. Here's concurrency of 16. It handle it. Here's concurrency of 16. It just keeps going. 700 tokens per second.
-
just keeps going. 700 tokens per second. just keeps going. 700 tokens per second. Now 750. We're over 800 tokens per Now 750. We're over 800 tokens per Now 750. We're over 800 tokens per second with concurrency of 16. And now second with concurrency of 16. And now second with concurrency of 16. And now concurrency 32. Ladies and gentlemen, we concurrency 32. Ladies and gentlemen, we concurrency 32. Ladies and gentlemen, we just hit 1,000 tokens per second. And just hit 1,000 tokens per second. And just hit 1,000 tokens per second. And yeah, we have more. We have more. 64 yeah, we have more. We have more. 64 yeah, we have more. We have more. 64 concurrency. That's 64 users or 64 concurrency. That's 64 users or 64 concurrency. That's 64 users or 64 agents. We're at 1,500 tokens per agents. We're at 1,500 tokens per agents. We're at 1,500 tokens per second. All that being generated right second. All that being generated right second. All that being generated right here on the fly. 1,700 1,800. We're at here on the fly. 1,700 1,800. We're at here on the fly. 1,700 1,800. We're at 128 concurrency. That's 128 requests. 128 concurrency. That's 128 requests. 128 concurrency. That's 128 requests. And it just keeps going. This is just And it just keeps going. This is just And it just keeps going. This is just insane. We're on Neatron 12B, one of the insane. We're on Neatron 12B, one of the insane. We're on Neatron 12B, one of the really fast models that runs on this. really fast models that runs on this. really fast models that runs on this. Really tuned well with NVFP4 and we've Really tuned well with NVFP4 and we've Really tuned well with NVFP4 and we've hit 2600 tokens per second. Now, it's hit 2600 tokens per second. Now, it's hit 2600 tokens per second. Now, it's not perfect. Each individual request not perfect. Each individual request not perfect. Each individual request takes a little bit longer. By the way, takes a little bit longer. By the way, takes a little bit longer. By the way, during that little experiment, we've during that little experiment, we've during that little experiment, we've generated 254,000 generated 254,000 generated 254,000 tokens in just about a minute or so. And tokens in just about a minute or so. And tokens in just about a minute or so. And the best peak was 4,96 tokens per the best peak was 4,96 tokens per the best peak was 4,96 tokens per second. That's continuous batching. The second. That's continuous batching. The second. That's continuous batching. The GPU packs all those requests together GPU packs all those requests together GPU packs all those requests together and runs them all in one pass. I did the and runs them all in one pass. I did the and runs them all in one pass. I did the other models, too. There's Neatron 12B other models, too. There's Neatron 12B other models, too. There's Neatron 12B that you just saw live. Here's Quen. We that you just saw live. Here's Quen. We that you just saw live. Here's Quen. We went to about the same, actually, 5,000, went to about the same, actually, 5,000, went to about the same, actually, 5,000, just over 5,000 tokens per second for just over 5,000 tokens per second for just over 5,000 tokens per second for 128 users. And this is how concurrency 128 users. And this is how concurrency 128 users. And this is how concurrency scales. Notice there is a little bit of scales. Notice there is a little bit of scales. Notice there is a little bit of a dip on both Neatron and Quen at 32, a dip on both Neatron and Quen at 32, a dip on both Neatron and Quen at 32, which is kind of weird because I would which is kind of weird because I would which is kind of weird because I would think that would be a little blip, an think that would be a little blip, an think that would be a little blip, an error, but both models did it several
-
error, but both models did it several error, but both models did it several times in a row, so that must be real. times in a row, so that must be real. times in a row, so that must be real. Deepseek does pretty well with that, Deepseek does pretty well with that, Deepseek does pretty well with that, too. 64 seems to be the sweet spot for too. 64 seems to be the sweet spot for too. 64 seems to be the sweet spot for DeepSseek on this machine. And GLM, DeepSseek on this machine. And GLM, DeepSseek on this machine. And GLM, remember GLM is offloading to the slower remember GLM is offloading to the slower remember GLM is offloading to the slower memory, but still it can all fit 35 memory, but still it can all fit 35 memory, but still it can all fit 35 tokens per second at concurrency of four tokens per second at concurrency of four tokens per second at concurrency of four and 56 at concurrency of 16. I didn't do and 56 at concurrency of 16. I didn't do and 56 at concurrency of 16. I didn't do eight. And reading the prompts does the eight. And reading the prompts does the eight. And reading the prompts does the exact same thing. Climbing to about exact same thing. Climbing to about exact same thing. Climbing to about 41,000 tokens per second for Quen 235, 41,000 tokens per second for Quen 235, 41,000 tokens per second for Quen 235, about 35,000 tokens per second on about 35,000 tokens per second on about 35,000 tokens per second on Neatron. And we do see a dive in Neatron. And we do see a dive in Neatron. And we do see a dive in DeepSeek for prompt processing. 64 is DeepSeek for prompt processing. 64 is DeepSeek for prompt processing. 64 is really nice for DeepSeek. So this tells really nice for DeepSeek. So this tells really nice for DeepSeek. So this tells you that it depends on what kind of you that it depends on what kind of you that it depends on what kind of model you're using too. That matters. model you're using too. That matters. model you're using too. That matters. Even GLM is happy. 1,800 tokens per Even GLM is happy. 1,800 tokens per Even GLM is happy. 1,800 tokens per second at concurrency of 16. So it's a second at concurrency of 16. So it's a second at concurrency of 16. So it's a trade-off. A few users keeps everyone trade-off. A few users keeps everyone trade-off. A few users keeps everyone snappy. Pack it full and you get maximum snappy. Pack it full and you get maximum snappy. Pack it full and you get maximum total throughput. And the longer the total throughput. And the longer the total throughput. And the longer the answers, the higher the total clients. answers, the higher the total clients. answers, the higher the total clients. But who actually needs thousands of But who actually needs thousands of But who actually needs thousands of tokens a second across dozens of tokens a second across dozens of tokens a second across dozens of parallel requests? H, that's not a parallel requests? H, that's not a parallel requests? H, that's not a person, that's an agent.
-
person, that's an agent. person, that's an agent. Running a lot of agents hides two Running a lot of agents hides two Running a lot of agents hides two different ideas. One is architecture. different ideas. One is architecture. different ideas. One is architecture. How do you wire agents together? Any How do you wire agents together? Any How do you wire agents together? Any laptop can do that, right? The other is laptop can do that, right? The other is laptop can do that, right? The other is throughput. How many actually run at the throughput. How many actually run at the throughput. How many actually run at the same time? That is hardware. The DJX same time? That is hardware. The DJX same time? That is hardware. The DJX doesn't change your design. It just doesn't change your design. It just doesn't change your design. It just raises your ceiling a little bit. So, raises your ceiling a little bit. So, raises your ceiling a little bit. So, the real question is, how many agents the real question is, how many agents the real question is, how many agents can this one box actually run at once? can this one box actually run at once? can this one box actually run at once? Let's get there. First, just one agent. Let's get there. First, just one agent. Let's get there. First, just one agent. Here's Visual Studio Code with its agent Here's Visual Studio Code with its agent Here's Visual Studio Code with its agent configured to use Neatron. I'm going to configured to use Neatron. I'm going to configured to use Neatron. I'm going to ask what is this project about? Boom. ask what is this project about? Boom. ask what is this project about? Boom. It's evaluating all the files, It's evaluating all the files, It's evaluating all the files, processing, and some of this is the processing, and some of this is the processing, and some of this is the thinking stage. This is an agent, so thinking stage. This is an agent, so thinking stage. This is an agent, so it's not just doing one thing at a time. it's not just doing one thing at a time. it's not just doing one thing at a time. It's constantly sending back and forth. It's constantly sending back and forth. It's constantly sending back and forth. And look how fast that goes. Wow. And I And look how fast that goes. Wow. And I And look how fast that goes. Wow. And I found out what it does. So, it has tool found out what it does. So, it has tool found out what it does. So, it has tool use, everything an agent needs. This use, everything an agent needs. This use, everything an agent needs. This project is a command line task tracker project is a command line task tracker project is a command line task tracker called TRK. That's right. I told you I called TRK. That's right. I told you I called TRK. That's right. I told you I tested four models. This is how fast tested four models. This is how fast tested four models. This is how fast each of the models is, including GLM each of the models is, including GLM each of the models is, including GLM 5.2, 2 Quan, Deep Seek, Neotron, 5.2, 2 Quan, Deep Seek, Neotron, 5.2, 2 Quan, Deep Seek, Neotron, GLM.
-
GLM. GLM. Notice GLM 5.2 is just a little bit Notice GLM 5.2 is just a little bit Notice GLM 5.2 is just a little bit slower. That's because it's spilled slower. That's because it's spilled slower. That's because it's spilled over. It can't all fit inside high over. It can't all fit inside high over. It can't all fit inside high bandwidth memory like Neatron can. Here, bandwidth memory like Neatron can. Here, bandwidth memory like Neatron can. Here, I hooked it up to Codeex. That's I hooked it up to Codeex. That's I hooked it up to Codeex. That's OpenAI's agent CLI. Model is Nvidia OpenAI's agent CLI. Model is Nvidia OpenAI's agent CLI. Model is Nvidia Neimatron Super 12B. And I'm just going Neimatron Super 12B. And I'm just going Neimatron Super 12B. And I'm just going to do a softball here. Write a story. to do a softball here. Write a story. to do a softball here. Write a story. Boom. And there is our proof that it's Boom. And there is our proof that it's Boom. And there is our proof that it's actually working. I'm running Envy Top actually working. I'm running Envy Top actually working. I'm running Envy Top on the actual machine watching it and on the actual machine watching it and on the actual machine watching it and the power is going up to about 715 the power is going up to about 715 the power is going up to about 715 watts, about 550 on average. 227 GB out watts, about 550 on average. 227 GB out watts, about 550 on average. 227 GB out of the 250 on high bandwidth memory. of the 250 on high bandwidth memory. of the 250 on high bandwidth memory. It's all fitting inside on the GB300. It's all fitting inside on the GB300. It's all fitting inside on the GB300. But one agent is easy. A laptop can do But one agent is easy. A laptop can do But one agent is easy. A laptop can do it. Even your Mac Mini can do it. Maybe it. Even your Mac Mini can do it. Maybe it. Even your Mac Mini can do it. Maybe not that fast. But here's the real test. not that fast. But here's the real test. not that fast. But here's the real test. Can this one box run a whole crowd of Can this one box run a whole crowd of Can this one box run a whole crowd of agents at the same time? Let's find out. agents at the same time? Let's find out. agents at the same time? Let's find out. So here I got a little agent view. This So here I got a little agent view. This So here I got a little agent view. This is the power bar. Right now it's at 205 is the power bar. Right now it's at 205 is the power bar. Right now it's at 205 watts. Temperature is 33. GPU is at 0%. watts. Temperature is 33. GPU is at 0%. watts. Temperature is 33. GPU is at 0%. And I can pick how many agents I want to And I can pick how many agents I want to And I can pick how many agents I want to launch. Let's start with 16, shall we? launch. Let's start with 16, shall we? launch. Let's start with 16, shall we? And launch. Here we go. GPU is up to And launch. Here we go. GPU is up to And launch. Here we go. GPU is up to 100%. Power is hitting about 750 watts.
-
100%. Power is hitting about 750 watts. 100%. Power is hitting about 750 watts. Temperatures gone up a little bit. Temperatures gone up a little bit. Temperatures gone up a little bit. Everything seems to be working and we're Everything seems to be working and we're Everything seems to be working and we're cranking those numbers out. Look at cranking those numbers out. Look at cranking those numbers out. Look at that. Live 16 agents, 1,400 tokens per that. Live 16 agents, 1,400 tokens per that. Live 16 agents, 1,400 tokens per second. And this is how many total second. And this is how many total second. And this is how many total tokens we've done so far. A lot of work tokens we've done so far. A lot of work tokens we've done so far. A lot of work is being done right now. A lot of work is being done right now. A lot of work is being done right now. A lot of work we are not going to use or see or we are not going to use or see or we are not going to use or see or whatever, but you know, it's a demo. We whatever, but you know, it's a demo. We whatever, but you know, it's a demo. We got different kinds of agents, too. A got different kinds of agents, too. A got different kinds of agents, too. A coder, a researcher, an analyst, a coder, a researcher, an analyst, a coder, a researcher, an analyst, a writer, a planner, a tester, an writer, a planner, a tester, an writer, a planner, a tester, an architect. You get the idea. Let's kick architect. You get the idea. Let's kick architect. You get the idea. Let's kick it up a notch. Let's go to 32. You know it up a notch. Let's go to 32. You know it up a notch. Let's go to 32. You know what? Let's go to 64. Boom. what? Let's go to 64. Boom. what? Let's go to 64. Boom. 900 watts. Wow. 3,500 tokens per second. 900 watts. Wow. 3,500 tokens per second. 900 watts. Wow. 3,500 tokens per second. And we're generating 35 37 40,000 And we're generating 35 37 40,000 And we're generating 35 37 40,000 tokens. I can't even keep up speaking tokens. I can't even keep up speaking tokens. I can't even keep up speaking it. Wow, look at them all go. These are it. Wow, look at them all go. These are it. Wow, look at them all go. These are all generating all at the same time. All all generating all at the same time. All all generating all at the same time. All doing their own thing. But we're not doing their own thing. But we're not doing their own thing. But we're not done yet. Let's try 128. And of course, done yet. Let's try 128. And of course, done yet. Let's try 128. And of course, you might have guessed this is going to you might have guessed this is going to you might have guessed this is going to work. And it does work. But look at the work. And it does work. But look at the work. And it does work. But look at the speed at which this is happening. We're speed at which this is happening. We're speed at which this is happening. We're not hitting the 1300 watts that this uh not hitting the 1300 watts that this uh not hitting the 1300 watts that this uh processor is capable of, but we're processor is capable of, but we're processor is capable of, but we're getting pretty close. We're about 1,000 getting pretty close. We're about 1,000 getting pretty close. We're about 1,000 watts right now if is what I've seen as watts right now if is what I've seen as watts right now if is what I've seen as maximum. Temperature is about 55° and so maximum. Temperature is about 55° and so maximum. Temperature is about 55° and so far since I've been talking, we've far since I've been talking, we've far since I've been talking, we've already generated almost 150,000 tokens.
-
already generated almost 150,000 tokens. already generated almost 150,000 tokens. This is a local agent swarm on the This is a local agent swarm on the This is a local agent swarm on the GB300. Now, the sweet spot's around 32 GB300. Now, the sweet spot's around 32 GB300. Now, the sweet spot's around 32 to 64 for this model. This again is to 64 for this model. This again is to 64 for this model. This again is Neatron right here. It's good enough to Neatron right here. It's good enough to Neatron right here. It's good enough to peg the GPU while every agent stays peg the GPU while every agent stays peg the GPU while every agent stays responsive and everyone gets the same responsive and everyone gets the same responsive and everyone gets the same big model. That's what a small machine big model. That's what a small machine big model. That's what a small machine can't do. 24/7, 100 agents, nothing can't do. 24/7, 100 agents, nothing can't do. 24/7, 100 agents, nothing leaves your desk. Zero per token cost. leaves your desk. Zero per token cost. leaves your desk. Zero per token cost. Of course, you have to buy the box. And Of course, you have to buy the box. And Of course, you have to buy the box. And these boxes are not cheap. Luckily, I these boxes are not cheap. Luckily, I these boxes are not cheap. Luckily, I didn't have to buy it. I'm borrowing it. didn't have to buy it. I'm borrowing it. didn't have to buy it. I'm borrowing it. Versus thousands of dollars a month in Versus thousands of dollars a month in Versus thousands of dollars a month in API bills. This is not bad. That's the API bills. This is not bad. That's the API bills. This is not bad. That's the pitch. Fast for one, scales for many, pitch. Fast for one, scales for many, pitch. Fast for one, scales for many, and 100 at once for free. Just pay your and 100 at once for free. Just pay your and 100 at once for free. Just pay your electricity bill, okay? Or they'll turn electricity bill, okay? Or they'll turn electricity bill, okay? Or they'll turn it off on you. Now, you can also put a it off on you. Now, you can also put a it off on you. Now, you can also put a GPU in one of those PCIe slots. And an GPU in one of those PCIe slots. And an GPU in one of those PCIe slots. And an RTX Pro 6000 is one of the ones you can RTX Pro 6000 is one of the ones you can RTX Pro 6000 is one of the ones you can actually fit in there. It's a nice GPU actually fit in there. It's a nice GPU actually fit in there. It's a nice GPU that also scales. If you can't buy a big that also scales. If you can't buy a big that also scales. If you can't buy a big machine like this, you can buy just a machine like this, you can buy just a machine like this, you can buy just a GPU. And you'll see my video right here GPU. And you'll see my video right here GPU. And you'll see my video right here on how I do that. Thanks for watching on how I do that. Thanks for watching on how I do that. Thanks for watching and I'll see you next time.
Summary
The main theme is the advent of the DGX station, a powerful desktop machine for AI development, featuring the GB300 chip and substantial unified memory. Key subjects include LLMs, agents, and comparisons to Apple Silicon's memory architecture, along with its high-speed networking capabilities. The practical takeaway is that advanced AI compute is now accessible on a personal workstation, enabling significant development without reliance on data centers.