← Back
Aaron Zisk August 31, 2026 14m

I Gave Local AI and the Cloud the Exact Same Job

Read full transcript 12 segments
  1. This is a 2.8 bill, This is a 2.8 bill, sorry, this is a 2.8 trillion parameter sorry, this is a 2.8 trillion parameter sorry, this is a 2.8 trillion parameter model and it looks like it's going to be model and it looks like it's going to be model and it looks like it's going to be the biggest one for a while. Kimmy K3, the biggest one for a while. Kimmy K3, the biggest one for a while. Kimmy K3, 1.56 terabytes on disk. It doesn't fit 1.56 terabytes on disk. It doesn't fit 1.56 terabytes on disk. It doesn't fit on a single Mac Studio, not even in Q4, on a single Mac Studio, not even in Q4, on a single Mac Studio, not even in Q4, not even in Q2. That's the quantization not even in Q2. That's the quantization not even in Q2. That's the quantization I'm talking about when you shrink the I'm talking about when you shrink the I'm talking about when you shrink the model from its original weights. So, I model from its original weights. So, I model from its original weights. So, I wired up four Mac Studios together, each wired up four Mac Studios together, each wired up four Mac Studios together, each with 512 GB for a total of 2 TB of with 512 GB for a total of 2 TB of with 512 GB for a total of 2 TB of unified memory. 2 TB minus 1.56 Yeah, unified memory. 2 TB minus 1.56 Yeah, unified memory. 2 TB minus 1.56 Yeah, we're good. And I wanted to see if it we're good. And I wanted to see if it we're good. And I wanted to see if it can beat the new Abacus AI can beat the new Abacus AI can beat the new Abacus AI supercomputer. I'm going to have Kimmy supercomputer. I'm going to have Kimmy supercomputer. I'm going to have Kimmy K3 build a web app, a front-end web app, K3 build a web app, a front-end web app, K3 build a web app, a front-end web app, and I'm going to have supercomputer do and I'm going to have supercomputer do and I'm going to have supercomputer do the same. It's local hardware versus the same. It's local hardware versus the same. It's local hardware versus cloud death match, round two. Let's go. cloud death match, round two. Let's go. cloud death match, round two. Let's go. So, last time I did this it was one Mac So, last time I did this it was one Mac So, last time I did this it was one Mac Studio, 512 gigs against Abacus deep Studio, 512 gigs against Abacus deep Studio, 512 gigs against Abacus deep agent. And a lot of you mentioned in the agent. And a lot of you mentioned in the agent. And a lot of you mentioned in the comments, "Alex, that's not a real local comments, "Alex, that's not a real local comments, "Alex, that's not a real local setup." Fair enough. So, this time I setup." Fair enough. So, this time I setup." Fair enough. So, this time I brought four Mac top and yeah, it's brought four Mac top and yeah, it's brought four Mac top and yeah, it's running on it right now charted across running on it right now charted across running on it right now charted across all four Mac Studios connected with all four Mac Studios connected with all four Mac Studios connected with RDMA. Mac Studio 1, 208 GB used, almost RDMA. Mac Studio 1, 208 GB used, almost RDMA. Mac Studio 1, 208 GB used, almost 200 used on Mac Studio 2, 177 200 used on Mac Studio 2, 177 200 used on Mac Studio 2, 177 on Mac 3, and 215 on Mac 4. For context, on Mac 3, and 215 on Mac 4. For context, on Mac 3, and 215 on Mac 4. For context, Kimmy K2 was 1 trillion parameters and Kimmy K2 was 1 trillion parameters and Kimmy K2 was 1 trillion parameters and that already felt pretty ridiculous.

  2. that already felt pretty ridiculous. that already felt pretty ridiculous. Kimmy K3's almost three times that. And Kimmy K3's almost three times that. And Kimmy K3's almost three times that. And the more parameters you have, the more the more parameters you have, the more the more parameters you have, the more the model can actually do. It's smarter. the model can actually do. It's smarter. the model can actually do. It's smarter. That's the whole trade-off. A 1 billion That's the whole trade-off. A 1 billion That's the whole trade-off. A 1 billion parameter model gives you gibberish. No parameter model gives you gibberish. No parameter model gives you gibberish. No that's Now, even though 1.56 that's Now, even though 1.56 that's Now, even though 1.56 TB in the original 8-bit precision, I'm TB in the original 8-bit precision, I'm TB in the original 8-bit precision, I'm still running a smaller quantization, still running a smaller quantization, still running a smaller quantization, but it's unpruned. I'm running all 896 but it's unpruned. I'm running all 896 but it's unpruned. I'm running all 896 experts. That's a total of 817 GB of experts. That's a total of 817 GB of experts. That's a total of 817 GB of weights sitting on disk. Now, you might weights sitting on disk. Now, you might weights sitting on disk. Now, you might think, "Oh, each one of those is 512, think, "Oh, each one of those is 512, think, "Oh, each one of those is 512, you can fit that on two, right?" Well, you can fit that on two, right?" Well, you can fit that on two, right?" Well, not exactly because I'm giving it a not exactly because I'm giving it a not exactly because I'm giving it a large amount of context, you're going to large amount of context, you're going to large amount of context, you're going to need a lot more space. And you could need a lot more space. And you could need a lot more space. And you could say, "Well, why not use three then? say, "Well, why not use three then? say, "Well, why not use three then? Because tensor parallelism splits the Because tensor parallelism splits the Because tensor parallelism splits the model dimensions across nodes. And those model dimensions across nodes. And those model dimensions across nodes. And those dimensions have to be divided evenly. dimensions have to be divided evenly. dimensions have to be divided evenly. So, it's either one, two, or four, or So, it's either one, two, or four, or So, it's either one, two, or four, or eight. Kimik4, maybe next year. And all eight. Kimik4, maybe next year. And all eight. Kimik4, maybe next year. And all four of these are connected through four of these are connected through four of these are connected through Thunderbolt 5 mesh on the back here. Six Thunderbolt 5 mesh on the back here. Six Thunderbolt 5 mesh on the back here. Six cables total connecting every machine to cables total connecting every machine to cables total connecting every machine to every other machine. I made a whole every other machine. I made a whole every other machine. I made a whole video describing how I did that and how video describing how I did that and how video describing how I did that and how I ran it with MLX distributed. You can I ran it with MLX distributed. You can I ran it with MLX distributed. You can check that out. I'll link to it down check that out. I'll link to it down check that out. I'll link to it down below. All right.

  3. below. All right. below. All right. Let's see your pretty face. Now, let's Let's see your pretty face. Now, let's Let's see your pretty face. Now, let's meet the other contender. This is Abacus meet the other contender. This is Abacus meet the other contender. This is Abacus AI, and they have a new tool called AI, and they have a new tool called AI, and they have a new tool called Supercomputer. You've probably already Supercomputer. You've probably already Supercomputer. You've probably already seen that they sponsored some of my seen that they sponsored some of my seen that they sponsored some of my content in the past, and they're content in the past, and they're content in the past, and they're sponsoring this video, too. But, I keep sponsoring this video, too. But, I keep sponsoring this video, too. But, I keep coming back to it on my own time because coming back to it on my own time because coming back to it on my own time because essentially it gives me one login to be essentially it gives me one login to be essentially it gives me one login to be able to do and use every frontier model able to do and use every frontier model able to do and use every frontier model out there on day one. There's over 100 out there on day one. There's over 100 out there on day one. There's over 100 models here, and route LLM basically models here, and route LLM basically models here, and route LLM basically routes you to the right one based on routes you to the right one based on routes you to the right one based on your context. So, there's GPT 5.6 Soul. your context. So, there's GPT 5.6 Soul. your context. So, there's GPT 5.6 Soul. That's pretty new as of this video. That's pretty new as of this video. That's pretty new as of this video. Llama 2 5, Cloud Fable 5, Grok 4.5. You Llama 2 5, Cloud Fable 5, Grok 4.5. You Llama 2 5, Cloud Fable 5, Grok 4.5. You know, they got image generation and know, they got image generation and know, they got image generation and things like that. They also have Kimik3 things like that. They also have Kimik3 things like that. They also have Kimik3 here, too. DeepSeek V4 Flash and Pro GLM here, too. DeepSeek V4 Flash and Pro GLM here, too. DeepSeek V4 Flash and Pro GLM 5.2. Open-source models, close-source 5.2. Open-source models, close-source 5.2. Open-source models, close-source models, everything. Kind of cool cuz models, everything. Kind of cool cuz models, everything. Kind of cool cuz you're not locked into a single you're not locked into a single you're not locked into a single ecosystem. Now, what's different about ecosystem. Now, what's different about ecosystem. Now, what's different about Supercomputer is that it's actually a Supercomputer is that it's actually a Supercomputer is that it's actually a machine, a virtual machine. It's not machine, a virtual machine. It's not machine, a virtual machine. It's not just a chat window. It's a cloud box just a chat window. It's a cloud box just a chat window. It's a cloud box that's always on that's also happens to that's always on that's also happens to that's always on that's also happens to be integrated into their agent. You be integrated into their agent. You be integrated into their agent. You don't wait for a cold start. You don't don't wait for a cold start. You don't don't wait for a cold start. You don't wait for a spin-up. You never see wait for a spin-up. You never see wait for a spin-up. You never see container provisioning because it's container provisioning because it's container provisioning because it's always on. Scheduled tasks, databases, always on. Scheduled tasks, databases, always on. Scheduled tasks, databases, storage, GitHub integration, SSH. And storage, GitHub integration, SSH. And storage, GitHub integration, SSH. And also, they got Hermes agent now. So, also, they got Hermes agent now. So, also, they got Hermes agent now. So, here I've got Hermes running, and I can here I've got Hermes running, and I can here I've got Hermes running, and I can change models to whatever I want. So, change models to whatever I want. So, change models to whatever I want. So, I've selected Kimik3. I can select I've selected Kimik3. I can select I've selected Kimik3. I can select whatever other model I want to talk to whatever other model I want to talk to whatever other model I want to talk to directly, like GPT 5 Luna.

  4. directly, like GPT 5 Luna. directly, like GPT 5 Luna. Say hi to Luna. So, you got Hermes. You Say hi to Luna. So, you got Hermes. You Say hi to Luna. So, you got Hermes. You don't need to install it locally. It's don't need to install it locally. It's don't need to install it locally. It's always on and running over there. I'll always on and running over there. I'll always on and running over there. I'll talk about the pricing in a bit. So, the talk about the pricing in a bit. So, the talk about the pricing in a bit. So, the fight is pretty simple. They're going to fight is pretty simple. They're going to fight is pretty simple. They're going to get the same prompt, both sides, and I get the same prompt, both sides, and I get the same prompt, both sides, and I got the prompt as a gist. You can check got the prompt as a gist. You can check got the prompt as a gist. You can check it out. I'll link to it down below. It's it out. I'll link to it down below. It's it out. I'll link to it down below. It's a pretty extensive long prompt, but it's a pretty extensive long prompt, but it's a pretty extensive long prompt, but it's specifically front end for a silicon specifically front end for a silicon specifically front end for a silicon compute exchange front end. I want to compute exchange front end. I want to compute exchange front end. I want to make sure we're using TypeScript with make sure we're using TypeScript with make sure we're using TypeScript with the latest and greatest front end tech. the latest and greatest front end tech. the latest and greatest front end tech. We're using mock data, not real data, We're using mock data, not real data, We're using mock data, not real data, but this can easily be swapped out but this can easily be swapped out but this can easily be swapped out later. Dark mode, microinteractions, later. Dark mode, microinteractions, later. Dark mode, microinteractions, accessibility, we should have a pretty accessibility, we should have a pretty accessibility, we should have a pretty decent-looking site in the end, decent-looking site in the end, decent-looking site in the end, hopefully. It's going to have some unit hopefully. It's going to have some unit hopefully. It's going to have some unit testing, and it needs to check its own testing, and it needs to check its own testing, and it needs to check its own work. So, this is the agentic part of work. So, this is the agentic part of work. So, this is the agentic part of it. It's going to be running all these it. It's going to be running all these it. It's going to be running all these commands to be able to verify and commands to be able to verify and commands to be able to verify and validate that it's working. Okay, let's validate that it's working. Okay, let's validate that it's working. Okay, let's get it. Raw copy. I'm going to head over get it. Raw copy. I'm going to head over get it. Raw copy. I'm going to head over to supercomputer and just paste this in to supercomputer and just paste this in to supercomputer and just paste this in into the agent. Now, here it's going to into the agent. Now, here it's going to into the agent. Now, here it's going to be set to auto, but you know what? I'm be set to auto, but you know what? I'm be set to auto, but you know what? I'm going to go with Opus 5 high and GPT 5.6 going to go with Opus 5 high and GPT 5.6 going to go with Opus 5 high and GPT 5.6 soul. Oh, yeah, you have access to the soul. Oh, yeah, you have access to the soul. Oh, yeah, you have access to the full virtual machine here on the right, full virtual machine here on the right, full virtual machine here on the right, too. You got the files that it's going too. You got the files that it's going too. You got the files that it's going to generate, access [music] directly to to generate, access [music] directly to to generate, access [music] directly to the terminal, and access to the desktop.

  5. the terminal, and access to the desktop. the terminal, and access to the desktop. Now, I don't want to give it a leg up, Now, I don't want to give it a leg up, Now, I don't want to give it a leg up, so I'm going to go and start this off in so I'm going to go and start this off in so I'm going to go and start this off in my Mac Studio cluster, too. And the way my Mac Studio cluster, too. And the way my Mac Studio cluster, too. And the way I set this up locally is I just have I set this up locally is I just have I set this up locally is I just have Open Code, which is the local agent, Open Code, which is the local agent, Open Code, which is the local agent, pretty popular, but you can point it to pretty popular, but you can point it to pretty popular, but you can point it to any local model that you want. And I've any local model that you want. And I've any local model that you want. And I've pointed it to the 4x Mac Studio cluster pointed it to the 4x Mac Studio cluster pointed it to the 4x Mac Studio cluster running Kimiko 3. I actually made it a running Kimiko 3. I actually made it a running Kimiko 3. I actually made it a separate video on how to set this up for separate video on how to set this up for separate video on how to set this up for members of the channel. By the way, members of the channel. By the way, members of the channel. By the way, thank you to the members for supporting thank you to the members for supporting thank you to the members for supporting the channel. Let's paste in that prompt, the channel. Let's paste in that prompt, the channel. Let's paste in that prompt, and boom. and boom. and boom. There we go. [music] It's doing stuff, I There we go. [music] It's doing stuff, I There we go. [music] It's doing stuff, I hope. How do I know? Here's all four of hope. How do I know? Here's all four of hope. How do I know? Here's all four of them, and the GPU usage is pegging 100% them, and the GPU usage is pegging 100% them, and the GPU usage is pegging 100% on all four of these machines. Do I hear on all four of these machines. Do I hear on all four of these machines. Do I hear it? No, I don't hear it. Do I feel it? It's No, I don't hear it. Do I feel it? It's warm. And if I put my ear right up into warm. And if I put my ear right up into warm. And if I put my ear right up into it, then I'll hear a little little it, then I'll hear a little little it, then I'll hear a little little twinkle twinkle little star kind of twinkle twinkle little star kind of twinkle twinkle little star kind of sound in there. Maybe the fans will kick sound in there. Maybe the fans will kick sound in there. Maybe the fans will kick up later, but it's working. Do we have up later, but it's working. Do we have up later, but it's working. Do we have anything built yet? No, it's thinking. anything built yet? No, it's thinking. anything built yet? No, it's thinking. Where is my output? All right. All Where is my output? All right. All Where is my output? All right. All right, it's doing it. I'm going to start right, it's doing it. I'm going to start right, it's doing it. I'm going to start the other one. Boom, I'll build that for the other one. Boom, I'll build that for the other one. Boom, I'll build that for you right now. This will basically do you right now. This will basically do you right now. This will basically do the same thing. It's building it. Let's the same thing. It's building it. Let's the same thing. It's building it. Let's pop this open on the right so we can see pop this open on the right so we can see pop this open on the right so we can see what's happening. By the way, this is what's happening. By the way, this is what's happening. By the way, this is the first time I'm actually running that the first time I'm actually running that the first time I'm actually running that prompt. I don't know what's going to prompt. I don't know what's going to prompt. I don't know what's going to happen.

  6. happen. happen. Hopefully it works. It says if you Hopefully it works. It says if you Hopefully it works. It says if you connect your GitHub, I can work with connect your GitHub, I can work with connect your GitHub, I can work with your repositories directly. That's your repositories directly. That's your repositories directly. That's pretty cool. Cloning, pushing, commits, pretty cool. Cloning, pushing, commits, pretty cool. Cloning, pushing, commits, and pull requests. Want me to set that and pull requests. Want me to set that and pull requests. Want me to set that up? Go ahead. So now Abacus AI agent up? Go ahead. So now Abacus AI agent up? Go ahead. So now Abacus AI agent takes over and we've seen it at work takes over and we've seen it at work takes over and we've seen it at work before, but the agent is now well before, but the agent is now well before, but the agent is now well integrated into the supercomputer integrated into the supercomputer integrated into the supercomputer concept that's always on and it's going concept that's always on and it's going concept that's always on and it's going to be deploying to this VM. Building a to be deploying to this VM. Building a to be deploying to this VM. Building a production-ready Next.js web app. Now production-ready Next.js web app. Now production-ready Next.js web app. Now compared to some of the other builds I compared to some of the other builds I compared to some of the other builds I I've done here on the channel with I've done here on the channel with I've done here on the channel with Nvidia and DGX Sparks, the prefill or Nvidia and DGX Sparks, the prefill or Nvidia and DGX Sparks, the prefill or the prompt processing stage of inference the prompt processing stage of inference the prompt processing stage of inference is something that Max are not as good at is something that Max are not as good at is something that Max are not as good at as, for example, Nvidia GPUs. Macs are as, for example, Nvidia GPUs. Macs are as, for example, Nvidia GPUs. Macs are really good at generating once they've really good at generating once they've really good at generating once they've processed because they have extremely processed because they have extremely processed because they have extremely good high memory bandwidth. And with good high memory bandwidth. And with good high memory bandwidth. And with Kimiko 3, I measured about a 238 tokens Kimiko 3, I measured about a 238 tokens Kimiko 3, I measured about a 238 tokens per second chewing through the prompt, per second chewing through the prompt, per second chewing through the prompt, which is not super fast. Oh. which is not super fast. Oh. which is not super fast. Oh. Now I can hear it. There's work to be Now I can hear it. There's work to be Now I can hear it. There's work to be had. Generation happens at about 14.7 had. Generation happens at about 14.7 had. Generation happens at about 14.7 tokens per second with Kimiko 3 here.

  7. tokens per second with Kimiko 3 here. tokens per second with Kimiko 3 here. It's not going to be fast, folks, but it It's not going to be fast, folks, but it It's not going to be fast, folks, but it is a big model. What's happening on the is a big model. What's happening on the is a big model. What's happening on the Abacus side? Well, we already have some Abacus side? Well, we already have some Abacus side? Well, we already have some HTML generated. Now while it's doing HTML generated. Now while it's doing HTML generated. Now while it's doing that, we can take a little peek under that, we can take a little peek under that, we can take a little peek under the hood and see what's happening. So the hood and see what's happening. So the hood and see what's happening. So I'm going to go in here and let's do uh I'm going to go in here and let's do uh I'm going to go in here and let's do uh this is just Ubuntu over here. We got this is just Ubuntu over here. We got this is just Ubuntu over here. We got skills. This is just a generic skills skills. This is just a generic skills skills. This is just a generic skills repo. It doesn't have any custom skills repo. It doesn't have any custom skills repo. It doesn't have any custom skills installed. So there is Silicon Exchange. installed. So there is Silicon Exchange. installed. So there is Silicon Exchange. Let's go in there. That's our app. Let's go in there. That's our app. Let's go in there. That's our app. Next.js space. And yeah, there's our Next.js space. And yeah, there's our Next.js space. And yeah, there's our scripts and Tailwind config, TypeScript scripts and Tailwind config, TypeScript scripts and Tailwind config, TypeScript configuration, Next.js config. I wonder configuration, Next.js config. I wonder configuration, Next.js config. I wonder what's in that style guide. I wonder if what's in that style guide. I wonder if what's in that style guide. I wonder if this is something that it generated just this is something that it generated just this is something that it generated just now based on my prompt or now based on my prompt or now based on my prompt or just a generic one for Next.js projects. just a generic one for Next.js projects. just a generic one for Next.js projects. Anybody know? >> [snorts] >> [snorts] >> Please, sir, I want a token. Now, >> Please, sir, I want a token. Now, >> Please, sir, I want a token. Now, nobody's going to believe me that this nobody's going to believe me that this nobody's going to believe me that this Kimmy actually works. I swear, I tried Kimmy actually works. I swear, I tried Kimmy actually works. I swear, I tried it. Didn't try this prompt, but I tried it. Didn't try this prompt, but I tried it. Didn't try this prompt, but I tried using it and it worked fine. Maybe this using it and it worked fine. Maybe this using it and it worked fine. Maybe this prompt is just too hard. Now, let's take prompt is just too hard. Now, let's take prompt is just too hard. Now, let's take that 14 tokens per second figure. We can that 14 tokens per second figure. We can that 14 tokens per second figure. We can use something called Abacus AI Desktop, use something called Abacus AI Desktop, use something called Abacus AI Desktop, which is their desktop companion app.

  8. which is their desktop companion app. which is their desktop companion app. It's got chat, co-work, and code. And It's got chat, co-work, and code. And It's got chat, co-work, and code. And this is basically like a local agent. this is basically like a local agent. this is basically like a local agent. They have a CLI, too, but they have a They have a CLI, too, but they have a They have a CLI, too, but they have a graphical interface. It's kind of like graphical interface. It's kind of like graphical interface. It's kind of like Codex. The difference is you can pick Codex. The difference is you can pick Codex. The difference is you can pick the model that you want right here from the model that you want right here from the model that you want right here from this drop-down. Fable 5 is there, Opus this drop-down. Fable 5 is there, Opus this drop-down. Fable 5 is there, Opus 5, GPT 5.6, Soul, Terra, Luna, Grok is 5, GPT 5.6, Soul, Terra, Luna, Grok is 5, GPT 5.6, Soul, Terra, Luna, Grok is there, and Kimmy K3 is there, too. So, there, and Kimmy K3 is there, too. So, there, and Kimmy K3 is there, too. So, let's see what we get here. Of course, let's see what we get here. Of course, let's see what we get here. Of course, I'm using code here, so I'm going to I'm using code here, so I'm going to I'm using code here, so I'm going to need to select a workspace. Write me a need to select a workspace. Write me a need to select a workspace. Write me a simple JavaScript Node application for simple JavaScript Node application for simple JavaScript Node application for adding two numbers. Boom. Kimmy K3. Now, adding two numbers. Boom. Kimmy K3. Now, adding two numbers. Boom. Kimmy K3. Now, this is not using my Kimmy K3. This is this is not using my Kimmy K3. This is this is not using my Kimmy K3. This is using the Kimmy K3 that Abacus uses. But using the Kimmy K3 that Abacus uses. But using the Kimmy K3 that Abacus uses. But it's writing a local application. So, we it's writing a local application. So, we it's writing a local application. So, we need to accept a few things here. Allow need to accept a few things here. Allow need to accept a few things here. Allow it. Let's do a bypass, and there it it. Let's do a bypass, and there it it. Let's do a bypass, and there it goes. Now, that is pretty fast. I think goes. Now, that is pretty fast. I think goes. Now, that is pretty fast. I think it's done. Here you can access the VS it's done. Here you can access the VS it's done. Here you can access the VS Code extension, browser extension. So, Code extension, browser extension. So, Code extension, browser extension. So, this desktop app has a lot of extra this desktop app has a lot of extra this desktop app has a lot of extra abilities that it adds. Let's see. Node, abilities that it adds. Let's see. Node, abilities that it adds. Let's see. Node, add five and four. Nine. Is that right? add five and four. Nine. Is that right? add five and four. Nine. Is that right? Think so. What's happening over here Think so. What's happening over here Think so. What's happening over here with our Kimmy? Oh, no token yet.

  9. with our Kimmy? Oh, no token yet. with our Kimmy? Oh, no token yet. Well, it's been a while, so I had to Well, it's been a while, so I had to Well, it's been a while, so I had to check what's going on, and check what's going on, and check what's going on, and it wasn't printing out any tokens at it wasn't printing out any tokens at it wasn't printing out any tokens at all. So, I'm going back into it to see all. So, I'm going back into it to see all. So, I'm going back into it to see if I can even get anything out of it. if I can even get anything out of it. if I can even get anything out of it. I'm going to just say hi. Does look like I'm going to just say hi. Does look like I'm going to just say hi. Does look like it's working. The GPUs are at 100%, so it's working. The GPUs are at 100%, so it's working. The GPUs are at 100%, so yeah. Hi, how can I help you today? It yeah. Hi, how can I help you today? It yeah. Hi, how can I help you today? It worked. Maybe I just wasn't patient worked. Maybe I just wasn't patient worked. Maybe I just wasn't patient enough for such a large prompt. Let's do enough for such a large prompt. Let's do enough for such a large prompt. Let's do this. Create a simple web app with a this. Create a simple web app with a this. Create a simple web app with a text box and a button. Well, look at text box and a button. Well, look at text box and a button. Well, look at that. that. that. It just printed out index.html for me. It just printed out index.html for me. It just printed out index.html for me. Only took about 3 minutes to 147 lines. Only took about 3 minutes to 147 lines. Only took about 3 minutes to 147 lines. It's a beautiful, beautiful app. Wow. It's a beautiful, beautiful app. Wow. It's a beautiful, beautiful app. Wow. Test, submit, and it works. Still Test, submit, and it works. Still Test, submit, and it works. Still printing out. It's not done yet. I've printing out. It's not done yet. I've printing out. It's not done yet. I've created a simple web app for you. It's created a simple web app for you. It's created a simple web app for you. It's It's doing the explainer. Thank you. It's doing the explainer. Thank you. It's doing the explainer. Thank you. Thank you for that, Kimmy. I mean, you Thank you for that, Kimmy. I mean, you Thank you for that, Kimmy. I mean, you did a fantastic job. I just wish you did did a fantastic job. I just wish you did did a fantastic job. I just wish you did it faster. it faster. it faster. Okay, after all that, I couldn't just Okay, after all that, I couldn't just Okay, after all that, I couldn't just leave it like that. I couldn't leave you leave it like that. I couldn't leave you leave it like that. I couldn't leave you hanging. So, I stopped recording, had hanging. So, I stopped recording, had hanging. So, I stopped recording, had some coffee, went back and gave Kimmy K3 some coffee, went back and gave Kimmy K3 some coffee, went back and gave Kimmy K3 the full prompt again. But, this time I the full prompt again. But, this time I the full prompt again. But, this time I didn't sit there and watch it because, didn't sit there and watch it because, didn't sit there and watch it because, you know, watching something It's kind you know, watching something It's kind you know, watching something It's kind of like watching grass grow. You don't of like watching grass grow. You don't of like watching grass grow. You don't think anything's happening, but think anything's happening, but think anything's happening, but something is. I had to walk away and something is. I had to walk away and something is. I had to walk away and have more coffee. Guess what happened 4 have more coffee. Guess what happened 4 have more coffee. Guess what happened 4 hours later? We have a full app. Here it hours later? We have a full app. Here it hours later? We have a full app. Here it is.

  10. is. is. And it looks pretty good. This was done And it looks pretty good. This was done And it looks pretty good. This was done by Kimmy on these machines. There's the by Kimmy on these machines. There's the by Kimmy on these machines. There's the browse. We have live filtering, reset browse. We have live filtering, reset browse. We have live filtering, reset all filters, inspect, compare. We have all filters, inspect, compare. We have all filters, inspect, compare. We have charts. We have scheduling. This is a charts. We have scheduling. This is a charts. We have scheduling. This is a beautiful interface. And it can do light beautiful interface. And it can do light beautiful interface. And it can do light mode. It can do dark mode. Sorry about mode. It can do dark mode. Sorry about mode. It can do dark mode. Sorry about that. What do you think? I think it did that. What do you think? I think it did that. What do you think? I think it did a fantastic job here. So, the question, a fantastic job here. So, the question, a fantastic job here. So, the question, can you do this locally? The answer is can you do this locally? The answer is can you do this locally? The answer is yes. Yes, you can. Let's take a look at yes. Yes, you can. Let's take a look at yes. Yes, you can. Let's take a look at what Abacus came up with. It looks like what Abacus came up with. It looks like what Abacus came up with. It looks like it's actually finished. So, let's pop it's actually finished. So, let's pop it's actually finished. So, let's pop open our VM here. Still looking at open our VM here. Still looking at open our VM here. Still looking at exchange, and then what's the name of exchange, and then what's the name of exchange, and then what's the name of that? Next.js app. And we're going to do that? Next.js app. And we're going to do that? Next.js app. And we're going to do npm run dev. Okay, it built and started npm run dev. Okay, it built and started npm run dev. Okay, it built and started it at this URL. Let's take a look at it at this URL. Let's take a look at it at this URL. Let's take a look at that. Okay. Wow. Can't get too excited that. Okay. Wow. Can't get too excited that. Okay. Wow. Can't get too excited cuz it is a sponsored video. Wow. Oh my cuz it is a sponsored video. Wow. Oh my cuz it is a sponsored video. Wow. Oh my gosh. gosh. gosh. >> [laughter] >> [laughter] >> [laughter] >> I mean, it is pretty cool. It's very >> I mean, it is pretty cool. It's very >> I mean, it is pretty cool. It's very cool. It's a beautiful application. The cool. It's a beautiful application. The cool. It's a beautiful application. The design is quite something. The dashboard design is quite something. The dashboard design is quite something. The dashboard is beautiful. You got filters here on is beautiful. You got filters here on is beautiful. You got filters here on the left doing live filtering with the left doing live filtering with the left doing live filtering with animation. This is from a single prompt.

  11. animation. This is from a single prompt. animation. This is from a single prompt. I mean, this is not hooked up to a back I mean, this is not hooked up to a back I mean, this is not hooked up to a back end, but it wouldn't be that difficult end, but it wouldn't be that difficult end, but it wouldn't be that difficult to take it and extend it there. Your to take it and extend it there. Your to take it and extend it there. Your reservations, compare boards, browse reservations, compare boards, browse reservations, compare boards, browse inventory goes back to that page. Wait, inventory goes back to that page. Wait, inventory goes back to that page. Wait, what does this do over here? Let's do what does this do over here? Let's do what does this do over here? Let's do Apple M3 Ultra available. Simple web Apple M3 Ultra available. Simple web Apple M3 Ultra available. Simple web app. Look, I'm sure the capability is app. Look, I'm sure the capability is app. Look, I'm sure the capability is there of this model there of this model there of this model with the right hardware. I'm capable to with the right hardware. I'm capable to with the right hardware. I'm capable to run it. It's just not very fast. And I run it. It's just not very fast. And I run it. It's just not very fast. And I like to code. And the way I use the like to code. And the way I use the like to code. And the way I use the agents is to do these one-off tools that agents is to do these one-off tools that agents is to do these one-off tools that I use. My GitHub lately has been a bunch I use. My GitHub lately has been a bunch I use. My GitHub lately has been a bunch of different tools that I implemented of different tools that I implemented of different tools that I implemented using the help of agents. These are using the help of agents. These are using the help of agents. These are things that I want to knock out quickly things that I want to knock out quickly things that I want to knock out quickly and get working. And for that kind of and get working. And for that kind of and get working. And for that kind of stuff, Abacus supercomputer is perfect. stuff, Abacus supercomputer is perfect. stuff, Abacus supercomputer is perfect. Same prompt, same app, both finished it. Same prompt, same app, both finished it. Same prompt, same app, both finished it. 15 minutes against 4 hours. Now, each 15 minutes against 4 hours. Now, each 15 minutes against 4 hours. Now, each one of those Mac Studios was about one of those Mac Studios was about one of those Mac Studios was about $16,000, but now I think they're not for $16,000, but now I think they're not for $16,000, but now I think they're not for sale anymore, so it's a lot more. And I sale anymore, so it's a lot more. And I sale anymore, so it's a lot more. And I don't know how much the new ones are don't know how much the new ones are don't know how much the new ones are going to cost, but you get the idea. going to cost, but you get the idea. going to cost, but you get the idea. It's expensive hardware. You buy it one It's expensive hardware. You buy it one It's expensive hardware. You buy it one time, then it's yours, and then you do time, then it's yours, and then you do time, then it's yours, and then you do whatever you want with it. Abacus starts whatever you want with it. Abacus starts whatever you want with it. Abacus starts at 10 bucks a month. Actually, $7 for at 10 bucks a month. Actually, $7 for at 10 bucks a month. Actually, $7 for the first month now. You could run that the first month now. You could run that the first month now. You could run that subscription for over 300 years before subscription for over 300 years before subscription for over 300 years before you spend that kind of money. And you you spend that kind of money. And you you spend that kind of money. And you can use their web UI, you can use the can use their web UI, you can use the can use their web UI, you can use the desktop application that I showed you.

  12. desktop application that I showed you. desktop application that I showed you. You can use their APIs. You're also not You can use their APIs. You're also not You can use their APIs. You're also not stuck at one price because there's stuck at one price because there's stuck at one price because there's custom routers that let you push the custom routers that let you push the custom routers that let you push the easy work to cheaper models and save the easy work to cheaper models and save the easy work to cheaper models and save the expensive ones for the harder stuff. So, expensive ones for the harder stuff. So, expensive ones for the harder stuff. So, the bottom line is, and it's not what the bottom line is, and it's not what the bottom line is, and it's not what sponsors usually want to hear, Yes, you sponsors usually want to hear, Yes, you sponsors usually want to hear, Yes, you can do all the stuff locally, and that's can do all the stuff locally, and that's can do all the stuff locally, and that's pretty remarkable. That wasn't possible pretty remarkable. That wasn't possible pretty remarkable. That wasn't possible just a year ago. But, if you want high just a year ago. But, if you want high just a year ago. But, if you want high quality, and you want it done fast, and quality, and you want it done fast, and quality, and you want it done fast, and you want to be able to host it all in you want to be able to host it all in you want to be able to host it all in one shot, Abacus's cloud solution wins one shot, Abacus's cloud solution wins one shot, Abacus's cloud solution wins this one. If your data can't leave the this one. If your data can't leave the this one. If your data can't leave the building, or you just love running this building, or you just love running this building, or you just love running this stuff yourself, like I do, Cluster. stuff yourself, like I do, Cluster. stuff yourself, like I do, Cluster. There's definitely a market for both, There's definitely a market for both, There's definitely a market for both, not one or the other. To see me build not one or the other. To see me build not one or the other. To see me build and configure that cluster and all the and configure that cluster and all the and configure that cluster and all the details of how I run it, watch this details of how I run it, watch this details of how I run it, watch this video right here. And watch this video video right here. And watch this video video right here. And watch this video when I compare a Mac Studio against when I compare a Mac Studio against when I compare a Mac Studio against Abacus's AI. By the way, that website is Abacus's AI. By the way, that website is Abacus's AI. By the way, that website is still up and running. still up and running. still up and running. >> [music] >> [music] >> [music] >> Thanks for watching, and I'll see you >> Thanks for watching, and I'll see you >> Thanks for watching, and I'll see you next time.

Summary

The analysis discusses the technical challenges and setup for running a massive 2.8 trillion parameter AI model, Kimmy K3, on local hardware by connecting four Mac Studios. It highlights the immense storage requirements, even with quantization, and explores the concept of tensor parallelism for distributing the model across machines. The practical takeaway is that achieving high-performance local AI execution of such large models requires sophisticated hardware configurations and network setups to overcome cloud limitations.

View original episode ↗