← Back
Theo August 29, 2026 25m

NVIDIA Just Lost Their Lead

Read full transcript 21 segments
  1. Nvidia's a weird company. They Nvidia's a weird company. They originally started making chips for originally started making chips for originally started making chips for gamers to have fancy 3D graphics in gamers to have fancy 3D graphics in gamers to have fancy 3D graphics in their computer games, and now they are their computer games, and now they are their computer games, and now they are the most valuable company in the world the most valuable company in the world the most valuable company in the world powering the majority of intelligence powering the majority of intelligence powering the majority of intelligence that we get from all the fancy AI tools that we get from all the fancy AI tools that we get from all the fancy AI tools we use every single day. There's a lot we use every single day. There's a lot we use every single day. There's a lot of pieces that make Nvidia's borderline of pieces that make Nvidia's borderline of pieces that make Nvidia's borderline monopoly on compute for intelligence monopoly on compute for intelligence monopoly on compute for intelligence just impossible to crack. Things like just impossible to crack. Things like just impossible to crack. Things like CUDA, which is the language and system CUDA, which is the language and system CUDA, which is the language and system of choice for the vast majority of AI of choice for the vast majority of AI of choice for the vast majority of AI research, to the crazy deals that research, to the crazy deals that research, to the crazy deals that they're cutting with everybody from the they're cutting with everybody from the they're cutting with everybody from the government itself in the United States government itself in the United States government itself in the United States to businesses like of course OpenAI, to businesses like of course OpenAI, to businesses like of course OpenAI, Anthropic, XAI, and more. And it seems Anthropic, XAI, and more. And it seems Anthropic, XAI, and more. And it seems like everyone is realizing that Nvidia like everyone is realizing that Nvidia like everyone is realizing that Nvidia is a not great core dependency to have is a not great core dependency to have is a not great core dependency to have on your business's success. OpenAI could on your business's success. OpenAI could on your business's success. OpenAI could not function if Nvidia decided they not function if Nvidia decided they not function if Nvidia decided they wanted to start charging them 10 times wanted to start charging them 10 times wanted to start charging them 10 times more. The US government would be pissed more. The US government would be pissed more. The US government would be pissed if their whole plan to build this crazy if their whole plan to build this crazy if their whole plan to build this crazy set of data centers fails because Nvidia set of data centers fails because Nvidia set of data centers fails because Nvidia changes their mind. A whole of these changes their mind. A whole of these changes their mind. A whole of these businesses are struggling to work with businesses are struggling to work with businesses are struggling to work with the terms Nvidia's giving them today. the terms Nvidia's giving them today. the terms Nvidia's giving them today. And with the promise that those terms And with the promise that those terms And with the promise that those terms will get worse tomorrow, they're all will get worse tomorrow, they're all will get worse tomorrow, they're all looking for ways out. But there's one looking for ways out. But there's one looking for ways out. But there's one particular place that is more motivated particular place that is more motivated particular place that is more motivated than anywhere else to solve this.

  2. than anywhere else to solve this. than anywhere else to solve this. China. Nvidia chips have largely been China. Nvidia chips have largely been China. Nvidia chips have largely been banned from being sent to China by the banned from being sent to China by the banned from being sent to China by the US government. Some have been cleared US government. Some have been cleared US government. Some have been cleared for export, but it's a very very small for export, but it's a very very small for export, but it's a very very small set, and it's not the giant powerful set, and it's not the giant powerful set, and it's not the giant powerful GPUs that we are using for training in GPUs that we are using for training in GPUs that we are using for training in the US. All of this is pretty much the US. All of this is pretty much the US. All of this is pretty much forced the Chinese labs like ZAI to rely forced the Chinese labs like ZAI to rely forced the Chinese labs like ZAI to rely heavily on chips that are being made by heavily on chips that are being made by heavily on chips that are being made by Chinese companies like Huawei. But it's Chinese companies like Huawei. But it's Chinese companies like Huawei. But it's not even just the Chinese labs anymore. not even just the Chinese labs anymore. not even just the Chinese labs anymore. Even OpenAI has started developing their Even OpenAI has started developing their Even OpenAI has started developing their own chips to get faster, cheaper own chips to get faster, cheaper own chips to get faster, cheaper inference and training available to them inference and training available to them inference and training available to them internally. All of this has Nvidia internally. All of this has Nvidia internally. All of this has Nvidia acting pretty scared and very very acting pretty scared and very very acting pretty scared and very very weird. From crashing out on Jim Cramer's weird. From crashing out on Jim Cramer's weird. From crashing out on Jim Cramer's show to buying Hugging Face, a little show to buying Hugging Face, a little show to buying Hugging Face, a little bit absurd. But as you guys can probably bit absurd. But as you guys can probably bit absurd. But as you guys can probably guess, I have a lot of thoughts on this. guess, I have a lot of thoughts on this. guess, I have a lot of thoughts on this. I'm a big nerd about chips. I'm a big I'm a big nerd about chips. I'm a big I'm a big nerd about chips. I'm a big nerd about Nvidia in particular, and nerd about Nvidia in particular, and nerd about Nvidia in particular, and never been the biggest fan of them. And never been the biggest fan of them. And never been the biggest fan of them. And the opportunity to point out how their the opportunity to point out how their the opportunity to point out how their monopoly is about to be crushed and monopoly is about to be crushed and monopoly is about to be crushed and crumbled is one I have to take. But crumbled is one I have to take. But crumbled is one I have to take. But unlike my friends who invested in Nvidia unlike my friends who invested in Nvidia unlike my friends who invested in Nvidia early, I don't have billions of dollars early, I don't have billions of dollars early, I don't have billions of dollars to spare, so I'm going to have to take a to spare, so I'm going to have to take a to spare, so I'm going to have to take a quick break for today's sponsor. I got a quick break for today's sponsor. I got a quick break for today's sponsor. I got a hot take for you guys. The knowledge in hot take for you guys. The knowledge in hot take for you guys. The knowledge in existing LLMs is nowhere near enough for existing LLMs is nowhere near enough for existing LLMs is nowhere near enough for us to do real work and get our jobs us to do real work and get our jobs us to do real work and get our jobs done. Thankfully, the labs have noticed done. Thankfully, the labs have noticed done. Thankfully, the labs have noticed this, too, and that's why the majority this, too, and that's why the majority this, too, and that's why the majority of the context your agents work with of the context your agents work with of the context your agents work with isn't stuff that they already knew, it's isn't stuff that they already knew, it's isn't stuff that they already knew, it's stuff that they got from the internet.

  3. stuff that they got from the internet. stuff that they got from the internet. The vast majority of the context in our The vast majority of the context in our The vast majority of the context in our real context windows is things that the real context windows is things that the real context windows is things that the agents fetched from online. What I'm agents fetched from online. What I'm agents fetched from online. What I'm trying to say is that your agents need trying to say is that your agents need trying to say is that your agents need access to the internet, and if they want access to the internet, and if they want access to the internet, and if they want the best possible way to do it, you the best possible way to do it, you the best possible way to do it, you should probably use today's sponsor, should probably use today's sponsor, should probably use today's sponsor, Browserbase. They provide all of the Browserbase. They provide all of the Browserbase. They provide all of the pieces your agents need to be smarter pieces your agents need to be smarter pieces your agents need to be smarter and have modern knowledge. From their and have modern knowledge. From their and have modern knowledge. From their search API, which is literally a single search API, which is literally a single search API, which is literally a single curl request that allows your agents to curl request that allows your agents to curl request that allows your agents to look things up and find real context look things up and find real context look things up and find real context from the web, to their fetch API, which from the web, to their fetch API, which from the web, to their fetch API, which lets agents give a URL to Browserbase lets agents give a URL to Browserbase lets agents give a URL to Browserbase and get back markdown that can actually and get back markdown that can actually and get back markdown that can actually parse and understand from pretty much parse and understand from pretty much parse and understand from pretty much any website across the entire web. And any website across the entire web. And any website across the entire web. And don't forget about the browser as a don't forget about the browser as a don't forget about the browser as a service, which allows your agents to service, which allows your agents to service, which allows your agents to actually navigate the web like a human actually navigate the web like a human actually navigate the web like a human would with a real Chromium browser, would with a real Chromium browser, would with a real Chromium browser, allowing them to fill out forms, sign allowing them to fill out forms, sign allowing them to fill out forms, sign into pages, and do real actions on into pages, and do real actions on into pages, and do real actions on users' behalfs. Over 85% of the APIs on users' behalfs. Over 85% of the APIs on users' behalfs. Over 85% of the APIs on the web can't be accessed by simply the web can't be accessed by simply the web can't be accessed by simply curling them. If you're okay with only curling them. If you're okay with only curling them. If you're okay with only having that 15%, stick to curl. But if having that 15%, stick to curl. But if having that 15%, stick to curl. But if you want to unblock your agents and give you want to unblock your agents and give you want to unblock your agents and give them the whole web, do it today at them the whole web, do it today at them the whole web, do it today at soda.link/browserbase. soda.link/browserbase. soda.link/browserbase. Before we can talk about how Nvidia Before we can talk about how Nvidia Before we can talk about how Nvidia loses, it's important to understand how loses, it's important to understand how loses, it's important to understand how they won. There are a couple core they won. There are a couple core they won. There are a couple core pieces, but I want to just fixate on the pieces, but I want to just fixate on the pieces, but I want to just fixate on the two that I'm most interested in right two that I'm most interested in right two that I'm most interested in right now. Those pieces are, of course, the now. Those pieces are, of course, the now. Those pieces are, of course, the chips that are being used by all of chips that are being used by all of chips that are being used by all of these companies that Nvidia makes.

  4. these companies that Nvidia makes. these companies that Nvidia makes. But the other piece is CUDA. We'll come But the other piece is CUDA. We'll come But the other piece is CUDA. We'll come back to this one in a bit because this back to this one in a bit because this back to this one in a bit because this is where things will get complicated. is where things will get complicated. is where things will get complicated. But for now, I want to focus on the But for now, I want to focus on the But for now, I want to focus on the chips side. GPUs are uniquely tailored chips side. GPUs are uniquely tailored chips side. GPUs are uniquely tailored for these types of AI tasks because they for these types of AI tasks because they for these types of AI tasks because they have millions of small cores instead of have millions of small cores instead of have millions of small cores instead of a handful of big ones. In order to a handful of big ones. In order to a handful of big ones. In order to traverse these gigantic piles of weights traverse these gigantic piles of weights traverse these gigantic piles of weights and parameters that a model is, you have and parameters that a model is, you have and parameters that a model is, you have to trail all of the data inside of it to to trail all of the data inside of it to to trail all of the data inside of it to make a response. You get a lot of value make a response. You get a lot of value make a response. You get a lot of value from having a chip that has lots of from having a chip that has lots of from having a chip that has lots of different processors on it that can different processors on it that can different processors on it that can handle that type of large amounts of handle that type of large amounts of handle that type of large amounts of data and crazy matrix transform data and crazy matrix transform data and crazy matrix transform [ __ ] Lots of small brains are a lot [ __ ] Lots of small brains are a lot [ __ ] Lots of small brains are a lot more powerful than one giant one in more powerful than one giant one in more powerful than one giant one in these cases. You also need to have these cases. You also need to have these cases. You also need to have enough memory to hold the weights that enough memory to hold the weights that enough memory to hold the weights that are being traversed, which is another are being traversed, which is another are being traversed, which is another real fun problem that Nvidia is real fun problem that Nvidia is real fun problem that Nvidia is consistently causing for themselves. The consistently causing for themselves. The consistently causing for themselves. The two pieces of a given I'll change this two pieces of a given I'll change this two pieces of a given I'll change this to be GPUs. The two pieces that matter to be GPUs. The two pieces that matter to be GPUs. The two pieces that matter for a GPU are the actual like process for a GPU are the actual like process for a GPU are the actual like process itself. itself. itself. I'll say the chip and the high-bandwidth I'll say the chip and the high-bandwidth I'll say the chip and the high-bandwidth memory are probably the easiest way to memory are probably the easiest way to memory are probably the easiest way to do this. The GPU has these two parts.

  5. do this. The GPU has these two parts. do this. The GPU has these two parts. The actual silicon that has all of the The actual silicon that has all of the The actual silicon that has all of the things on it that can process all of things on it that can process all of things on it that can process all of this data and then the high-bandwidth this data and then the high-bandwidth this data and then the high-bandwidth memory, which actually holds said data. memory, which actually holds said data. memory, which actually holds said data. Nvidia has had a lot of fun with this Nvidia has had a lot of fun with this Nvidia has had a lot of fun with this split and arguably using it to fix split and arguably using it to fix split and arguably using it to fix prices in their favor. I won't say it's prices in their favor. I won't say it's prices in their favor. I won't say it's cheap, but getting a powerful Nvidia cheap, but getting a powerful Nvidia cheap, but getting a powerful Nvidia chip is not the hardest or most chip is not the hardest or most chip is not the hardest or most expensive thing in the world. One of the expensive thing in the world. One of the expensive thing in the world. One of the highest-end Nvidia GPUs available in highest-end Nvidia GPUs available in highest-end Nvidia GPUs available in terms of its actual like GPU performance terms of its actual like GPU performance terms of its actual like GPU performance and throughput is the RTX 5090. The and throughput is the RTX 5090. The and throughput is the RTX 5090. The retail price for the RTX 5090 was retail price for the RTX 5090 was retail price for the RTX 5090 was originally $2,000. They now consistently originally $2,000. They now consistently originally $2,000. They now consistently go for around 4,500 because the demand go for around 4,500 because the demand go for around 4,500 because the demand is so absurd. There's also a server card is so absurd. There's also a server card is so absurd. There's also a server card they put out called the RTX Pro 6000 they put out called the RTX Pro 6000 they put out called the RTX Pro 6000 that ranges between 12,000 and $16,000. that ranges between 12,000 and $16,000. that ranges between 12,000 and $16,000. You would assume this chip must be way, You would assume this chip must be way, You would assume this chip must be way, way faster than the 5090 if it's going way faster than the 5090 if it's going way faster than the 5090 if it's going to be six to 10 times more expensive. to be six to 10 times more expensive. to be six to 10 times more expensive. But guess what? It is the same speed or But guess what? It is the same speed or But guess what? It is the same speed or slower. This might sound crazy. You slower. This might sound crazy. You slower. This might sound crazy. You might be confused. How is the chip that might be confused. How is the chip that might be confused. How is the chip that costs six to eight times more slower?

  6. costs six to eight times more slower? costs six to eight times more slower? Why would I ever get that instead of the Why would I ever get that instead of the Why would I ever get that instead of the 5090? Well, it turns out the chip isn't 5090? Well, it turns out the chip isn't 5090? Well, it turns out the chip isn't the only thing that matters. the only thing that matters. the only thing that matters. The 96 gigs of GDDR7 The 96 gigs of GDDR7 The 96 gigs of GDDR7 are what mattered here. The 5090 only are what mattered here. The 5090 only are what mattered here. The 5090 only gets 32 gigs of GDDR7 and it's not error gets 32 gigs of GDDR7 and it's not error gets 32 gigs of GDDR7 and it's not error correcting. Technically, the RTX Pro has correcting. Technically, the RTX Pro has correcting. Technically, the RTX Pro has more CUDA cores, but I believe the more CUDA cores, but I believe the more CUDA cores, but I believe the process on the 5090 is slightly like process on the 5090 is slightly like process on the 5090 is slightly like newer and better. It should be within newer and better. It should be within newer and better. It should be within spitting distance for the raw spitting distance for the raw spitting distance for the raw performance and throughput there, but performance and throughput there, but performance and throughput there, but the RAM difference is why they're able the RAM difference is why they're able the RAM difference is why they're able to charge so much more. to charge so much more. to charge so much more. And just to be explicitly clear here, And just to be explicitly clear here, And just to be explicitly clear here, they were doing this way before RAM got they were doing this way before RAM got they were doing this way before RAM got more expensive. They do this because more expensive. They do this because more expensive. They do this because they know the people who need a fast they know the people who need a fast they know the people who need a fast chip and a lot of RAM are much more chip and a lot of RAM are much more chip and a lot of RAM are much more willing to spend lots of money than a willing to spend lots of money than a willing to spend lots of money than a video game player is. And we're already video game player is. And we're already video game player is. And we're already seeing my favorite question in chat, seeing my favorite question in chat, seeing my favorite question in chat, which is why not just connect them? which is why not just connect them? which is why not just connect them? I'm going to use my favorite analogy I I'm going to use my favorite analogy I I'm going to use my favorite analogy I always use for chip-related stuff here. always use for chip-related stuff here. always use for chip-related stuff here. A kitchen. Imagine you have a kitchen A kitchen. Imagine you have a kitchen A kitchen. Imagine you have a kitchen that serves a bunch of customers at your that serves a bunch of customers at your that serves a bunch of customers at your restaurant. You have three really restaurant. You have three really restaurant. You have three really talented chefs in there, but your talented chefs in there, but your talented chefs in there, but your freezer is full and your refrigerator is freezer is full and your refrigerator is freezer is full and your refrigerator is full. You're running out of space. So, full. You're running out of space. So, full. You're running out of space. So, you decide you need another fridge.

  7. you decide you need another fridge. you decide you need another fridge. You're out of space in your restaurant You're out of space in your restaurant You're out of space in your restaurant though, but you need that other fridge though, but you need that other fridge though, but you need that other fridge pretty bad. So, you buy a building a pretty bad. So, you buy a building a pretty bad. So, you buy a building a mile away and you put the fridge there. mile away and you put the fridge there. mile away and you put the fridge there. Maybe you put 10 fridges and five giant Maybe you put 10 fridges and five giant Maybe you put 10 fridges and five giant deep freezers there. You massively deep freezers there. You massively deep freezers there. You massively increase the amount of storage you have, increase the amount of storage you have, increase the amount of storage you have, but now every time the chef needs but now every time the chef needs but now every time the chef needs something from those fridges or something from those fridges or something from those fridges or freezers, they have to run a mile to the freezers, they have to run a mile to the freezers, they have to run a mile to the other place, grab it, and then run back. other place, grab it, and then run back. other place, grab it, and then run back. You can't just plug the chips in You can't just plug the chips in You can't just plug the chips in together. If you have a model that is 40 together. If you have a model that is 40 together. If you have a model that is 40 gigs, hell, if you have a model that's gigs, hell, if you have a model that's gigs, hell, if you have a model that's 33 gigs and it doesn't fit on your 5090, 33 gigs and it doesn't fit on your 5090, 33 gigs and it doesn't fit on your 5090, you now have to split the data across you now have to split the data across you now have to split the data across the two GPUs, which means ideally you the two GPUs, which means ideally you the two GPUs, which means ideally you have some way of predicting which data have some way of predicting which data have some way of predicting which data is needed when and how so that GPU one is needed when and how so that GPU one is needed when and how so that GPU one only has the data it needs and two only only has the data it needs and two only only has the data it needs and two only needs the data that it needs. Good luck. needs the data that it needs. Good luck. needs the data that it needs. Good luck. Have fun. Not [ __ ] happening. There Have fun. Not [ __ ] happening. There Have fun. Not [ __ ] happening. There are techniques around things like are techniques around things like are techniques around things like mixture of experts models that allow you mixture of experts models that allow you mixture of experts models that allow you to more easily some amount assign the to more easily some amount assign the to more easily some amount assign the work across stuff. But there is no way work across stuff. But there is no way work across stuff. But there is no way on consumer hardware to get reasonable on consumer hardware to get reasonable on consumer hardware to get reasonable bandwidth between two 5090s to allow bandwidth between two 5090s to allow bandwidth between two 5090s to allow them to share memory. By the way, that them to share memory. By the way, that them to share memory. By the way, that memory that they're using for this is as memory that they're using for this is as memory that they're using for this is as much as 48 gigabits per second. That much as 48 gigabits per second. That much as 48 gigabits per second. That means you need something even faster in means you need something even faster in means you need something even faster in order to transfer the data between the order to transfer the data between the order to transfer the data between the GPUs. Once again, GPUs. Once again, GPUs. Once again, not happening. Don't worry though, not happening. Don't worry though, not happening. Don't worry though, Nvidia has an answer for you. If you Nvidia has an answer for you. If you Nvidia has an answer for you. If you need more memory, they're more than need more memory, they're more than need more memory, they're more than willing to sell you something. And no, willing to sell you something. And no, willing to sell you something. And no, I'm not referring to the RTX Pro. They I'm not referring to the RTX Pro. They I'm not referring to the RTX Pro. They know that there are people who can't put

  8. know that there are people who can't put know that there are people who can't put out 18 grand plus on a GPU that they're out 18 grand plus on a GPU that they're out 18 grand plus on a GPU that they're just going to use to run shitty models just going to use to run shitty models just going to use to run shitty models locally. And that's why they made the locally. And that's why they made the locally. And that's why they made the DGX Spark. This small little box costs DGX Spark. This small little box costs DGX Spark. This small little box costs four grand and I hope you don't plan to four grand and I hope you don't plan to four grand and I hope you don't plan to use it as a computer. That is what it is use it as a computer. That is what it is use it as a computer. That is what it is to be clear. It's a computer. It shows to be clear. It's a computer. It shows to be clear. It's a computer. It shows up running a botched install of Ubuntu up running a botched install of Ubuntu up running a botched install of Ubuntu that they filled with [ __ ] and it that they filled with [ __ ] and it that they filled with [ __ ] and it runs a 20 core arm chip. And having runs a 20 core arm chip. And having runs a 20 core arm chip. And having spent far far far too much time fighting spent far far far too much time fighting spent far far far too much time fighting arm Linux in my life, I promise you arm Linux in my life, I promise you arm Linux in my life, I promise you you're not going to use this computer as you're not going to use this computer as you're not going to use this computer as a computer. Linux on arm is hell. Don't a computer. Linux on arm is hell. Don't a computer. Linux on arm is hell. Don't bother for this for anything other than bother for this for anything other than bother for this for anything other than inference. So this arm chip in this inference. So this arm chip in this inference. So this arm chip in this random $4,000 mini PC that you can't use random $4,000 mini PC that you can't use random $4,000 mini PC that you can't use as a computer does have a benefit. 128 as a computer does have a benefit. 128 as a computer does have a benefit. 128 gigs of unified memory. That memory is gigs of unified memory. That memory is gigs of unified memory. That memory is LPDDR5 memory though, which says a LPDDR5 memory though, which says a LPDDR5 memory though, which says a different number here. There's a lot of different number here. There's a lot of different number here. There's a lot of layers to how those numbers are layers to how those numbers are layers to how those numbers are measured. LPDDR5 is meaningfully slower measured. LPDDR5 is meaningfully slower measured. LPDDR5 is meaningfully slower than GDDR7. So the RAM is slower, than GDDR7. So the RAM is slower, than GDDR7. So the RAM is slower, but you have way more of it. But most but you have way more of it. But most but you have way more of it. But most importantly, you have CUDA cores, cores importantly, you have CUDA cores, cores importantly, you have CUDA cores, cores that can be used with CUDA-backed stuff that can be used with CUDA-backed stuff that can be used with CUDA-backed stuff that are actually capable of running that are actually capable of running that are actually capable of running real workflows. But you get 6,000 of real workflows. But you get 6,000 of real workflows. But you get 6,000 of them instead of the 24,000 plus you get them instead of the 24,000 plus you get them instead of the 24,000 plus you get on hardware like a 5090 or an RTX Pro.

  9. on hardware like a 5090 or an RTX Pro. on hardware like a 5090 or an RTX Pro. So, your options here are the world's So, your options here are the world's So, your options here are the world's worst computer environment possibly that worst computer environment possibly that worst computer environment possibly that you can buy today for real amounts of you can buy today for real amounts of you can buy today for real amounts of money. But you get a real amount of RAM money. But you get a real amount of RAM money. But you get a real amount of RAM you can use for models at the cost of an you can use for models at the cost of an you can use for models at the cost of an absolute [ __ ] garbage chip that runs absolute [ __ ] garbage chip that runs absolute [ __ ] garbage chip that runs terribly slow. Or you can get a really terribly slow. Or you can get a really terribly slow. Or you can get a really powerful, capable chip like an RTX 5090, powerful, capable chip like an RTX 5090, powerful, capable chip like an RTX 5090, and now you're entirely gimped on RAM. and now you're entirely gimped on RAM. and now you're entirely gimped on RAM. And if you want RAM and a high-end chip, And if you want RAM and a high-end chip, And if you want RAM and a high-end chip, you're paying the 16,000 plus dollars you're paying the 16,000 plus dollars you're paying the 16,000 plus dollars for that RTX Pro, sadly. Do you see what for that RTX Pro, sadly. Do you see what for that RTX Pro, sadly. Do you see what they've done here? they've done here? they've done here? They have the cheap option if you need They have the cheap option if you need They have the cheap option if you need more RAM. They have the cheap option if more RAM. They have the cheap option if more RAM. They have the cheap option if you want a fast chip. But you're paying you want a fast chip. But you're paying you want a fast chip. But you're paying up to 10 times plus more if you need up to 10 times plus more if you need up to 10 times plus more if you need both. They did do one really nice thing both. They did do one really nice thing both. They did do one really nice thing with the Sparks. They gave them a 200 with the Sparks. They gave them a 200 with the Sparks. They gave them a 200 gigabit per second NIC so that you can gigabit per second NIC so that you can gigabit per second NIC so that you can connect them over SFP. So, you can have connect them over SFP. So, you can have connect them over SFP. So, you can have the different Sparks transfer data the different Sparks transfer data the different Sparks transfer data between each other relatively fast. It's between each other relatively fast. It's between each other relatively fast. It's hell to set up, but you can at least do hell to set up, but you can at least do hell to set up, but you can at least do it. So, we've addressed all the fun it. So, we've addressed all the fun it. So, we've addressed all the fun things here for where Nvidia's at and things here for where Nvidia's at and things here for where Nvidia's at and how they can charge these absurd prices.

  10. how they can charge these absurd prices. how they can charge these absurd prices. What we haven't yet addressed is why I'm What we haven't yet addressed is why I'm What we haven't yet addressed is why I'm filming this today. There's a couple of filming this today. There's a couple of filming this today. There's a couple of things that are inspiring me to take things that are inspiring me to take things that are inspiring me to take this video on. The first is an anonymous this video on. The first is an anonymous this video on. The first is an anonymous model that dropped last week called Aux model that dropped last week called Aux model that dropped last week called Aux Alpha. I already have a video about Alpha. I already have a video about Alpha. I already have a video about this. It's probably already out on the this. It's probably already out on the this. It's probably already out on the channel, by the way. You should check channel, by the way. You should check channel, by the way. You should check that out. And when you're on the way to that out. And when you're on the way to that out. And when you're on the way to it, you should hit that subscribe button it, you should hit that subscribe button it, you should hit that subscribe button cuz it costs you nothing and with these cuz it costs you nothing and with these cuz it costs you nothing and with these videos take a lot of work. So, Aux Alpha videos take a lot of work. So, Aux Alpha videos take a lot of work. So, Aux Alpha was an anonymous model that came out on was an anonymous model that came out on was an anonymous model that came out on open router and open code, both of which open router and open code, both of which open router and open code, both of which had an absurd amount of free throughput had an absurd amount of free throughput had an absurd amount of free throughput on them. I believe it was 100 trillion on them. I believe it was 100 trillion on them. I believe it was 100 trillion tokens a day for free. And the model was tokens a day for free. And the model was tokens a day for free. And the model was great. A bunch of people were using it. great. A bunch of people were using it. great. A bunch of people were using it. I was using it a bunch. I have a video I was using it a bunch. I have a video I was using it a bunch. I have a video about how much I love it. Great model. about how much I love it. Great model. about how much I love it. Great model. Turned out to be GLM 53 flash. And the Turned out to be GLM 53 flash. And the Turned out to be GLM 53 flash. And the reason they could serve it so reason they could serve it so reason they could serve it so aggressively is both because it's a aggressively is both because it's a aggressively is both because it's a relatively small model, but also because relatively small model, but also because relatively small model, but also because they served it entirely on Chinese they served it entirely on Chinese they served it entirely on Chinese chips. Huawei made the chips that they chips. Huawei made the chips that they chips. Huawei made the chips that they served all this traffic on, and it's served all this traffic on, and it's served all this traffic on, and it's still working great. It's genuinely still working great. It's genuinely still working great. It's genuinely impressive that without any Nvidia in impressive that without any Nvidia in impressive that without any Nvidia in their stack, they were able to do this their stack, they were able to do this their stack, they were able to do this absurd level of traffic. I also got absurd level of traffic. I also got absurd level of traffic. I also got called out here because people were called out here because people were called out here because people were saying only Frontier Labs have this much saying only Frontier Labs have this much saying only Frontier Labs have this much compute. When I said that, I assumed compute. When I said that, I assumed compute. When I said that, I assumed that they were still serving Nvidia that they were still serving Nvidia that they were still serving Nvidia because we basically expected that. And because we basically expected that. And because we basically expected that. And also this wasn't a small model. To be also this wasn't a small model. To be also this wasn't a small model. To be fair, when I tweeted that, my secret fair, when I tweeted that, my secret fair, when I tweeted that, my secret personal belief was that Xiaomi was personal belief was that Xiaomi was personal belief was that Xiaomi was doing this on Chinese chips somehow.

  11. doing this on Chinese chips somehow. doing this on Chinese chips somehow. Then it turned out to be GLM shipping Then it turned out to be GLM shipping Then it turned out to be GLM shipping again in like a 2-week window, also on again in like a 2-week window, also on again in like a 2-week window, also on those same chips. All of the traffic was those same chips. All of the traffic was those same chips. All of the traffic was served on Chinese chips attaining served on Chinese chips attaining served on Chinese chips attaining hardware efficiency and per token cost hardware efficiency and per token cost hardware efficiency and per token cost comparable to Nvidia GPUs. The CUDA mode comparable to Nvidia GPUs. The CUDA mode comparable to Nvidia GPUs. The CUDA mode is being tested once again after is being tested once again after is being tested once again after Jalapeno's announcement yesterday. Jalapeno's announcement yesterday. Jalapeno's announcement yesterday. Jalapeno's Open AI chip that they've Jalapeno's Open AI chip that they've Jalapeno's Open AI chip that they've been working on in order to get out of been working on in order to get out of been working on in order to get out of the hell of relying on Nvidia so much. the hell of relying on Nvidia so much. the hell of relying on Nvidia so much. They are partnering with Cerebrus, who's They are partnering with Cerebrus, who's They are partnering with Cerebrus, who's a company that makes faster inference a company that makes faster inference a company that makes faster inference chips and host models that way in order chips and host models that way in order chips and host models that way in order to everything they can to massively to everything they can to massively to everything they can to massively increase speeds and reduce costs. But increase speeds and reduce costs. But increase speeds and reduce costs. But that is not the only thing that was just that is not the only thing that was just that is not the only thing that was just announced to scare Nvidia. Apple, out of announced to scare Nvidia. Apple, out of announced to scare Nvidia. Apple, out of nowhere, announced upgrades to the Mac nowhere, announced upgrades to the Mac nowhere, announced upgrades to the Mac mini, but more importantly, the Mac mini, but more importantly, the Mac mini, but more importantly, the Mac Studio, which has not seen upgrade since Studio, which has not seen upgrade since Studio, which has not seen upgrade since the M3 era. And the M4 did kind of come the M3 era. And the M4 did kind of come the M3 era. And the M4 did kind of come out, but they never made an M4 Ultra, out, but they never made an M4 Ultra, out, but they never made an M4 Ultra, they only made an M4 Max and it was they only made an M4 Max and it was they only made an M4 Max and it was gimped on how much RAM it could have. So gimped on how much RAM it could have. So gimped on how much RAM it could have. So you were stuck with like a 3 and a half you were stuck with like a 3 and a half you were stuck with like a 3 and a half plus year old chip if you wanted a lot plus year old chip if you wanted a lot plus year old chip if you wanted a lot of RAM on a Mac Studio. Obnoxious, dumb, of RAM on a Mac Studio. Obnoxious, dumb, of RAM on a Mac Studio. Obnoxious, dumb, solved. The Mac Studio now has M5 Max solved. The Mac Studio now has M5 Max solved. The Mac Studio now has M5 Max and Ultra.

  12. and Ultra. and Ultra. The first section here is for the Max, The first section here is for the Max, The first section here is for the Max, which you have 128 gigs of RAM and 614 which you have 128 gigs of RAM and 614 which you have 128 gigs of RAM and 614 gigabytes per second of memory gigabytes per second of memory gigabytes per second of memory bandwidth. But the Ultra is basically bandwidth. But the Ultra is basically bandwidth. But the Ultra is basically just two of these chips stapled to each just two of these chips stapled to each just two of these chips stapled to each other, which means it can do 36 cores of other, which means it can do 36 cores of other, which means it can do 36 cores of CPU, 80 cores of GPU, up to 512 gigs of CPU, 80 cores of GPU, up to 512 gigs of CPU, 80 cores of GPU, up to 512 gigs of memory, 1.2 terabytes a second of memory, 1.2 terabytes a second of memory, 1.2 terabytes a second of bandwidth, and a 32 core neural engine. bandwidth, and a 32 core neural engine. bandwidth, and a 32 core neural engine. As a video nerd, I love the fact that As a video nerd, I love the fact that As a video nerd, I love the fact that the new Ultra chip can do 33 streams of the new Ultra chip can do 33 streams of the new Ultra chip can do 33 streams of 8K ProRes 422 at 30 FPS playback. Like 8K ProRes 422 at 30 FPS playback. Like 8K ProRes 422 at 30 FPS playback. Like that's insane. But that's not what we're that's insane. But that's not what we're that's insane. But that's not what we're here for, let's be real. here for, let's be real. here for, let's be real. We're here because the time to first We're here because the time to first We're here because the time to first token on a Mac Studio with M5 Ultra is token on a Mac Studio with M5 Ultra is token on a Mac Studio with M5 Ultra is 10 times faster than it was on a Mac 10 times faster than it was on a Mac 10 times faster than it was on a Mac Studio with M1 Ultra. That is a massive Studio with M1 Ultra. That is a massive Studio with M1 Ultra. That is a massive increase in performance and this largely increase in performance and this largely increase in performance and this largely comes down to the improvements in the comes down to the improvements in the comes down to the improvements in the memory. But the important thing to note memory. But the important thing to note memory. But the important thing to note here with the memory isn't even just the here with the memory isn't even just the here with the memory isn't even just the amount, it's this word here, unified. amount, it's this word here, unified. amount, it's this word here, unified. Because that means it works as normal Because that means it works as normal Because that means it works as normal RAM, but also as VRAM, which means the RAM, but also as VRAM, which means the RAM, but also as VRAM, which means the GPUs can use it for inference. When you GPUs can use it for inference. When you GPUs can use it for inference. When you look at the Mac Studio compared to look at the Mac Studio compared to look at the Mac Studio compared to Nvidia's options, you see just how Nvidia's options, you see just how Nvidia's options, you see just how compelling it gets. You can get a 590 compelling it gets. You can get a 590 compelling it gets. You can get a 590 which has crazy GPU performance, which which has crazy GPU performance, which which has crazy GPU performance, which isn't indicated here how fast the isn't indicated here how fast the isn't indicated here how fast the computer is on it. You only get 32 gigs computer is on it. You only get 32 gigs computer is on it. You only get 32 gigs of RAM though at the benefit of the way of RAM though at the benefit of the way of RAM though at the benefit of the way Nvidia implements it, effective 1800 GB Nvidia implements it, effective 1800 GB Nvidia implements it, effective 1800 GB per second memory bandwidth. The RTX per second memory bandwidth. The RTX per second memory bandwidth. The RTX Pro, similar bandwidth, 3x the RAM, much

  13. Pro, similar bandwidth, 3x the RAM, much Pro, similar bandwidth, 3x the RAM, much more than 3x the price. The DGX Spark, more than 3x the price. The DGX Spark, more than 3x the price. The DGX Spark, way more RAM up to 128 gigs, but it's way more RAM up to 128 gigs, but it's way more RAM up to 128 gigs, but it's unified LPDDR5, that is a seventh or so unified LPDDR5, that is a seventh or so unified LPDDR5, that is a seventh or so the speed, but suddenly we get a thing the speed, but suddenly we get a thing the speed, but suddenly we get a thing with no compromises. No compromise on with no compromises. No compromise on with no compromises. No compromise on the amount of RAM, no compromise on the the amount of RAM, no compromise on the the amount of RAM, no compromise on the memory bandwidth, and depending on how memory bandwidth, and depending on how memory bandwidth, and depending on how it performs when it comes out, not as it performs when it comes out, not as it performs when it comes out, not as big of a compromise on the compute big of a compromise on the compute big of a compromise on the compute itself. So for 10 grand MSRP and the itself. So for 10 grand MSRP and the itself. So for 10 grand MSRP and the current price still that cuz it isn't current price still that cuz it isn't current price still that cuz it isn't coming out soon, you get way more RAM coming out soon, you get way more RAM coming out soon, you get way more RAM than any other option as offered by than any other option as offered by than any other option as offered by Nvidia, a chip that is meaningfully Nvidia, a chip that is meaningfully Nvidia, a chip that is meaningfully better than the DGX Spark, comically so better than the DGX Spark, comically so better than the DGX Spark, comically so even. Not necessarily better than the even. Not necessarily better than the even. Not necessarily better than the GPUs in the 5090 and the 6000 Blackwell. GPUs in the 5090 and the 6000 Blackwell. GPUs in the 5090 and the 6000 Blackwell. We don't know yet until it's actually We don't know yet until it's actually We don't know yet until it's actually out, but I'm guessing almost certainly out, but I'm guessing almost certainly out, but I'm guessing almost certainly not. But most importantly, you get 256 not. But most importantly, you get 256 not. But most importantly, you get 256 gigs of their unified memory with gigs of their unified memory with gigs of their unified memory with bandwidth a hell of a lot closer to what bandwidth a hell of a lot closer to what bandwidth a hell of a lot closer to what Nvidia sees on their GPUs than to what Nvidia sees on their GPUs than to what Nvidia sees on their GPUs than to what that you would get on a Spark. Almost that you would get on a Spark. Almost that you would get on a Spark. Almost six times faster than the Spark and only six times faster than the Spark and only six times faster than the Spark and only like 30 to 40% slower than the 5090.

  14. like 30 to 40% slower than the 5090. like 30 to 40% slower than the 5090. That is insane. This kills almost all of That is insane. This kills almost all of That is insane. This kills almost all of the reasons you could ever justify the reasons you could ever justify the reasons you could ever justify buying the DGX Spark, which is a whole buying the DGX Spark, which is a whole buying the DGX Spark, which is a whole category of Nvidia devices killed. And category of Nvidia devices killed. And category of Nvidia devices killed. And all the people who upgraded from a Spark all the people who upgraded from a Spark all the people who upgraded from a Spark to a RTX Pro 6000 or skipped the Spark to a RTX Pro 6000 or skipped the Spark to a RTX Pro 6000 or skipped the Spark because they wanted speeds that weren't because they wanted speeds that weren't because they wanted speeds that weren't trash, they can now get way, way better trash, they can now get way, way better trash, they can now get way, way better experiences for meaningfully cheaper experiences for meaningfully cheaper experiences for meaningfully cheaper by just getting the M5 Ultra instead. by just getting the M5 Ultra instead. by just getting the M5 Ultra instead. That's crazy. This is Apple's first real That's crazy. This is Apple's first real That's crazy. This is Apple's first real play in the AI space. Them coming in and play in the AI space. Them coming in and play in the AI space. Them coming in and saying, "Sorry guys, you're [ __ ] saying, "Sorry guys, you're [ __ ] saying, "Sorry guys, you're [ __ ] around too much. We're going to put an around too much. We're going to put an around too much. We're going to put an end to that." And yes, the DGX Spark's end to that." And yes, the DGX Spark's end to that." And yes, the DGX Spark's effective memory bandwidth is actually effective memory bandwidth is actually effective memory bandwidth is actually this pathetic. It's a [ __ ] chip. It's a this pathetic. It's a [ __ ] chip. It's a this pathetic. It's a [ __ ] chip. It's a [ __ ] system. The DGX Spark is my least [ __ ] system. The DGX Spark is my least [ __ ] system. The DGX Spark is my least favorite computer in this apartment. favorite computer in this apartment. favorite computer in this apartment. What about Jalapeno? Well, according to What about Jalapeno? Well, according to What about Jalapeno? Well, according to SemiAnalysis, it is coming out to be SemiAnalysis, it is coming out to be SemiAnalysis, it is coming out to be better than Blackwell. Their better than Blackwell. Their better than Blackwell. Their self-designed ASIC comparing with Rubin, self-designed ASIC comparing with Rubin, self-designed ASIC comparing with Rubin, which is Jalapeno's TCO through per which is Jalapeno's TCO through per which is Jalapeno's TCO through per megawatt and spicy deeds. Cool. Open AI megawatt and spicy deeds. Cool. Open AI megawatt and spicy deeds. Cool. Open AI actually invited SemiAnalysis come take actually invited SemiAnalysis come take actually invited SemiAnalysis come take a look early. In general, first energy a look early. In general, first energy a look early. In general, first energy and trips are not competitive, but Open and trips are not competitive, but Open and trips are not competitive, but Open AI bucks that trend by being industry AI bucks that trend by being industry AI bucks that trend by being industry leading and beating every Nvidia, AMD, leading and beating every Nvidia, AMD, leading and beating every Nvidia, AMD, and Google chip we have been able to and Google chip we have been able to and Google chip we have been able to test on multiple top open source models.

  15. test on multiple top open source models. test on multiple top open source models. Open AI does this with extreme hardware Open AI does this with extreme hardware Open AI does this with extreme hardware software code design. Traditionally, software code design. Traditionally, software code design. Traditionally, these bespoke chips tend to super these bespoke chips tend to super these bespoke chips tend to super fixated specialize on specific things. fixated specialize on specific things. fixated specialize on specific things. They have not actually done that. They They have not actually done that. They They have not actually done that. They built a really good general chip that built a really good general chip that built a really good general chip that delivers high performance in all delivers high performance in all delivers high performance in all scenarios. Apparently, their chip goes scenarios. Apparently, their chip goes scenarios. Apparently, their chip goes as high as 216 GB of HBM, they only as high as 216 GB of HBM, they only as high as 216 GB of HBM, they only require 700 W, and it's doing 13.4 require 700 W, and it's doing 13.4 require 700 W, and it's doing 13.4 petaflops for FP4. And yeah, those are petaflops for FP4. And yeah, those are petaflops for FP4. And yeah, those are insane numbers. Apparently, it doesn't insane numbers. Apparently, it doesn't insane numbers. Apparently, it doesn't have FP16 numbers, but for FP8, it is have FP16 numbers, but for FP8, it is have FP16 numbers, but for FP8, it is slightly behind what you're seeing on slightly behind what you're seeing on slightly behind what you're seeing on GB2000 and 3000s from Nvidia, but it's GB2000 and 3000s from Nvidia, but it's GB2000 and 3000s from Nvidia, but it's meaningfully ahead of the H100 and 200 meaningfully ahead of the H100 and 200 meaningfully ahead of the H100 and 200 already, which is pretty crazy. And already, which is pretty crazy. And already, which is pretty crazy. And also, more RAM and way higher bandwidth. also, more RAM and way higher bandwidth. also, more RAM and way higher bandwidth. Even just this chart should be enough to Even just this chart should be enough to Even just this chart should be enough to give Nvidia a heart attack. It's still give Nvidia a heart attack. It's still give Nvidia a heart attack. It's still not quite Reuben levels, but it's not quite Reuben levels, but it's not quite Reuben levels, but it's trading blows at a way lower wattage. A trading blows at a way lower wattage. A trading blows at a way lower wattage. A lot of the media coverage of this chip lot of the media coverage of this chip lot of the media coverage of this chip has followed a few throwaway comments has followed a few throwaway comments has followed a few throwaway comments from OpenAI that claim the chip will be from OpenAI that claim the chip will be from OpenAI that claim the chip will be optimized for their models in a way that optimized for their models in a way that optimized for their models in a way that other chips aren't. This is wrong. other chips aren't. This is wrong. other chips aren't. This is wrong. Jalapeno is a generalized inference chip Jalapeno is a generalized inference chip Jalapeno is a generalized inference chip capable of running all sorts of models capable of running all sorts of models capable of running all sorts of models in all sorts of workloads, including our in all sorts of workloads, including our in all sorts of workloads, including our benchmark inference X, where we ran the benchmark inference X, where we ran the benchmark inference X, where we ran the benchmark with OpenAI engineers in the benchmark with OpenAI engineers in the benchmark with OpenAI engineers in the lab. As a joke, OpenAI even showed us it lab. As a joke, OpenAI even showed us it lab. As a joke, OpenAI even showed us it running Doom, which was ported to their running Doom, which was ported to their running Doom, which was ported to their chip with just Codex prompts. Of course, chip with just Codex prompts. Of course, chip with just Codex prompts. Of course, they have it running Doom. The following they have it running Doom. The following they have it running Doom. The following is the headline performance per watt is the headline performance per watt is the headline performance per watt result. So, let's see what the numbers result. So, let's see what the numbers result. So, let's see what the numbers looked like. It's looking pretty insane looked like. It's looking pretty insane looked like. It's looking pretty insane token per second per megawatt token per second per megawatt token per second per megawatt performance here. They are crushing performance here. They are crushing performance here. They are crushing everything else in efficiency in terms everything else in efficiency in terms everything else in efficiency in terms of electricity. Jalapeno is beating

  16. of electricity. Jalapeno is beating of electricity. Jalapeno is beating Black Hole on performance per watt Black Hole on performance per watt Black Hole on performance per watt across almost all scenarios without across almost all scenarios without across almost all scenarios without being tuned for any specific point in being tuned for any specific point in being tuned for any specific point in the curve. It excels not only at low the curve. It excels not only at low the curve. It excels not only at low latency scenarios, but also in high latency scenarios, but also in high latency scenarios, but also in high throughput scenarios. A more throughput scenarios. A more throughput scenarios. A more apples-to-apples comparison is against apples-to-apples comparison is against apples-to-apples comparison is against single token prediction results. It single token prediction results. It single token prediction results. It knocks every competitor out of the knocks every competitor out of the knocks every competitor out of the water. At low concurrency scenarios, water. At low concurrency scenarios, water. At low concurrency scenarios, Jalapeno demonstrates remarkable Jalapeno demonstrates remarkable Jalapeno demonstrates remarkable interactivity, hitting over 700 TPS per interactivity, hitting over 700 TPS per interactivity, hitting over 700 TPS per user at concurrency one on the DeepSeek user at concurrency one on the DeepSeek user at concurrency one on the DeepSeek R1 model. This is all being achieved R1 model. This is all being achieved R1 model. This is all being achieved with single token prediction, no with single token prediction, no with single token prediction, no speculative decoding, and no prefilled speculative decoding, and no prefilled speculative decoding, and no prefilled decode disaggregation. They got Kim 25 decode disaggregation. They got Kim 25 decode disaggregation. They got Kim 25 and GPT-OSS running at 1400 tokens per and GPT-OSS running at 1400 tokens per and GPT-OSS running at 1400 tokens per second per user. And they confirmed that second per user. And they confirmed that second per user. And they confirmed that the GSM-8KE valves attained results on the GSM-8KE valves attained results on the GSM-8KE valves attained results on par with Nvidia chips, so So, not par with Nvidia chips, so So, not par with Nvidia chips, so So, not nerfing the models when they run them. nerfing the models when they run them. nerfing the models when they run them. Since they're using HBM4, which is the Since they're using HBM4, which is the Since they're using HBM4, which is the new generation of high-bandwidth memory, new generation of high-bandwidth memory, new generation of high-bandwidth memory, they get a pretty substantial win over they get a pretty substantial win over they get a pretty substantial win over things on older memory. Rubin is the new things on older memory. Rubin is the new things on older memory. Rubin is the new Nvidia line that will use HBM4, but Nvidia line that will use HBM4, but Nvidia line that will use HBM4, but those chips are not really actually out those chips are not really actually out those chips are not really actually out yet. Blackwell is the line that most yet. Blackwell is the line that most yet. Blackwell is the line that most things are buying and using. There are things are buying and using. There are things are buying and using. There are people who have put in huge orders for people who have put in huge orders for people who have put in huge orders for Blackwell chips that will finally show Blackwell chips that will finally show Blackwell chips that will finally show up in like 2 to 3 years. It's a while up in like 2 to 3 years. It's a while up in like 2 to 3 years. It's a while before Rubin's going to matter. There's before Rubin's going to matter. There's before Rubin's going to matter. There's also the callout that this is not large also the callout that this is not large also the callout that this is not large model performance. All the models model performance. All the models model performance. All the models they've tested so far are relatively they've tested so far are relatively they've tested so far are relatively easy to run in terms of size. So, we easy to run in terms of size. So, we easy to run in terms of size. So, we don't know how these will perform when don't know how these will perform when don't know how these will perform when you give them huge models like Kimmy K3 you give them huge models like Kimmy K3 you give them huge models like Kimmy K3 or Deep Seek before Pro. OpenAI is or Deep Seek before Pro. OpenAI is or Deep Seek before Pro. OpenAI is designing for performance per watt. The designing for performance per watt. The designing for performance per watt. The reason is simple. OpenAI is currently reason is simple. OpenAI is currently reason is simple. OpenAI is currently limited by data center power, not by limited by data center power, not by limited by data center power, not by budget or floor space, and thus tokens budget or floor space, and thus tokens budget or floor space, and thus tokens per megawatt is paramount. At Computex per megawatt is paramount. At Computex per megawatt is paramount. At Computex 2026, Jensen said the performance per

  17. 2026, Jensen said the performance per 2026, Jensen said the performance per watt, reliability, and long lifetimes watt, reliability, and long lifetimes watt, reliability, and long lifetimes are the core features of future GPUs. To are the core features of future GPUs. To are the core features of future GPUs. To quote, "If you have 1 gigawatt of power, quote, "If you have 1 gigawatt of power, quote, "If you have 1 gigawatt of power, then throughput per watt is revenue." then throughput per watt is revenue." then throughput per watt is revenue." Yep. I've been yelling this for a while. Yep. I've been yelling this for a while. Yep. I've been yelling this for a while. Electricity's going to be a big deal. Electricity's going to be a big deal. Electricity's going to be a big deal. So, OpenAI focusing on that side is a So, OpenAI focusing on that side is a So, OpenAI focusing on that side is a huge deal. Don't worry though, Jensen's huge deal. Don't worry though, Jensen's huge deal. Don't worry though, Jensen's definitely not scared. Nvidia has a definitely not scared. Nvidia has a definitely not scared. Nvidia has a plan. They're going to be totally fine. plan. They're going to be totally fine. plan. They're going to be totally fine. Not like he's going to say something Not like he's going to say something Not like he's going to say something stupid like they're just going to cancel stupid like they're just going to cancel stupid like they're just going to cancel the development of the thing that the development of the thing that the development of the thing that destroys our business model. I haven't destroys our business model. I haven't destroys our business model. I haven't actually seen this clip, so we get to actually seen this clip, so we get to actually seen this clip, so we get to watch it together. watch it together. watch it together. >> And I said, "You know what? I The whole >> And I said, "You know what? I The whole >> And I said, "You know what? I The whole time I've been working on on something, time I've been working on on something, time I've been working on on something, a chip that I think could really hurt a chip that I think could really hurt a chip that I think could really hurt you. It's called jalapeno. I don't know you. It's called jalapeno. I don't know you. It's called jalapeno. I don't know what let's call it I don't know what what let's call it I don't know what what let's call it I don't know what something to call something to call something to call and you said to me, "Well, that's fine." and you said to me, "Well, that's fine." and you said to me, "Well, that's fine." Would you really say that's fine cuz I Would you really say that's fine cuz I Would you really say that's fine cuz I personally would be hurt if OpenAI did personally would be hurt if OpenAI did personally would be hurt if OpenAI did to what they do to Nvidia if they did to to what they do to Nvidia if they did to to what they do to Nvidia if they did to me. me. me. >> You know, I'm okay with it, Jim. There's >> You know, I'm okay with it, Jim. There's >> You know, I'm okay with it, Jim. There's so many XPUs that are being announced, so many XPUs that are being announced, so many XPUs that are being announced, and as we know, it's not easy doing what and as we know, it's not easy doing what and as we know, it's not easy doing what we do. We've been doing this for 33 we do. We've been doing this for 33 we do. We've been doing this for 33 years.

  18. years. years. And so, lots of projects get started, a And so, lots of projects get started, a And so, lots of projects get started, a lots of projects get gets canceled. Um lots of projects get gets canceled. Um lots of projects get gets canceled. Um we're we're here we're here to support we're we're here we're here to support we're we're here we're here to support our partners and and um, we're going to our partners and and um, we're going to our partners and and um, we're going to build the world's best technology. I build the world's best technology. I build the world's best technology. I have every confidence in that. Uh, we're have every confidence in that. Uh, we're have every confidence in that. Uh, we're going to be the most productive going to be the most productive going to be the most productive infrastructure that they have. I have infrastructure that they have. I have infrastructure that they have. I have every confidence in that. Um, we have every confidence in that. Um, we have every confidence in that. Um, we have the supply chain and the technology the supply chain and the technology the supply chain and the technology scale to be their largest supplier. I scale to be their largest supplier. I scale to be their largest supplier. I have every confidence in that. And so, have every confidence in that. And so, have every confidence in that. And so, you know, I I don't I don't have to take you know, I I don't I don't have to take you know, I I don't I don't have to take anything personally because I've got so anything personally because I've got so anything personally because I've got so much confidence in what we we're able to much confidence in what we we're able to much confidence in what we we're able to deliver. And look at look at all of the deliver. And look at look at all of the deliver. And look at look at all of the XPU announcements and all the startups XPU announcements and all the startups XPU announcements and all the startups that have been announced and yet today that have been announced and yet today that have been announced and yet today Nvidia is increasing our market share of Nvidia is increasing our market share of Nvidia is increasing our market share of the AI market. the AI market. the AI market. We're our growth is accelerating. Our We're our growth is accelerating. Our We're our growth is accelerating. Our technology leadership is extending. And technology leadership is extending. And technology leadership is extending. And so, I'm very comfortable with all the so, I'm very comfortable with all the so, I'm very comfortable with all the competition. competition. competition. >> Okay, and I also >> Okay, and I also >> Okay, and I also >> Yeah. >> Yeah. >> Yeah. Definitely not scared at all, are you, Definitely not scared at all, are you, Definitely not scared at all, are you, bud? Well, on the bright side, if they bud? Well, on the bright side, if they bud? Well, on the bright side, if they own Hugging Face, they can make sure own Hugging Face, they can make sure own Hugging Face, they can make sure they suppress access to all of the they suppress access to all of the they suppress access to all of the versions of the models that run well on versions of the models that run well on versions of the models that run well on things that aren't CUDA. I'll say that I things that aren't CUDA. I'll say that I things that aren't CUDA. I'll say that I think the Hugging Face bid is an actual think the Hugging Face bid is an actual think the Hugging Face bid is an actual good faith play to try and bolster and good faith play to try and bolster and good faith play to try and bolster and fund the development of open source AI fund the development of open source AI fund the development of open source AI and encouraging more and more people to and encouraging more and more people to and encouraging more and more people to train because let's be real, training is train because let's be real, training is train because let's be real, training is still happening on CUDA. The more they still happening on CUDA. The more they still happening on CUDA. The more they encourage businesses to try and train encourage businesses to try and train encourage businesses to try and train and fine-tune and customize things and fine-tune and customize things and fine-tune and customize things themselves, the more customers they have themselves, the more customers they have themselves, the more customers they have for their chips, the longer-term their for their chips, the longer-term their for their chips, the longer-term their absurd saturation can go for. And absurd saturation can go for. And absurd saturation can go for. And Nvidia's still making a ton of money.

  19. Nvidia's still making a ton of money. Nvidia's still making a ton of money. They can justify doing [ __ ] like this. They can justify doing [ __ ] like this. They can justify doing [ __ ] like this. Good for them. If you want to spend less Good for them. If you want to spend less Good for them. If you want to spend less than $12.9 billion to do something that than $12.9 billion to do something that than $12.9 billion to do something that positions your business better, positions your business better, positions your business better, you have access to my DMs, Jensen. you have access to my DMs, Jensen. you have access to my DMs, Jensen. Somebody in chat mentioned, "You know Somebody in chat mentioned, "You know Somebody in chat mentioned, "You know what's funny, Theo? I bet SpaceX AI is what's funny, Theo? I bet SpaceX AI is what's funny, Theo? I bet SpaceX AI is also doing something similar now." Which also doing something similar now." Which also doing something similar now." Which reminded me, somehow I entirely forgot reminded me, somehow I entirely forgot reminded me, somehow I entirely forgot about Terafab, the most epic chip about Terafab, the most epic chip about Terafab, the most epic chip building effort ever, which is, by the building effort ever, which is, by the building effort ever, which is, by the way, the only thing in Elon Musk's bio way, the only thing in Elon Musk's bio way, the only thing in Elon Musk's bio right now, despite the fact that him and right now, despite the fact that him and right now, despite the fact that him and Jensen are buddy-buddy and Elon has some Jensen are buddy-buddy and Elon has some Jensen are buddy-buddy and Elon has some of the biggest and most lucrative of the biggest and most lucrative of the biggest and most lucrative contracts with Nvidia. He has more GPUs contracts with Nvidia. He has more GPUs contracts with Nvidia. He has more GPUs than Anthropic does. Anthropic is than Anthropic does. Anthropic is than Anthropic does. Anthropic is renting GPUs from Elon now because he renting GPUs from Elon now because he renting GPUs from Elon now because he was so quick on this. And he is still was so quick on this. And he is still was so quick on this. And he is still concerned about Nvidia's monopoly and concerned about Nvidia's monopoly and concerned about Nvidia's monopoly and just trying to get out of it. Terafab just trying to get out of it. Terafab just trying to get out of it. Terafab will close the gap between today's chip will close the gap between today's chip will close the gap between today's chip production and the future's demand, a production and the future's demand, a production and the future's demand, a future among the stars. And we do this future among the stars. And we do this future among the stars. And we do this by building a gigantic chip fab. The by building a gigantic chip fab. The by building a gigantic chip fab. The Terafab will be comically bigger than Terafab will be comically bigger than Terafab will be comically bigger than even giant things like the Giga Texas even giant things like the Giga Texas even giant things like the Giga Texas fab for Tesla, the US Pentagon, Mall of fab for Tesla, the US Pentagon, Mall of fab for Tesla, the US Pentagon, Mall of America, and more. It's 25 times the America, and more. It's 25 times the America, and more. It's 25 times the size of Apple Park. It's size of Apple Park. It's size of Apple Park. It's 20 times the size of the Pentagon. Kind 20 times the size of the Pentagon. Kind 20 times the size of the Pentagon. Kind of crazy. So, yeah.

  20. of crazy. So, yeah. of crazy. So, yeah. I guess you could say Elon and SpaceX AI I guess you could say Elon and SpaceX AI I guess you could say Elon and SpaceX AI and Tesla or whatever whatever business and Tesla or whatever whatever business and Tesla or whatever whatever business he has doing this he has doing this he has doing this are considering competing with Nvidia. are considering competing with Nvidia. are considering competing with Nvidia. It's almost like literally everyone is. It's almost like literally everyone is. It's almost like literally everyone is. I should include a I should include a I should include a callout here. I am an AMD investor. I am callout here. I am an AMD investor. I am callout here. I am an AMD investor. I am not a special early investor. I'm just not a special early investor. I'm just not a special early investor. I'm just buying their stocks, but I have not buying their stocks, but I have not buying their stocks, but I have not talked about AMD at any point here cuz talked about AMD at any point here cuz talked about AMD at any point here cuz I'll be real. I love them to death. I'll be real. I love them to death. I'll be real. I love them to death. They're very behind. It'll be a while They're very behind. It'll be a while They're very behind. It'll be a while before they can catch up to what these before they can catch up to what these before they can catch up to what these other things are that I'm talking about other things are that I'm talking about other things are that I'm talking about here. I hope that changes, but for now here. I hope that changes, but for now here. I hope that changes, but for now AMD is a potential big winner, but they AMD is a potential big winner, but they AMD is a potential big winner, but they have some time to before they get there. have some time to before they get there. have some time to before they get there. I think that's all I have to say on the I think that's all I have to say on the I think that's all I have to say on the chaos that is the current state of chaos that is the current state of chaos that is the current state of chips. Apparently, Intel is involved in chips. Apparently, Intel is involved in chips. Apparently, Intel is involved in Terafab, too. Fun. That'll be an Terafab, too. Fun. That'll be an Terafab, too. Fun. That'll be an interesting project. I have no idea interesting project. I have no idea interesting project. I have no idea where any of this will go. All I know is where any of this will go. All I know is where any of this will go. All I know is that Nvidia is scared, and they have that Nvidia is scared, and they have that Nvidia is scared, and they have good reason to be. The AI world good reason to be. The AI world good reason to be. The AI world constantly is changing, and these tools constantly is changing, and these tools constantly is changing, and these tools and technologies make it easier than and technologies make it easier than and technologies make it easier than ever to catch up. We finally now have ever to catch up. We finally now have ever to catch up. We finally now have models that are useful enough to help models that are useful enough to help models that are useful enough to help these manufacturers in their process. I these manufacturers in their process. I these manufacturers in their process. I think this is a big part of why OpenAI think this is a big part of why OpenAI think this is a big part of why OpenAI is catching up as quickly as they are.

  21. is catching up as quickly as they are. is catching up as quickly as they are. Their use of AI in their catch-up Their use of AI in their catch-up Their use of AI in their catch-up process has enabled them to do it more process has enabled them to do it more process has enabled them to do it more effectively than anyone would have effectively than anyone would have effectively than anyone would have anticipated, including Nvidia. And now anticipated, including Nvidia. And now anticipated, including Nvidia. And now we're quickly approaching a future where we're quickly approaching a future where we're quickly approaching a future where the labs are able to compete with Nvidia the labs are able to compete with Nvidia the labs are able to compete with Nvidia directly instead of relying on them with directly instead of relying on them with directly instead of relying on them with these trillion dollar contracts to get these trillion dollar contracts to get these trillion dollar contracts to get all of the chips they need. As I've said all of the chips they need. As I've said all of the chips they need. As I've said many times now, the future is going to many times now, the future is going to many times now, the future is going to be fought not on chips, but on be fought not on chips, but on be fought not on chips, but on electricity. And if we don't have ways electricity. And if we don't have ways electricity. And if we don't have ways to get the energy we need, then none of to get the energy we need, then none of to get the energy we need, then none of this ends up mattering in the end. But this ends up mattering in the end. But this ends up mattering in the end. But at the very least for now, it is super at the very least for now, it is super at the very least for now, it is super interesting and Nvidia's weird position interesting and Nvidia's weird position interesting and Nvidia's weird position in the market might not last as long as in the market might not last as long as in the market might not last as long as they probably think. Am I crazy for they probably think. Am I crazy for they probably think. Am I crazy for saying all of this or am I kind of onto saying all of this or am I kind of onto saying all of this or am I kind of onto something? Let me know how you guys feel something? Let me know how you guys feel something? Let me know how you guys feel about the future for Nvidia and this about the future for Nvidia and this about the future for Nvidia and this whole space in general in the comments. whole space in general in the comments. whole space in general in the comments. Until next time, Until next time, Until next time, peace nerds.

Summary

The main theme is Nvidia's powerful but precarious monopoly in AI compute, driven by its CUDA system and deals with major AI players and governments. The practical takeaway is that this dependence creates vulnerabilities, leading companies like OpenAI and even China, due to US sanctions, to seek alternatives, potentially leading to Nvidia's monopoly being challenged.

View original episode ↗