US AI Dominance Is Over: Here's Why
Read full transcript 19 segments
-
It's something that everyone has on It's something that everyone has on their lips right now. We've had this their lips right now. We've had this their lips right now. We've had this moment since DeepSeek. It feels like moment since DeepSeek. It feels like moment since DeepSeek. It feels like China has been coming on. There's been a China has been coming on. There's been a China has been coming on. There's been a lot of narrative recently with the lot of narrative recently with the lot of narrative recently with the release of Kimi K3 and Qwen 3.8. release of Kimi K3 and Qwen 3.8. release of Kimi K3 and Qwen 3.8. Basically, everyone's asking, are these Basically, everyone's asking, are these Basically, everyone's asking, are these models caught up? Do we see the model models caught up? Do we see the model models caught up? Do we see the model frontier actually closing? I want to frontier actually closing? I want to frontier actually closing? I want to take this video and just talk about the take this video and just talk about the take this video and just talk about the Chinese model phenomenon, what it looks Chinese model phenomenon, what it looks Chinese model phenomenon, what it looks like when you're deep in testing and you like when you're deep in testing and you like when you're deep in testing and you test Chinese models, how they differ test Chinese models, how they differ test Chinese models, how they differ from US models, how you should approach from US models, how you should approach from US models, how you should approach thinking about Chinese models as an thinking about Chinese models as an thinking about Chinese models as an individual, maybe as a company. So, individual, maybe as a company. So, individual, maybe as a company. So, we're going to get into all of that we're going to get into all of that we're going to get into all of that detail in this video. I'm really excited detail in this video. I'm really excited detail in this video. I'm really excited for it. I don't think Chinese models get for it. I don't think Chinese models get for it. I don't think Chinese models get the attention they should get, and I the attention they should get, and I the attention they should get, and I wanted to make a dedicated video just to wanted to make a dedicated video just to wanted to make a dedicated video just to jump into the whole Chinese model world. jump into the whole Chinese model world. jump into the whole Chinese model world. Last week, Moonshot AI released Kimi K3, Last week, Moonshot AI released Kimi K3, Last week, Moonshot AI released Kimi K3, a Chinese model with 2.8 trillion a Chinese model with 2.8 trillion a Chinese model with 2.8 trillion parameters and a million token context parameters and a million token context parameters and a million token context window. Its API charges a fair bit, window. Its API charges a fair bit, window. Its API charges a fair bit, right? $15 per million output tokens. right? $15 per million output tokens. right? $15 per million output tokens. Meanwhile, DeepSeek V4 Pro charges just Meanwhile, DeepSeek V4 Pro charges just Meanwhile, DeepSeek V4 Pro charges just 87 cents for the same unit of volume.
-
87 cents for the same unit of volume. 87 cents for the same unit of volume. Moonshot also calls K3 open weight. Now, Moonshot also calls K3 open weight. Now, Moonshot also calls K3 open weight. Now, as I record this, the weights are not as I record this, the weights are not as I record this, the weights are not downloadable, but apparently, very, very downloadable, but apparently, very, very downloadable, but apparently, very, very soon, maybe by the time this video is soon, maybe by the time this video is soon, maybe by the time this video is out, you will have the full weights out, you will have the full weights out, you will have the full weights released. Both DeepSeek and Kimi and released. Both DeepSeek and Kimi and released. Both DeepSeek and Kimi and also Qwen are Chinese frontier models. also Qwen are Chinese frontier models. also Qwen are Chinese frontier models. They don't share a price, they don't They don't share a price, they don't They don't share a price, they don't share a deployment path or a hardware share a deployment path or a hardware share a deployment path or a hardware burden, or even a best use. They just burden, or even a best use. They just burden, or even a best use. They just share a country of origin. And that's share a country of origin. And that's share a country of origin. And that's one of the things I want to emphasize at one of the things I want to emphasize at one of the things I want to emphasize at the top here. We say Chinese models, and the top here. We say Chinese models, and the top here. We say Chinese models, and we make this big simplifying assumption. we make this big simplifying assumption. we make this big simplifying assumption. That's the problem with the Chinese That's the problem with the Chinese That's the problem with the Chinese model conversation. We use Chinese model conversation. We use Chinese model conversation. We use Chinese models sometimes as a shorthand for models sometimes as a shorthand for models sometimes as a shorthand for cheap, or for open, or for local. cheap, or for open, or for local. cheap, or for open, or for local. Sometimes all three are true, and Sometimes all three are true, and Sometimes all three are true, and sometimes none of them are. This video sometimes none of them are. This video sometimes none of them are. This video is going to dive deeper. It's going to is going to dive deeper. It's going to is going to dive deeper. It's going to give you a practical way to choose give you a practical way to choose give you a practical way to choose whether to include a Chinese model in whether to include a Chinese model in whether to include a Chinese model in your stack. Which work should you give your stack. Which work should you give your stack. Which work should you give them? When do you use an API? When is them? When do you use an API? When is them? When do you use an API? When is self-hosting justified? I'm going to get self-hosting justified? I'm going to get self-hosting justified? I'm going to get into the distillation fight because it into the distillation fight because it into the distillation fight because it could change how quickly capability could change how quickly capability could change how quickly capability spreads. I'm going to talk about which spreads. I'm going to talk about which spreads. I'm going to talk about which models remain and are likely to remain models remain and are likely to remain models remain and are likely to remain available regardless. Now, before we go available regardless. Now, before we go available regardless. Now, before we go too far, my answer up front is yes.
-
too far, my answer up front is yes. too far, my answer up front is yes. Serious AI users should test Chinese Serious AI users should test Chinese Serious AI users should test Chinese models, but you need to plan to make models, but you need to plan to make models, but you need to plan to make separate decisions about the task, the separate decisions about the task, the separate decisions about the task, the model, and the deployment path for your model, and the deployment path for your model, and the deployment path for your model. So, let's not conflate those, and model. So, let's not conflate those, and model. So, let's not conflate those, and I'm going to get into all three of them I'm going to get into all three of them I'm going to get into all three of them here. here. here. If you are tackling high-volume, If you are tackling high-volume, If you are tackling high-volume, price-sensitive API work, DeepSeek price-sensitive API work, DeepSeek price-sensitive API work, DeepSeek absolutely belongs in your test set. It absolutely belongs in your test set. It absolutely belongs in your test set. It has an output price of, I think, yeah, has an output price of, I think, yeah, has an output price of, I think, yeah, 87 cents per million tokens, and it can 87 cents per million tokens, and it can 87 cents per million tokens, and it can change the economics of document change the economics of document change the economics of document processing, of code generation, of processing, of code generation, of processing, of code generation, of research pipelines, anything at volume. research pipelines, anything at volume. research pipelines, anything at volume. When you can cut the price down that When you can cut the price down that When you can cut the price down that much, 15x, 20x, 30x from whatever you're much, 15x, 20x, 30x from whatever you're much, 15x, 20x, 30x from whatever you're using now, it's a big deal. So, if using now, it's a big deal. So, if using now, it's a big deal. So, if you're looking at extraction, if you're you're looking at extraction, if you're you're looking at extraction, if you're looking at classification, if you're looking at classification, if you're looking at classification, if you're looking at first-pass research, at test looking at first-pass research, at test looking at first-pass research, at test generation, any other reviewable work in generation, any other reviewable work in generation, any other reviewable work in that line, the low price matters a ton that line, the low price matters a ton that line, the low price matters a ton because a machine can afford to check because a machine can afford to check because a machine can afford to check and recheck and check again, and you can and recheck and check again, and you can and recheck and check again, and you can get really good quality at a really get really good quality at a really get really good quality at a really affordable price. Now, I would be more affordable price. Now, I would be more affordable price. Now, I would be more cautious about using a tool like cautious about using a tool like cautious about using a tool like DeepSeek when one ambiguous answer DeepSeek when one ambiguous answer DeepSeek when one ambiguous answer drives an unrecoverable action. Right?
-
drives an unrecoverable action. Right? drives an unrecoverable action. Right? If there's something you can't easily If there's something you can't easily If there's something you can't easily verify, then you want to think about verify, then you want to think about verify, then you want to think about whether you are trading down quality in whether you are trading down quality in whether you are trading down quality in a way you can't accept. If you're a way you can't accept. If you're a way you can't accept. If you're looking at local work, if you're on looking at local work, if you're on looking at local work, if you're on personal hardware like this laptop right personal hardware like this laptop right personal hardware like this laptop right here, you want to look at some of the here, you want to look at some of the here, you want to look at some of the smaller Qwen models, some of the models smaller Qwen models, some of the models smaller Qwen models, some of the models DeepSeek distilled from R1. These are DeepSeek distilled from R1. These are DeepSeek distilled from R1. These are releases that are behind the real local releases that are behind the real local releases that are behind the real local AI story. They're not the same artifact AI story. They're not the same artifact AI story. They're not the same artifact as full DeepSeek R1, which had 671 as full DeepSeek R1, which had 671 as full DeepSeek R1, which had 671 billion total parameters, or DeepSeek V4 billion total parameters, or DeepSeek V4 billion total parameters, or DeepSeek V4 Pro, which has 1.6 trillion. Local Pro, which has 1.6 trillion. Local Pro, which has 1.6 trillion. Local models make sense for private notes, for models make sense for private notes, for models make sense for private notes, for offline work, for predictable offline work, for predictable offline work, for predictable availability, and workflows where you availability, and workflows where you availability, and workflows where you would rather trade some capability for a would rather trade some capability for a would rather trade some capability for a high high degree of control and privacy. high high degree of control and privacy. high high degree of control and privacy. A useful comparison might be maybe not a A useful comparison might be maybe not a A useful comparison might be maybe not a small Qwen model against Kimmy K3. It's small Qwen model against Kimmy K3. It's small Qwen model against Kimmy K3. It's more whether the small model is good more whether the small model is good more whether the small model is good enough for the exact job you want to enough for the exact job you want to enough for the exact job you want to keep on your machine. For long horizon keep on your machine. For long horizon keep on your machine. For long horizon coding, research, agent work, I would coding, research, agent work, I would coding, research, agent work, I would test a different set of Chinese models. test a different set of Chinese models. test a different set of Chinese models. I would look at the GLM, Kimmy, Qwen, I would look at the GLM, Kimmy, Qwen, I would look at the GLM, Kimmy, Qwen, and MiniMax against the best American and MiniMax against the best American and MiniMax against the best American frontier models using the same tools in frontier models using the same tools in frontier models using the same tools in the same acceptance standard. Don't the same acceptance standard. Don't the same acceptance standard. Don't change your bar. Some Chinese models are change your bar. Some Chinese models are change your bar. Some Chinese models are close to the frontier on selected tasks.
-
close to the frontier on selected tasks. close to the frontier on selected tasks. GLM 5.2 on coding is particularly GLM 5.2 on coding is particularly GLM 5.2 on coding is particularly strong, but that doesn't mean that they strong, but that doesn't mean that they strong, but that doesn't mean that they are automatically universal replacements are automatically universal replacements are automatically universal replacements in that domain. I would not use GLM 5.2 in that domain. I would not use GLM 5.2 in that domain. I would not use GLM 5.2 for every task in the coding universe for every task in the coding universe for every task in the coding universe just because it happens to be good at just because it happens to be good at just because it happens to be good at coding. A model can write an excellent coding. A model can write an excellent coding. A model can write an excellent patch, right? It can still lose a patch, right? It can still lose a patch, right? It can still lose a repository level task because it used repository level task because it used repository level task because it used the wrong tool or missed a test or the wrong tool or missed a test or the wrong tool or missed a test or failed after 20 minutes. And I'm I'm failed after 20 minutes. And I'm I'm failed after 20 minutes. And I'm I'm asking you to test because these things asking you to test because these things asking you to test because these things matter, right? A model might be strong matter, right? A model might be strong matter, right? A model might be strong at research in your area and weak at at research in your area and weak at at research in your area and weak at visual reasoning for you. Frontier visual reasoning for you. Frontier visual reasoning for you. Frontier becomes a useful phrase only after you becomes a useful phrase only after you becomes a useful phrase only after you understand what you're actually testing understand what you're actually testing understand what you're actually testing so you can measure where the frontier is so you can measure where the frontier is so you can measure where the frontier is for you. This is why benchmarks are for you. This is why benchmarks are for you. This is why benchmarks are notoriously easy to game and hard to notoriously easy to game and hard to notoriously easy to game and hard to trust. For sensitive work, I would trust. For sensitive work, I would trust. For sensitive work, I would decide where the data may go before I decide where the data may go before I decide where the data may go before I choose the model. A Chinese model choose the model. A Chinese model choose the model. A Chinese model running on your own server creates a running on your own server creates a running on your own server creates a completely different risk profile from completely different risk profile from completely different risk profile from the same model accessed through a the same model accessed through a the same model accessed through a first-party chat service in China. Those first-party chat service in China. Those first-party chat service in China. Those are very different decisions, right?
-
are very different decisions, right? are very different decisions, right? Let's start with DeepSeek at a slightly Let's start with DeepSeek at a slightly Let's start with DeepSeek at a slightly deeper level, pun intended. DeepSeek is deeper level, pun intended. DeepSeek is deeper level, pun intended. DeepSeek is pursuing extreme inference economics. pursuing extreme inference economics. pursuing extreme inference economics. Qwen is a broad family that runs from Qwen is a broad family that runs from Qwen is a broad family that runs from genuinely small open models to a genuinely small open models to a genuinely small open models to a separate hosted Max line whose weights separate hosted Max line whose weights separate hosted Max line whose weights are closed. Z.ai's GLM 5.2 is MIT are closed. Z.ai's GLM 5.2 is MIT are closed. Z.ai's GLM 5.2 is MIT licensed and very strong on long horizon licensed and very strong on long horizon licensed and very strong on long horizon coding. But, the released BF16 coding. But, the released BF16 coding. But, the released BF16 checkpoint is really large. It's like 1 checkpoint is really large. It's like 1 checkpoint is really large. It's like 1 and 1/2 terabytes. Moonshot is pricing and 1/2 terabytes. Moonshot is pricing and 1/2 terabytes. Moonshot is pricing Kimmy K3 as a premium frontier product. Kimmy K3 as a premium frontier product. Kimmy K3 as a premium frontier product. MiniMaxM3 combines coding and computer MiniMaxM3 combines coding and computer MiniMaxM3 combines coding and computer use and image and video understanding use and image and video understanding use and image and video understanding under a custom license that requires under a custom license that requires under a custom license that requires attribution and notice and requires attribution and notice and requires attribution and notice and requires prior written authorization above $20 prior written authorization above $20 prior written authorization above $20 million in annual revenue and prohibits million in annual revenue and prohibits million in annual revenue and prohibits military use. Sound different because military use. Sound different because military use. Sound different because they are. They're entirely different they are. They're entirely different they are. They're entirely different products and different strategies. products and different strategies. products and different strategies. Open is not one property, either. You Open is not one property, either. You Open is not one property, either. You need to ask a bunch of separate need to ask a bunch of separate need to ask a bunch of separate questions to understand what open really questions to understand what open really questions to understand what open really means. Can I call the model through an means. Can I call the model through an means. Can I call the model through an API? Can I download the weights? Does API? Can I download the weights? Does API? Can I download the weights? Does the license permit my use? Can my the license permit my use? Can my the license permit my use? Can my hardware run it at the context length hardware run it at the context length hardware run it at the context length that I need at the speed I need? A model that I need at the speed I need? A model that I need at the speed I need? A model can be open on one dimension and closed can be open on one dimension and closed can be open on one dimension and closed or or simply impractical on a bunch of or or simply impractical on a bunch of or or simply impractical on a bunch of the rest. Now, one technical reason the rest. Now, one technical reason the rest. Now, one technical reason enormous models can still be economical enormous models can still be economical enormous models can still be economical is this concept of a mixture of experts.
-
is this concept of a mixture of experts. is this concept of a mixture of experts. DeepSeekV4 Pro has 1.6 trillion total DeepSeekV4 Pro has 1.6 trillion total DeepSeekV4 Pro has 1.6 trillion total parameters. MiniMaxM3 has about 427 parameters. MiniMaxM3 has about 427 parameters. MiniMaxM3 has about 427 billion total and just 23 billion active billion total and just 23 billion active billion total and just 23 billion active per token. The router does not use the per token. The router does not use the per token. The router does not use the whole network for every token and part whole network for every token and part whole network for every token and part of the magic is in picking the right of the magic is in picking the right of the magic is in picking the right parameters to activate for a given parameters to activate for a given parameters to activate for a given token. And active parameters help token. And active parameters help token. And active parameters help explain how we handle compute per token. explain how we handle compute per token. explain how we handle compute per token. Total parameters still determine much of Total parameters still determine much of Total parameters still determine much of the storage, the memory, the networking, the storage, the memory, the networking, the storage, the memory, the networking, and the serving burden. And that's why a and the serving burden. And that's why a and the serving burden. And that's why a model can be inexpensive through an model can be inexpensive through an model can be inexpensive through an optimized API and absolutely absurd to optimized API and absolutely absurd to optimized API and absolutely absurd to run on your laptop. So, don't turn the run on your laptop. So, don't turn the run on your laptop. So, don't turn the active count into a simple hardware active count into a simple hardware active count into a simple hardware estimate. It's more complicated than estimate. It's more complicated than estimate. It's more complicated than that. You have to ask about checkpoint that. You have to ask about checkpoint that. You have to ask about checkpoint size, about precision, about context size, about precision, about context size, about precision, about context requirements, and about serving requirements, and about serving requirements, and about serving topology. And if all of that sounds topology. And if all of that sounds topology. And if all of that sounds really complicated, well, it kind of is. really complicated, well, it kind of is. really complicated, well, it kind of is. And that's exactly the advantage that And that's exactly the advantage that And that's exactly the advantage that frontier models have in the states is frontier models have in the states is frontier models have in the states is that you don't have to think about that that you don't have to think about that that you don't have to think about that stuff if you're just signing up for stuff if you're just signing up for stuff if you're just signing up for OpenAI. And that gives us one of our OpenAI. And that gives us one of our OpenAI. And that gives us one of our first hosting rules. If you want to run first hosting rules. If you want to run first hosting rules. If you want to run a smaller model locally when the a smaller model locally when the a smaller model locally when the checkpoint fits your hardware and the checkpoint fits your hardware and the checkpoint fits your hardware and the lower capability is enough for the task, lower capability is enough for the task, lower capability is enough for the task, absolutely you can do that. If you want absolutely you can do that. If you want absolutely you can do that. If you want to host a giant open weight model, make to host a giant open weight model, make to host a giant open weight model, make sure that you have hardware and data sure that you have hardware and data sure that you have hardware and data control and portability and you're set control and portability and you're set control and portability and you're set up to handle the scale you're talking up to handle the scale you're talking up to handle the scale you're talking about. Downloading the weights does not about. Downloading the weights does not about. Downloading the weights does not magically solve the hardware and team magically solve the hardware and team magically solve the hardware and team requirements you actually need to serve requirements you actually need to serve requirements you actually need to serve what is effectively an enterprise model
-
what is effectively an enterprise model what is effectively an enterprise model from China. Hardware and licensing can from China. Hardware and licensing can from China. Hardware and licensing can tell you what you have to deploy to get tell you what you have to deploy to get tell you what you have to deploy to get a model up, right? So, that's a part of a model up, right? So, that's a part of a model up, right? So, that's a part of it. But even then, they don't tell you it. But even then, they don't tell you it. But even then, they don't tell you whether the output is going to be good whether the output is going to be good whether the output is going to be good enough. For that, the gap has to be enough. For that, the gap has to be enough. For that, the gap has to be measured against the work that you are measured against the work that you are measured against the work that you are doing. Chinese labs can report near doing. Chinese labs can report near doing. Chinese labs can report near frontier or leading scores on coding and frontier or leading scores on coding and frontier or leading scores on coding and math and tool use and long context math and tool use and long context math and tool use and long context tasks, but that doesn't mean they work tasks, but that doesn't mean they work tasks, but that doesn't mean they work for you, right? You have to think about for you, right? You have to think about for you, right? You have to think about your harness, about your reasoning your harness, about your reasoning your harness, about your reasoning budget, and about your tools. And if budget, and about your tools. And if budget, and about your tools. And if you're not ready to go to that level of you're not ready to go to that level of you're not ready to go to that level of detail, you're probably not ready for a detail, you're probably not ready for a detail, you're probably not ready for a Chinese model conversation. In May, the Chinese model conversation. In May, the Chinese model conversation. In May, the US government CAISI evaluation called US government CAISI evaluation called US government CAISI evaluation called DeepSeek V4 Pro the most capable Chinese DeepSeek V4 Pro the most capable Chinese DeepSeek V4 Pro the most capable Chinese model it had tested, while estimating model it had tested, while estimating model it had tested, while estimating that it remained about eight months that it remained about eight months that it remained about eight months behind the leading US frontier. That behind the leading US frontier. That behind the leading US frontier. That capability gap was smaller on individual capability gap was smaller on individual capability gap was smaller on individual tasks, which tells us the model is spiky tasks, which tells us the model is spiky tasks, which tells us the model is spiky and reinforcement learned in specific and reinforcement learned in specific and reinforcement learned in specific areas. And of course, the economics are areas. And of course, the economics are areas. And of course, the economics are super different, right? DeepSeek would super different, right? DeepSeek would super different, right? DeepSeek would cost less, which is what everybody cost less, which is what everybody cost less, which is what everybody expects. Well, in other tasks, the expects. Well, in other tasks, the expects. Well, in other tasks, the cheaper tokens didn't end up producing cheaper tokens didn't end up producing cheaper tokens didn't end up producing the cheaper solution because either it the cheaper solution because either it the cheaper solution because either it had to spend so much tokens to get had to spend so much tokens to get had to spend so much tokens to get there, or it never got there at all. And there, or it never got there at all. And there, or it never got there at all. And so, what CAISI found is that DeepSeek so, what CAISI found is that DeepSeek so, what CAISI found is that DeepSeek ranged from 53% cheaper to 41% more ranged from 53% cheaper to 41% more ranged from 53% cheaper to 41% more expensive per correctly solved task expensive per correctly solved task expensive per correctly solved task across seven benchmarks. Have I said across seven benchmarks. Have I said across seven benchmarks. Have I said testing enough? You got to test.
-
testing enough? You got to test. testing enough? You got to test. Otherwise, you won't realize savings Otherwise, you won't realize savings Otherwise, you won't realize savings here. DeepSeek has changed their prices here. DeepSeek has changed their prices here. DeepSeek has changed their prices since as well. So, those percentages are since as well. So, those percentages are since as well. So, those percentages are not a permanent verdict. They're an not a permanent verdict. They're an not a permanent verdict. They're an indicator of where you need to test. indicator of where you need to test. indicator of where you need to test. They're proof that token price and They're proof that token price and They're proof that token price and finished work cost can often point you finished work cost can often point you finished work cost can often point you in opposite directions, and almost in opposite directions, and almost in opposite directions, and almost everybody cares about the finished work everybody cares about the finished work everybody cares about the finished work cost. Those results remind us why cost cost. Those results remind us why cost cost. Those results remind us why cost per accepted result is the gold per accepted result is the gold per accepted result is the gold standard, and you have to define your standard, and you have to define your standard, and you have to define your task clearly enough so you can do that. task clearly enough so you can do that. task clearly enough so you can do that. A cheap model can become expensive when A cheap model can become expensive when A cheap model can become expensive when it produces longer reasoning traces, it produces longer reasoning traces, it produces longer reasoning traces, which is literally what we see with which is literally what we see with which is literally what we see with Claude 3 versus Fable 5. When it makes Claude 3 versus Fable 5. When it makes Claude 3 versus Fable 5. When it makes unnecessary tool calls, when it fails unnecessary tool calls, when it fails unnecessary tool calls, when it fails late, or hands a human more cleanup. A late, or hands a human more cleanup. A late, or hands a human more cleanup. A premium model can end up being cheaper, premium model can end up being cheaper, premium model can end up being cheaper, like Fable 5, if it succeeds in one like Fable 5, if it succeeds in one like Fable 5, if it succeeds in one pass. And actually, Opus 5 has just come pass. And actually, Opus 5 has just come pass. And actually, Opus 5 has just come out, and it has the same token out, and it has the same token out, and it has the same token efficiency focus, which is how frontier efficiency focus, which is how frontier efficiency focus, which is how frontier models are starting to compete right models are starting to compete right models are starting to compete right now. So, if we back up a bit, for now. So, if we back up a bit, for now. So, if we back up a bit, for bounded, repeatable, high-volume jobs, bounded, repeatable, high-volume jobs, bounded, repeatable, high-volume jobs, Chinese models can be an extraordinary Chinese models can be an extraordinary Chinese models can be an extraordinary value. For ambiguous work, where the value. For ambiguous work, where the value. For ambiguous work, where the model has to infer intent, or recover model has to infer intent, or recover model has to infer intent, or recover from surprises, or make a a high-stakes from surprises, or make a a high-stakes from surprises, or make a a high-stakes judgment call, the strongest American judgment call, the strongest American judgment call, the strongest American frontier systems remain the baseline frontier systems remain the baseline frontier systems remain the baseline that I would assume you have to beat to that I would assume you have to beat to that I would assume you have to beat to justify. Leading American services often justify. Leading American services often justify. Leading American services often come with more mature tool integrations, come with more mature tool integrations, come with more mature tool integrations, with enterprise controls, with with enterprise controls, with with enterprise controls, with operational evidence. Some Chinese operational evidence. Some Chinese operational evidence. Some Chinese offerings will counter, and they'll offerings will counter, and they'll offerings will counter, and they'll often counter with things that we feel often counter with things that we feel often counter with things that we feel like we hear a lot in the narrative,
-
like we hear a lot in the narrative, like we hear a lot in the narrative, right, in the media. Lower prices, more right, in the media. Lower prices, more right, in the media. Lower prices, more open weight options, more deployment open weight options, more deployment open weight options, more deployment choice. And neither bundle's going to choice. And neither bundle's going to choice. And neither bundle's going to work for everybody. The relative value work for everybody. The relative value work for everybody. The relative value changes when your job values raw changes when your job values raw changes when your job values raw judgment, or values control, or values judgment, or values control, or values judgment, or values control, or values portability, or values cost. Now, there portability, or values cost. Now, there portability, or values cost. Now, there are two more complexifiers I want to are two more complexifiers I want to are two more complexifiers I want to mention here. mention here. mention here. How does capability move into the model? How does capability move into the model? How does capability move into the model? And where does your data go? And that And where does your data go? And that And where does your data go? And that takes us into policy and distillation. takes us into policy and distillation. takes us into policy and distillation. And it begins with those restrict And it begins with those restrict And it begins with those restrict specified advanced chips, manufacturing specified advanced chips, manufacturing specified advanced chips, manufacturing equipment, software, end users, and end equipment, software, end users, and end equipment, software, end users, and end uses, different from users, without uses, different from users, without uses, different from users, without telling us which hardware trained every telling us which hardware trained every telling us which hardware trained every model. model. model. I would not say controls invented I would not say controls invented I would not say controls invented Chinese efficiency. I think that Chinese Chinese efficiency. I think that Chinese Chinese efficiency. I think that Chinese efficiency is something that we see efficiency is something that we see efficiency is something that we see across a wide range of Chinese across a wide range of Chinese across a wide range of Chinese technologies over the last two or three technologies over the last two or three technologies over the last two or three decades. Mixture of experts, low decades. Mixture of experts, low decades. Mixture of experts, low precision training, sparse attention, precision training, sparse attention, precision training, sparse attention, quantization, and better hardware quantization, and better hardware quantization, and better hardware utilization are techniques that we do utilization are techniques that we do utilization are techniques that we do see in Chinese models that reflect that see in Chinese models that reflect that see in Chinese models that reflect that broader tendency. And the controls broader tendency. And the controls broader tendency. And the controls simply increased the strategic simply increased the strategic simply increased the strategic incentives to get more capability. The incentives to get more capability. The incentives to get more capability. The Americans don't have to worry about Americans don't have to worry about Americans don't have to worry about chips, they have to worry about power.
-
chips, they have to worry about power. chips, they have to worry about power. China is pushing at the opposite side of China is pushing at the opposite side of China is pushing at the opposite side of the system in most of these incentive the system in most of these incentive the system in most of these incentive sets versus what American AI companies sets versus what American AI companies sets versus what American AI companies are doing. It's AI plus policy supports are doing. It's AI plus policy supports are doing. It's AI plus policy supports open-source communities in the sharing open-source communities in the sharing open-source communities in the sharing of models, tools, and data sets. That of models, tools, and data sets. That of models, tools, and data sets. That doesn't prove that Beijing is ordering a doesn't prove that Beijing is ordering a doesn't prove that Beijing is ordering a lab to release weights. It's just a part lab to release weights. It's just a part lab to release weights. It's just a part of the community there. The labs have of the community there. The labs have of the community there. The labs have commercial reasons to build an commercial reasons to build an commercial reasons to build an ecosystem, while official policy can ecosystem, while official policy can ecosystem, while official policy can support broad adoption. Open weights can support broad adoption. Open weights can support broad adoption. Open weights can also reduce users dependence on American also reduce users dependence on American also reduce users dependence on American APIs, although the policy doesn't state APIs, although the policy doesn't state APIs, although the policy doesn't state that motive explicitly. Now, that motive explicitly. Now, that motive explicitly. Now, distillation, which has been a very distillation, which has been a very distillation, which has been a very hotly debated topic, uh in fact, Kimik 3 hotly debated topic, uh in fact, Kimik 3 hotly debated topic, uh in fact, Kimik 3 was publicly accused of distilling Fable was publicly accused of distilling Fable was publicly accused of distilling Fable 5 weights just last week. Distillation 5 weights just last week. Distillation 5 weights just last week. Distillation connects those stories. In the normal connects those stories. In the normal connects those stories. In the normal version, a powerful teacher model version, a powerful teacher model version, a powerful teacher model generates examples, answers, reasoning generates examples, answers, reasoning generates examples, answers, reasoning traces, or preferences. Another model, traces, or preferences. Another model, traces, or preferences. Another model, often smaller or more specialized, often smaller or more specialized, often smaller or more specialized, trains on those outputs and inherits trains on those outputs and inherits trains on those outputs and inherits some of the teachers' behavior without some of the teachers' behavior without some of the teachers' behavior without reproducing the teachers' full training reproducing the teachers' full training reproducing the teachers' full training program. OpenAI actually sells program. OpenAI actually sells program. OpenAI actually sells distillation as a product. Frontier Labs distillation as a product. Frontier Labs distillation as a product. Frontier Labs use it internally. There is a Mythos use it internally. There is a Mythos use it internally. There is a Mythos teacher model that teaches the released teacher model that teaches the released teacher model that teaches the released versions of Mythos and Fable. DeepSeek versions of Mythos and Fable. DeepSeek versions of Mythos and Fable. DeepSeek openly used R1 generated samples to openly used R1 generated samples to openly used R1 generated samples to train smaller models based on Qwen and train smaller models based on Qwen and train smaller models based on Qwen and on Llama. Its repository says that Qwen on Llama. Its repository says that Qwen on Llama. Its repository says that Qwen derived student models were fine-tuned derived student models were fine-tuned derived student models were fine-tuned on 800,000 samples curated from R1. That on 800,000 samples curated from R1. That on 800,000 samples curated from R1. That is super ordinary disclosed is super ordinary disclosed is super ordinary disclosed teacher-to-student model training for teacher-to-student model training for teacher-to-student model training for AI.
-
AI. AI. A distilled student is not the full A distilled student is not the full A distilled student is not the full teacher squeezed into a smaller file. It teacher squeezed into a smaller file. It teacher squeezed into a smaller file. It can inherit useful reasoning patterns on can inherit useful reasoning patterns on can inherit useful reasoning patterns on the kinds of examples it saw, but it the kinds of examples it saw, but it the kinds of examples it saw, but it loses breadth, reliability, and loses breadth, reliability, and loses breadth, reliability, and safeguards. This is why we talk about safeguards. This is why we talk about safeguards. This is why we talk about the idea of a big model smell the idea of a big model smell the idea of a big model smell for models that are teacher models. for models that are teacher models. for models that are teacher models. That is another reason the phrase That is another reason the phrase That is another reason the phrase DeepSeek runs locally, the local student DeepSeek runs locally, the local student DeepSeek runs locally, the local student and the full Frontier teacher are very and the full Frontier teacher are very and the full Frontier teacher are very different systems. The controversy about different systems. The controversy about different systems. The controversy about distillation begins when the teachers' distillation begins when the teachers' distillation begins when the teachers' outputs were never authorized for that outputs were never authorized for that outputs were never authorized for that use case. Anthropic says that DeepSeek, use case. Anthropic says that DeepSeek, use case. Anthropic says that DeepSeek, Moonshot, and MiniMax used about 24,000 Moonshot, and MiniMax used about 24,000 Moonshot, and MiniMax used about 24,000 fraudulent accounts to generate more fraudulent accounts to generate more fraudulent accounts to generate more than 16 million cloud exchanges while than 16 million cloud exchanges while than 16 million cloud exchanges while evading access restrictions. A later evading access restrictions. A later evading access restrictions. A later White House memo described White House memo described White House memo described industrial-scale extraction by industrial-scale extraction by industrial-scale extraction by China-based entities, but did not name China-based entities, but did not name China-based entities, but did not name those labs or validate Anthropic's those labs or validate Anthropic's those labs or validate Anthropic's numbers. Now, that's since changed numbers. Now, that's since changed numbers. Now, that's since changed because Kimmy K3 was named by the White because Kimmy K3 was named by the White because Kimmy K3 was named by the White House just last week. And these are very House just last week. And these are very House just last week. And these are very serious allegations, not court findings, serious allegations, not court findings, serious allegations, not court findings, right? But allegations. And the public right? But allegations. And the public right? But allegations. And the public record doesn't tell us how much release record doesn't tell us how much release record doesn't tell us how much release capability came from the alleged capability came from the alleged capability came from the alleged campaigns that Anthropic is describing.
-
campaigns that Anthropic is describing. campaigns that Anthropic is describing. I think a useful distinction here is I think a useful distinction here is I think a useful distinction here is authorization and conduct. A lab can authorization and conduct. A lab can authorization and conduct. A lab can distill its own or a licensed teacher. distill its own or a licensed teacher. distill its own or a licensed teacher. It can train on API outputs contrary to It can train on API outputs contrary to It can train on API outputs contrary to contractual restrictions or use fake contractual restrictions or use fake contractual restrictions or use fake accounts and access control evasion at accounts and access control evasion at accounts and access control evasion at scale. The technical method may overlap, scale. The technical method may overlap, scale. The technical method may overlap, but the contractual and potential legal but the contractual and potential legal but the contractual and potential legal claims really differ in this situation. claims really differ in this situation. claims really differ in this situation. And typically, Chinese entities are not And typically, Chinese entities are not And typically, Chinese entities are not authorized to distill American models. authorized to distill American models. authorized to distill American models. And and this is why the fight matters, And and this is why the fight matters, And and this is why the fight matters, even if you want to reserve judgment on even if you want to reserve judgment on even if you want to reserve judgment on the companies. Chips are physical and the companies. Chips are physical and the companies. Chips are physical and model outputs are information. model outputs are information. model outputs are information. Restricting a teacher model's GPU Restricting a teacher model's GPU Restricting a teacher model's GPU doesn't control everything that that doesn't control everything that that doesn't control everything that that model can teach through an API, right? model can teach through an API, right? model can teach through an API, right? This is reminding me, as a graybeard, of This is reminding me, as a graybeard, of This is reminding me, as a graybeard, of back in the '90s when we talked about back in the '90s when we talked about back in the '90s when we talked about Napster and music wanting to be free. Napster and music wanting to be free. Napster and music wanting to be free. Capability can move very easily across Capability can move very easily across Capability can move very easily across the global internet. It can move through the global internet. It can move through the global internet. It can move through synthetic data, through fine-tuning, synthetic data, through fine-tuning, synthetic data, through fine-tuning, through reinforcement learning, through through reinforcement learning, through through reinforcement learning, through distillation. The transfer may not be distillation. The transfer may not be distillation. The transfer may not be perfect, but it may not need to be, perfect, but it may not need to be, perfect, but it may not need to be, right? It can reproduce enough valuable right? It can reproduce enough valuable right? It can reproduce enough valuable behavior fast enough and cheaply enough behavior fast enough and cheaply enough behavior fast enough and cheaply enough that it's worth doing. And that means that it's worth doing. And that means that it's worth doing. And that means today's capability gap can move faster today's capability gap can move faster today's capability gap can move faster than hardware policy alone would than hardware policy alone would than hardware policy alone would suggest, and it gives frontier providers suggest, and it gives frontier providers suggest, and it gives frontier providers strong incentives to tighten their strong incentives to tighten their strong incentives to tighten their access, their monitoring, and their access, their monitoring, and their access, their monitoring, and their account controls, which is exactly what account controls, which is exactly what account controls, which is exactly what we see with Anthropic and the moonshot we see with Anthropic and the moonshot we see with Anthropic and the moonshot program, for example. Now, for a program, for example. Now, for a program, for example. Now, for a customer, that absolutely increases the customer, that absolutely increases the customer, that absolutely increases the value of owning your own evals and value of owning your own evals and value of owning your own evals and preserving an exit path. If you think
-
preserving an exit path. If you think preserving an exit path. If you think about it as this creates uncertainty, a about it as this creates uncertainty, a about it as this creates uncertainty, a workflow is much less exposed to one workflow is much less exposed to one workflow is much less exposed to one provider when the prompts, the tests, provider when the prompts, the tests, provider when the prompts, the tests, and the tools can easily be lifted and and the tools can easily be lifted and and the tools can easily be lifted and shifted to another model. Now, let's shifted to another model. Now, let's shifted to another model. Now, let's talk about the hosting question talk about the hosting question talk about the hosting question directly. There are three basic directly. There are three basic directly. There are three basic deployment choices that you have, right? deployment choices that you have, right? deployment choices that you have, right? You have a first-party API, that's the You have a first-party API, that's the You have a first-party API, that's the easiest, but then the provider controls easiest, but then the provider controls easiest, but then the provider controls the service, the provider controls the the service, the provider controls the the service, the provider controls the price, the logging, the availability. price, the logging, the availability. price, the logging, the availability. Now, a third-party host can run Now, a third-party host can run Now, a third-party host can run OpenWeights in a region and contract any OpenWeights in a region and contract any OpenWeights in a region and contract any way you prefer. Self-hosting gives you way you prefer. Self-hosting gives you way you prefer. Self-hosting gives you the most control over prompts and the most control over prompts and the most control over prompts and runtime behavior, but it makes you runtime behavior, but it makes you runtime behavior, but it makes you responsible for hardware, for security, responsible for hardware, for security, responsible for hardware, for security, for updates, for monitoring, for for updates, for monitoring, for for updates, for monitoring, for everything that goes with having a everything that goes with having a everything that goes with having a model. And self-host only makes sense, I model. And self-host only makes sense, I model. And self-host only makes sense, I think, when four things are true. The think, when four things are true. The think, when four things are true. The weights need to be available and the weights need to be available and the weights need to be available and the license needs to permit your use, like license needs to permit your use, like license needs to permit your use, like that's fairly intuitive, I think. The that's fairly intuitive, I think. The that's fairly intuitive, I think. The data or sovereignty requirement has to data or sovereignty requirement has to data or sovereignty requirement has to be real enough that you need to have the be real enough that you need to have the be real enough that you need to have the model on your premises. The workload model on your premises. The workload model on your premises. The workload should justify the economic should justify the economic should justify the economic infrastructure you need, the AI infrastructure you need, the AI infrastructure you need, the AI infrastructure to run, and you need to infrastructure to run, and you need to infrastructure to run, and you need to have a team that can operate the stack have a team that can operate the stack have a team that can operate the stack safely. If those conditions aren't safely. If those conditions aren't safely. If those conditions aren't present, you're going to have to look at present, you're going to have to look at present, you're going to have to look at either an API or a managed third-party either an API or a managed third-party either an API or a managed third-party deployment, and that's usually going to deployment, and that's usually going to deployment, and that's usually going to be a better solution for you.
-
be a better solution for you. be a better solution for you. Self-hosting also needs a named owner Self-hosting also needs a named owner Self-hosting also needs a named owner for some of the high-liability stuff. for some of the high-liability stuff. for some of the high-liability stuff. And so, that gets at the team side, And so, that gets at the team side, And so, that gets at the team side, right? For authentication, for model right? For authentication, for model right? For authentication, for model server patches, for capacity and server patches, for capacity and server patches, for capacity and observability. Without a team that has observability. Without a team that has observability. Without a team that has that kind of accountability, you're just that kind of accountability, you're just that kind of accountability, you're just sort of converting your vendor risk into sort of converting your vendor risk into sort of converting your vendor risk into an operating problem that you may only an operating problem that you may only an operating problem that you may only discover when something goes wrong. discover when something goes wrong. discover when something goes wrong. So, don't assume that Chinese model So, don't assume that Chinese model So, don't assume that Chinese model means the data path or the governing means the data path or the governing means the data path or the governing regime is set. China's interim rules regime is set. China's interim rules regime is set. China's interim rules govern generative AI services offered to govern generative AI services offered to govern generative AI services offered to the public inside mainland China. They the public inside mainland China. They the public inside mainland China. They don't automatically govern raw model don't automatically govern raw model don't automatically govern raw model weights privately operated outside weights privately operated outside weights privately operated outside China. So, Deep Seek's first-party China. So, Deep Seek's first-party China. So, Deep Seek's first-party privacy policy says personal data is privacy policy says personal data is privacy policy says personal data is processed and stored in China, and processed and stored in China, and processed and stored in China, and inputs may be used to improve its inputs may be used to improve its inputs may be used to improve its technology. technology. technology. That's not surprising. Alibaba says That's not surprising. Alibaba says That's not surprising. Alibaba says Model Studio customer data is not used Model Studio customer data is not used Model Studio customer data is not used for model training and offers deployment for model training and offers deployment for model training and offers deployment scopes that can exclude Mainland China.
-
scopes that can exclude Mainland China. scopes that can exclude Mainland China. And that's something not a lot of people And that's something not a lot of people And that's something not a lot of people know. And you can go even further here, know. And you can go even further here, know. And you can go even further here, right? An American third-party host right? An American third-party host right? An American third-party host running Quen weights creates yet another running Quen weights creates yet another running Quen weights creates yet another path, right? Where that would be under path, right? Where that would be under path, right? Where that would be under American jurisdiction. A properly American jurisdiction. A properly American jurisdiction. A properly configured self-hosted model can keep configured self-hosted model can keep configured self-hosted model can keep prompts away from the original lab and prompts away from the original lab and prompts away from the original lab and remove provider side runtime filters. remove provider side runtime filters. remove provider side runtime filters. But, it does not remove refusal behavior But, it does not remove refusal behavior But, it does not remove refusal behavior or political bias learned in the or political bias learned in the or political bias learned in the weights. You may have to fine-tune, weights. You may have to fine-tune, weights. You may have to fine-tune, right? You may have to check inference right? You may have to check inference right? You may have to check inference calls, tools, telemetry, and logs to calls, tools, telemetry, and logs to calls, tools, telemetry, and logs to figure that out. And then a lot of figure that out. And then a lot of figure that out. And then a lot of people run a fine-tune on a Chinese people run a fine-tune on a Chinese people run a fine-tune on a Chinese model to adjust it for that American model to adjust it for that American model to adjust it for that American sensibility. Moving the weights onto sensibility. Moving the weights onto sensibility. Moving the weights onto your server changes who can see the your server changes who can see the your server changes who can see the data. It doesn't make the rest of the data. It doesn't make the rest of the data. It doesn't make the rest of the software supply chain automatically software supply chain automatically software supply chain automatically trustworthy by magic, just to be clear. trustworthy by magic, just to be clear. trustworthy by magic, just to be clear. First, I name the job. Is it bounded First, I name the job. Is it bounded First, I name the job. Is it bounded volume work? Is it local assistance? Is volume work? Is it local assistance? Is volume work? Is it local assistance? Is it sensitive internal work? Or is it it sensitive internal work? Or is it it sensitive internal work? Or is it frontier judgment? I do not choose a frontier judgment? I do not choose a frontier judgment? I do not choose a model before I choose what I can model before I choose what I can model before I choose what I can tolerate from a failure perspective.
-
tolerate from a failure perspective. tolerate from a failure perspective. Second, I name the artifact in the Second, I name the artifact in the Second, I name the artifact in the deployment that I'm expecting. Is it an deployment that I'm expecting. Is it an deployment that I'm expecting. Is it an API? Is it weights I can download? Is API? Is it weights I can download? Is API? Is it weights I can download? Is some kind of promised release that's some kind of promised release that's some kind of promised release that's coming in the future? What license coming in the future? What license coming in the future? What license applies to all of that? What What is the applies to all of that? What What is the applies to all of that? What What is the full model require at a realistic full model require at a realistic full model require at a realistic context length for my work? Third, I context length for my work? Third, I context length for my work? Third, I measure cost per accepted result, which measure cost per accepted result, which measure cost per accepted result, which we talked about earlier. I include input we talked about earlier. I include input we talked about earlier. I include input and output and reasoning and tool calls and output and reasoning and tool calls and output and reasoning and tool calls and retries and latency, the entire and retries and latency, the entire and retries and latency, the entire chain and the infrastructure when I chain and the infrastructure when I chain and the infrastructure when I measure that. I don't want hidden costs measure that. I don't want hidden costs measure that. I don't want hidden costs here. Fourth, I trace the data and the here. Fourth, I trace the data and the here. Fourth, I trace the data and the exit path really clearly. I identify exit path really clearly. I identify exit path really clearly. I identify which company receives the prompt, where which company receives the prompt, where which company receives the prompt, where it's stored and processed, the governing it's stored and processed, the governing it's stored and processed, the governing law and contract there, and the law and contract there, and the law and contract there, and the retention and training terms. I then ask retention and training terms. I then ask retention and training terms. I then ask whether I can move the workflow and keep whether I can move the workflow and keep whether I can move the workflow and keep my prompts, my tools, my evaluations if my prompts, my tools, my evaluations if my prompts, my tools, my evaluations if the provider changes its price, the provider changes its price, the provider changes its price, whether I can keep my model and access whether I can keep my model and access whether I can keep my model and access policy. Those are all things that you policy. Those are all things that you policy. Those are all things that you have to think about if you're thinking have to think about if you're thinking have to think about if you're thinking about taking a serious bet on Chinese about taking a serious bet on Chinese about taking a serious bet on Chinese models. Now, if you're just an models. Now, if you're just an models. Now, if you're just an individual, you can run this test individual, you can run this test individual, you can run this test without building a laboratory, right?
-
without building a laboratory, right? without building a laboratory, right? You can take 20 examples of real work, You can take 20 examples of real work, You can take 20 examples of real work, you can include ugly edge cases, you can you can include ugly edge cases, you can you can include ugly edge cases, you can pick a Chinese model, and you can run a pick a Chinese model, and you can run a pick a Chinese model, and you can run a test through it and and find out whether test through it and and find out whether test through it and and find out whether it works. And I did this, right? I did it works. And I did this, right? I did it works. And I did this, right? I did this when I was building the Ringer this when I was building the Ringer this when I was building the Ringer multi-agent that I launched a week or multi-agent that I launched a week or multi-agent that I launched a week or two ago, where you can actually have two ago, where you can actually have two ago, where you can actually have multi-agent swarms that deliver multi-agent swarms that deliver multi-agent swarms that deliver frontier-level intelligence for much frontier-level intelligence for much frontier-level intelligence for much less than frontier cost, where Fable less than frontier cost, where Fable less than frontier cost, where Fable orchestrates and much cheaper models orchestrates and much cheaper models orchestrates and much cheaper models actually do the work. And what I found actually do the work. And what I found actually do the work. And what I found was Qwen was really helpful as part of was Qwen was really helpful as part of was Qwen was really helpful as part of an agent swarm. And so, by testing, I an agent swarm. And so, by testing, I an agent swarm. And so, by testing, I determined that for a lot of use cases, determined that for a lot of use cases, determined that for a lot of use cases, Qwen was a helpful swarm agent Qwen was a helpful swarm agent Qwen was a helpful swarm agent participant, but of course I configured participant, but of course I configured participant, but of course I configured Ringers so you can use it however you Ringers so you can use it however you Ringers so you can use it however you want. want. want. Similarly, you need to do your own Similarly, you need to do your own Similarly, you need to do your own testing. You need to decide where you testing. You need to decide where you testing. You need to decide where you need frontier model intelligence, where need frontier model intelligence, where need frontier model intelligence, where you would put GLM or Kimmy or Qwen or you would put GLM or Kimmy or Qwen or you would put GLM or Kimmy or Qwen or Minimax on your particular work. Minimax on your particular work. Minimax on your particular work. And absolutely, you need to think about And absolutely, you need to think about And absolutely, you need to think about sensitive workloads. You want to think sensitive workloads. You want to think sensitive workloads. You want to think about where is the data architecture about where is the data architecture about where is the data architecture that allows you to preserve your that allows you to preserve your that allows you to preserve your liability in a situation where people liability in a situation where people liability in a situation where people may expect that you processed your data may expect that you processed your data may expect that you processed your data using an intelligence based in your using an intelligence based in your using an intelligence based in your jurisdiction. So, should you use Chinese jurisdiction. So, should you use Chinese jurisdiction. So, should you use Chinese models? My answer is yes, selectively. I models? My answer is yes, selectively. I models? My answer is yes, selectively. I mean, I do, and sometimes aggressively.
-
mean, I do, and sometimes aggressively. mean, I do, and sometimes aggressively. Use them as specialists, use them as Use them as specialists, use them as Use them as specialists, use them as challengers, use them in some workflows challengers, use them in some workflows challengers, use them in some workflows as the default, but don't be sloppy in as the default, but don't be sloppy in as the default, but don't be sloppy in your analysis. Don't use Chinese as your analysis. Don't use Chinese as your analysis. Don't use Chinese as shorthand for a particular capability shorthand for a particular capability shorthand for a particular capability score, for a particular security score, for a particular security score, for a particular security verdict, for a typical deployment plan. verdict, for a typical deployment plan. verdict, for a typical deployment plan. Start with the work. Check what you're Start with the work. Check what you're Start with the work. Check what you're actually doing. Make sure you understand actually doing. Make sure you understand actually doing. Make sure you understand the full hardware burden of what you're the full hardware burden of what you're the full hardware burden of what you're going after. Understand the data path. going after. Understand the data path. going after. Understand the data path. Understand the real cost for the work Understand the real cost for the work Understand the real cost for the work you want done. Country of origin is just you want done. Country of origin is just you want done. Country of origin is just the very beginning of having a the very beginning of having a the very beginning of having a conversation about due diligence. It conversation about due diligence. It conversation about due diligence. It doesn't substitute for the full testing doesn't substitute for the full testing doesn't substitute for the full testing regime that you need to do. So, if regime that you need to do. So, if regime that you need to do. So, if you're serious about doing Chinese model you're serious about doing Chinese model you're serious about doing Chinese model work, you need to take the work, you need to take the work, you need to take the responsibility to do that testing. Now, responsibility to do that testing. Now, responsibility to do that testing. Now, if you want to see how I tested, how I if you want to see how I tested, how I if you want to see how I tested, how I think about these models, I have way think about these models, I have way think about these models, I have way more detail in the article for today. I more detail in the article for today. I more detail in the article for today. I know it's a hot topic, so I have a deep know it's a hot topic, so I have a deep know it's a hot topic, so I have a deep dive on these models, and I will show dive on these models, and I will show dive on these models, and I will show you, if you're interested, how to put you, if you're interested, how to put you, if you're interested, how to put them into that multi-agent swarm them into that multi-agent swarm them into that multi-agent swarm approach, that ringer approach that I approach, that ringer approach that I approach, that ringer approach that I used and introduced a couple of videos used and introduced a couple of videos used and introduced a couple of videos ago. I hope you had fun with this one, ago. I hope you had fun with this one, ago. I hope you had fun with this one, and I will be back soon with whatever's and I will be back soon with whatever's and I will be back soon with whatever's next in the world of AI.
Summary
The main theme is the analysis of emerging Chinese large language models like Kimi K3 and DeepSeek V4 Pro, questioning their current capabilities and frontier closure. The discussion highlights their significant parameters, diverse pricing, and open-weight status, contrasting them with US models. The practical takeaway is to provide individuals and companies with a method to critically evaluate and choose whether to integrate these Chinese models into their technology stacks, considering API usage versus self-hosting.