← Back
Theo June 30, 2026 31m

Why is OpenAI so much more efficient?

Read full transcript 24 segments
  1. There's a lot of debate about which There's a lot of debate about which models are the smartest and the most models are the smartest and the most models are the smartest and the most capable, especially for things like capable, especially for things like capable, especially for things like writing code. The winner right now is writing code. The winner right now is writing code. The winner right now is Fable, but we don't have access to it, Fable, but we don't have access to it, Fable, but we don't have access to it, so we have to debate between Opus 4 8, so we have to debate between Opus 4 8, so we have to debate between Opus 4 8, Gemini, and whatever is going on with Gemini, and whatever is going on with Gemini, and whatever is going on with opening eye, which right now seems to be opening eye, which right now seems to be opening eye, which right now seems to be GPT-55. I've heard rumors something's GPT-55. I've heard rumors something's GPT-55. I've heard rumors something's coming, but who knows when that will coming, but who knows when that will coming, but who knows when that will happen. There is one thing that we can't happen. There is one thing that we can't happen. There is one thing that we can't really argue about though, and that is really argue about though, and that is really argue about though, and that is the efficiency of these different the efficiency of these different the efficiency of these different models. Because as smart as models like models. Because as smart as models like models. Because as smart as models like Gemini and Fable can be, neither are Gemini and Fable can be, neither are Gemini and Fable can be, neither are particularly efficient when you compare particularly efficient when you compare particularly efficient when you compare them to what OpenAI can do on much much them to what OpenAI can do on much much them to what OpenAI can do on much much smaller token budgets. Charts like this smaller token budgets. Charts like this smaller token budgets. Charts like this one from Deep SWE really emphasize what one from Deep SWE really emphasize what one from Deep SWE really emphasize what I'm trying to talk about here, where I'm trying to talk about here, where I'm trying to talk about here, where GPT-55 medium got an incredible score on GPT-55 medium got an incredible score on GPT-55 medium got an incredible score on their benchmark with only 20K tokens, their benchmark with only 20K tokens, their benchmark with only 20K tokens, whereas a model like Opus 4 8 scored whereas a model like Opus 4 8 scored whereas a model like Opus 4 8 scored lower at 50K tokens. And even the lower at 50K tokens. And even the lower at 50K tokens. And even the heaviest 55 run on X high was only heaviest 55 run on X high was only heaviest 55 run on X high was only 46,000 tokens. The efficiency that 46,000 tokens. The efficiency that 46,000 tokens. The efficiency that OpenAI is when pulling off with their OpenAI is when pulling off with their OpenAI is when pulling off with their models and the capabilities that they models and the capabilities that they models and the capabilities that they have demonstrated on really really low have demonstrated on really really low have demonstrated on really really low reasoning budgets are insane. And I reasoning budgets are insane. And I reasoning budgets are insane. And I think there's a lot we can learn about think there's a lot we can learn about think there's a lot we can learn about if we try to figure out how they are if we try to figure out how they are if we try to figure out how they are doing that. From what these tokens are doing that. From what these tokens are doing that. From what these tokens are even being used for to how they're even being used for to how they're even being used for to how they're affecting the intelligence of the models affecting the intelligence of the models affecting the intelligence of the models to how they affect the actual quality of to how they affect the actual quality of to how they affect the actual quality of the outputs we are getting. This isn't the outputs we are getting. This isn't the outputs we are getting. This isn't as simple as a number on a chart. It as simple as a number on a chart. It as simple as a number on a chart. It goes much deeper and I'm going to do my goes much deeper and I'm going to do my goes much deeper and I'm going to do my best to explain why after a really quick best to explain why after a really quick best to explain why after a really quick word from today's sponsor. While AI has word from today's sponsor. While AI has word from today's sponsor. While AI has gotten much better at writing code, I gotten much better at writing code, I gotten much better at writing code, I never thought I would see the day where never thought I would see the day where never thought I would see the day where I could actually use a computer I could actually use a computer I could actually use a computer properly. Sure, if it can write properly. Sure, if it can write properly. Sure, if it can write commands, it can do a lot, but what commands, it can do a lot, but what commands, it can do a lot, but what happens when it needs to navigate a happens when it needs to navigate a happens when it needs to navigate a complex webpage and click buttons and complex webpage and click buttons and complex webpage and click buttons and all of those types of things. It turns

  2. all of those types of things. It turns all of those types of things. It turns out OpenAI and Anthropic have actually out OpenAI and Anthropic have actually out OpenAI and Anthropic have actually put a lot of time in and made agents put a lot of time in and made agents put a lot of time in and made agents really good at that. If you don't really good at that. If you don't really good at that. If you don't believe me, open up Codex and tell it to believe me, open up Codex and tell it to believe me, open up Codex and tell it to go to configure something on the Google go to configure something on the Google go to configure something on the Google Cloud dashboard for you. I never thought Cloud dashboard for you. I never thought Cloud dashboard for you. I never thought I would see the day it works, but it I would see the day it works, but it I would see the day it works, but it does now. But what does this mean? It does now. But what does this mean? It does now. But what does this mean? It means your agents need a browser in means your agents need a browser in means your agents need a browser in order to take advantage of that order to take advantage of that order to take advantage of that capability. And sure, it's nice and fun capability. And sure, it's nice and fun capability. And sure, it's nice and fun when you're running it on your MacBook, when you're running it on your MacBook, when you're running it on your MacBook, but what happens when you want to put but what happens when you want to put but what happens when you want to put out a real service to real users that out a real service to real users that out a real service to real users that requires your agents be able to navigate requires your agents be able to navigate requires your agents be able to navigate the web? Well, I hope you know about the web? Well, I hope you know about the web? Well, I hope you know about today's sponsor before you build that today's sponsor before you build that today's sponsor before you build that because BrowserBase makes it 100 times because BrowserBase makes it 100 times because BrowserBase makes it 100 times easier. These guys built the perfect easier. These guys built the perfect easier. These guys built the perfect browser for your agents and they made it browser for your agents and they made it browser for your agents and they made it as easy as possible to set up. You can as easy as possible to set up. You can as easy as possible to set up. You can literally just click the setup for literally just click the setup for literally just click the setup for agents, paste it in your agent of agents, paste it in your agent of agents, paste it in your agent of choice, and then you're good to go. choice, and then you're good to go. choice, and then you're good to go. Because BrowserBase provides everything Because BrowserBase provides everything Because BrowserBase provides everything from the SDK to the web infrastructure from the SDK to the web infrastructure from the SDK to the web infrastructure allowing your agents to use the allowing your agents to use the allowing your agents to use the internet. Fun fact, did you know 85% of internet. Fun fact, did you know 85% of internet. Fun fact, did you know 85% of the web isn't exposed over traditional the web isn't exposed over traditional the web isn't exposed over traditional APIs? They're just exposed for specific APIs? They're just exposed for specific APIs? They're just exposed for specific web services when you use them. That 85% web services when you use them. That 85% web services when you use them. That 85% of the web is suddenly accessible when of the web is suddenly accessible when of the web is suddenly accessible when you use a service like BrowserBase you use a service like BrowserBase you use a service like BrowserBase because the agents can actually explore because the agents can actually explore because the agents can actually explore the pages directly. You can use this for the pages directly. You can use this for the pages directly. You can use this for your own services to catch bugs. You can your own services to catch bugs. You can your own services to catch bugs. You can use this for platforms you're building use this for platforms you're building use this for platforms you're building on top of to integrate them better and on top of to integrate them better and on top of to integrate them better and so much more. But if the things you're so much more. But if the things you're so much more. But if the things you're doing are simpler and you just want to, doing are simpler and you just want to, doing are simpler and you just want to, I don't know, grab the content of a page I don't know, grab the content of a page I don't know, grab the content of a page in markdown or go search the web for in markdown or go search the web for in markdown or go search the web for certain content, they now provide that, certain content, they now provide that, certain content, they now provide that, too, with their super simple search API too, with their super simple search API too, with their super simple search API and their fetch API that let you get and their fetch API that let you get and their fetch API that let you get results back in HTML, JSON, or markdown, results back in HTML, JSON, or markdown, results back in HTML, JSON, or markdown, the greatest language ever. Jokes aside, the greatest language ever. Jokes aside, the greatest language ever. Jokes aside, they built something awesome here. And they built something awesome here. And they built something awesome here. And if you want to see your agents really if you want to see your agents really if you want to see your agents really use the web, check them out now. It's use the web, check them out now. It's use the web, check them out now. It's soydev.link/browserbase.

  3. soydev.link/browserbase. soydev.link/browserbase. Before we go too deep into the why, I Before we go too deep into the why, I Before we go too deep into the why, I want to make sure we understand what want to make sure we understand what want to make sure we understand what we're even looking at here. I also want we're even looking at here. I also want we're even looking at here. I also want to turn back on the Gemini models to turn back on the Gemini models to turn back on the Gemini models because it makes the chart much funnier. because it makes the chart much funnier. because it makes the chart much funnier. OpenAI can score meaningfully higher on OpenAI can score meaningfully higher on OpenAI can score meaningfully higher on Deep SWE with 20K tokens than the best Deep SWE with 20K tokens than the best Deep SWE with 20K tokens than the best Gemini models can score with 270K Gemini models can score with 270K Gemini models can score with 270K tokens. That is a 12 to 14x increase in tokens. That is a 12 to 14x increase in tokens. That is a 12 to 14x increase in the number of tokens used to solve the the number of tokens used to solve the the number of tokens used to solve the problems. And the score is almost half problems. And the score is almost half problems. And the score is almost half what the score was for the OpenAI what the score was for the OpenAI what the score was for the OpenAI equivalent. Insanity. The inefficiency equivalent. Insanity. The inefficiency equivalent. Insanity. The inefficiency of the Gemini models is crazy. And I of the Gemini models is crazy. And I of the Gemini models is crazy. And I think people get too fixated on token think people get too fixated on token think people get too fixated on token costs and not enough on efficiency. So, costs and not enough on efficiency. So, costs and not enough on efficiency. So, we're going to do a quick breakdown of we're going to do a quick breakdown of we're going to do a quick breakdown of all of that so you can better understand all of that so you can better understand all of that so you can better understand because I I'll be frank, I'm tired of because I I'll be frank, I'm tired of because I I'll be frank, I'm tired of seeing people in the comment section not seeing people in the comment section not seeing people in the comment section not understanding how these things are understanding how these things are understanding how these things are built. In order for this all to make built. In order for this all to make built. In order for this all to make sense, we first need to talk about the sense, we first need to talk about the sense, we first need to talk about the two types of tokens. I'm going to put a two types of tokens. I'm going to put a two types of tokens. I'm going to put a star next to two cuz it's not as simple star next to two cuz it's not as simple star next to two cuz it's not as simple as it might seem. We have input and as it might seem. We have input and as it might seem. We have input and output. A token is a chunk of text the output. A token is a chunk of text the output. A token is a chunk of text the way that the models break them up in way that the models break them up in way that the models break them up in order to process information. Input order to process information. Input order to process information. Input tokens are the things that you submit to tokens are the things that you submit to tokens are the things that you submit to the model, whether that is a prompt you the model, whether that is a prompt you the model, whether that is a prompt you wrote, text that it's reading from your wrote, text that it's reading from your wrote, text that it's reading from your code base, something that it's running code base, something that it's running code base, something that it's running as a command that it gets text out of as a command that it gets text out of as a command that it gets text out of that it then puts back in. All of the that it then puts back in. All of the that it then puts back in. All of the content being ingested by the model is content being ingested by the model is content being ingested by the model is tokenized and broken down into these tokenized and broken down into these tokenized and broken down into these input tokens. Even when you're ingesting input tokens. Even when you're ingesting input tokens. Even when you're ingesting PDFs or images or any other format, it PDFs or images or any other format, it PDFs or images or any other format, it gets broken down into these tokens.

  4. gets broken down into these tokens. gets broken down into these tokens. Output tokens are what the model Output tokens are what the model Output tokens are what the model generates. Those are the things that it generates. Those are the things that it generates. Those are the things that it uses to create a response as well as the uses to create a response as well as the uses to create a response as well as the response itself. As I said though, it's response itself. As I said though, it's response itself. As I said though, it's not as simple as two types because both not as simple as two types because both not as simple as two types because both of these break down a little bit. Input of these break down a little bit. Input of these break down a little bit. Input isn't as simple as just input tokens isn't as simple as just input tokens isn't as simple as just input tokens because there is uncached and cached because there is uncached and cached because there is uncached and cached inputs. And output is also not as simple inputs. And output is also not as simple inputs. And output is also not as simple because there is both reasoning tokens because there is both reasoning tokens because there is both reasoning tokens and the actual output tokens that are and the actual output tokens that are and the actual output tokens that are the content you're reading or the code the content you're reading or the code the content you're reading or the code that it's writing. So, let's start by that it's writing. So, let's start by that it's writing. So, let's start by going through a fake chat history. Let's going through a fake chat history. Let's going through a fake chat history. Let's say we start with some simple prompt say we start with some simple prompt say we start with some simple prompt like, "Hey, I want to add dark mode to like, "Hey, I want to add dark mode to like, "Hey, I want to add dark mode to my app." Then the model responds to my app." Then the model responds to my app." Then the model responds to something like, "Okay, first I have to something like, "Okay, first I have to something like, "Okay, first I have to understand how styles are handled in understand how styles are handled in understand how styles are handled in this app." Then it goes and fires some this app." Then it goes and fires some this app." Then it goes and fires some grep request or a find in order to read grep request or a find in order to read grep request or a find in order to read the different CSS files and other style the different CSS files and other style the different CSS files and other style related files in your project. Maybe it related files in your project. Maybe it related files in your project. Maybe it starts by just reading package.json and starts by just reading package.json and starts by just reading package.json and then it gets all of the content from the then it gets all of the content from the then it gets all of the content from the package.json file. I'm intentionally package.json file. I'm intentionally package.json file. I'm intentionally putting the package.json contents on the putting the package.json contents on the putting the package.json contents on the right side here because traditionally right side here because traditionally right side here because traditionally this is what we think of as input this is what we think of as input this is what we think of as input tokens, but more importantly, this is tokens, but more importantly, this is tokens, but more importantly, this is the context that the model is getting.

  5. the context that the model is getting. the context that the model is getting. So, when you make a prompt, the model So, when you make a prompt, the model So, when you make a prompt, the model starts generating, but it then realizes starts generating, but it then realizes starts generating, but it then realizes it needs more information, so it ends it needs more information, so it ends it needs more information, so it ends with a tool call and that tool call gets with a tool call and that tool call gets with a tool call and that tool call gets run on whatever machine you're prompting run on whatever machine you're prompting run on whatever machine you're prompting from or if you're doing this in a cloud from or if you're doing this in a cloud from or if you're doing this in a cloud environment or whatever in order to run environment or whatever in order to run environment or whatever in order to run the tool to get additional context that the tool to get additional context that the tool to get additional context that ends up becoming part of this history ends up becoming part of this history ends up becoming part of this history that the model uses to generate the next that the model uses to generate the next that the model uses to generate the next thing. So now that the model has the thing. So now that the model has the thing. So now that the model has the context of the package.json, maybe it context of the package.json, maybe it context of the package.json, maybe it sees that this is a Tailwind project and sees that this is a Tailwind project and sees that this is a Tailwind project and it says, "This project is using it says, "This project is using it says, "This project is using Tailwind. That means I have to" and then Tailwind. That means I have to" and then Tailwind. That means I have to" and then it goes and does whatever it needs to it goes and does whatever it needs to it goes and does whatever it needs to do. When this response gets generated, do. When this response gets generated, do. When this response gets generated, the input isn't just your first message. the input isn't just your first message. the input isn't just your first message. It isn't even just your first message It isn't even just your first message It isn't even just your first message and the package.json. The input is and the package.json. The input is and the package.json. The input is literally everything beforehand. The literally everything beforehand. The literally everything beforehand. The whole history goes into the model and is whole history goes into the model and is whole history goes into the model and is used to skew the weights and skew the used to skew the weights and skew the used to skew the weights and skew the parameters so that they are pointed parameters so that they are pointed parameters so that they are pointed towards what we want to see, which is an towards what we want to see, which is an towards what we want to see, which is an answer that generates this specific dark answer that generates this specific dark answer that generates this specific dark mode that we're requesting. mode that we're requesting. mode that we're requesting. So all output tokens become input tokens So all output tokens become input tokens So all output tokens become input tokens when the job has additional steps to run when the job has additional steps to run when the job has additional steps to run and when a step ends in a tool call, the and when a step ends in a tool call, the and when a step ends in a tool call, the tool call results also become input tool call results also become input tool call results also become input tokens. So you might think, "Oh, I only tokens. So you might think, "Oh, I only tokens. So you might think, "Oh, I only sent two sentences, so there isn't going sent two sentences, so there isn't going sent two sentences, so there isn't going to be that many input tokens." That is to be that many input tokens." That is to be that many input tokens." That is not the case at all.

  6. not the case at all. not the case at all. Anything in your history becomes an Anything in your history becomes an Anything in your history becomes an input token on the next generation, input token on the next generation, input token on the next generation, whether the generation is you sending a whether the generation is you sending a whether the generation is you sending a message requesting follow-ups or doing message requesting follow-ups or doing message requesting follow-ups or doing more work or if it is another tool call more work or if it is another tool call more work or if it is another tool call the model is doing in order to do the model is doing in order to do the model is doing in order to do additional stuff. additional stuff. additional stuff. But if we were to re-ingest all of this But if we were to re-ingest all of this But if we were to re-ingest all of this every single time, it would be super every single time, it would be super every single time, it would be super inefficient. Like that would just be inefficient. Like that would just be inefficient. Like that would just be absurdly expensive. So ideally we're not absurdly expensive. So ideally we're not absurdly expensive. So ideally we're not doing that. There are strategies to doing that. There are strategies to doing that. There are strategies to reduce how expensive all of this is. The reduce how expensive all of this is. The reduce how expensive all of this is. The one that we're all probably most one that we're all probably most one that we're all probably most familiar with is compaction where you familiar with is compaction where you familiar with is compaction where you ask the model to summarize everything ask the model to summarize everything ask the model to summarize everything that happened earlier and instead of that happened earlier and instead of that happened earlier and instead of this being the thousands of tokens it this being the thousands of tokens it this being the thousands of tokens it might be right now, it could become a might be right now, it could become a might be right now, it could become a much smaller amount cuz the model much smaller amount cuz the model much smaller amount cuz the model ingests it once, summarizes it, and now ingests it once, summarizes it, and now ingests it once, summarizes it, and now that is the history going forward. If that is the history going forward. If that is the history going forward. If you ever wonder what compaction was, you ever wonder what compaction was, you ever wonder what compaction was, that is it. It takes your history, it that is it. It takes your history, it that is it. It takes your history, it runs it through the model once and then runs it through the model once and then runs it through the model once and then generates a smaller thing that mostly generates a smaller thing that mostly generates a smaller thing that mostly represents the intent of the previous represents the intent of the previous represents the intent of the previous messages. That's also why you might lose messages. That's also why you might lose messages. That's also why you might lose details when compaction happens or it details when compaction happens or it details when compaction happens or it might lose track of what it's doing might lose track of what it's doing might lose track of what it's doing because the details that were much more because the details that were much more because the details that were much more specific before in that longer history specific before in that longer history specific before in that longer history get compacted into a smaller one. But get compacted into a smaller one. But get compacted into a smaller one. But that's only one of the two things that that's only one of the two things that that's only one of the two things that allows for your history to not be super allows for your history to not be super allows for your history to not be super expensive. The other is caching, which expensive. The other is caching, which expensive. The other is caching, which is a strategy most of the labs put in in is a strategy most of the labs put in in is a strategy most of the labs put in in some way, some in much more effective some way, some in much more effective some way, some in much more effective and easy to implement ways than others, and easy to implement ways than others, and easy to implement ways than others, where they will keep track of how your where they will keep track of how your where they will keep track of how your input affected the model so they don't input affected the model so they don't input affected the model so they don't have to regenerate that every single have to regenerate that every single have to regenerate that every single time a new thing comes in. So for time a new thing comes in. So for time a new thing comes in. So for example here, when it makes the read example here, when it makes the read example here, when it makes the read package JSON request, it knows what the package JSON request, it knows what the package JSON request, it knows what the state of the model was when it got state of the model was when it got state of the model was when it got there.

  7. there. there. So it saves that and then when this new So it saves that and then when this new So it saves that and then when this new piece comes in, all it has to piece comes in, all it has to piece comes in, all it has to recalculate is everything before when it recalculate is everything before when it recalculate is everything before when it cached cached cached and the stuff that has been added since. and the stuff that has been added since. and the stuff that has been added since. And now I can take this additional And now I can take this additional And now I can take this additional piece, add it to the cache, and then the piece, add it to the cache, and then the piece, add it to the cache, and then the next generation doesn't have to do all next generation doesn't have to do all next generation doesn't have to do all of that calculation again. So the two of that calculation again. So the two of that calculation again. So the two types of input tokens, to be very types of input tokens, to be very types of input tokens, to be very specific here, specific here, specific here, are cached and uncached. And if we look are cached and uncached. And if we look are cached and uncached. And if we look at a model like GPT-5.5, the pricing on at a model like GPT-5.5, the pricing on at a model like GPT-5.5, the pricing on a million input tokens is $5, but if a million input tokens is $5, but if a million input tokens is $5, but if they're cached, it's only 50 cents. So they're cached, it's only 50 cents. So they're cached, it's only 50 cents. So if you're caching efficiently, your if you're caching efficiently, your if you're caching efficiently, your input tokens get much cheaper, but the input tokens get much cheaper, but the input tokens get much cheaper, but the output tokens are still really output tokens are still really output tokens are still really expensive. I said there was two types of expensive. I said there was two types of expensive. I said there was two types of output tokens, but we're only seeing one output tokens, but we're only seeing one output tokens, but we're only seeing one here. And no, I'm not referring to short here. And no, I'm not referring to short here. And no, I'm not referring to short versus long context. I'm going to ignore versus long context. I'm going to ignore versus long context. I'm going to ignore that for now. It just makes things too that for now. It just makes things too that for now. It just makes things too complex. complex. complex. The two types of output tokens are The two types of output tokens are The two types of output tokens are reasoning reasoning reasoning and output. and output. and output. I know one of those is contradictory. I know one of those is contradictory. I know one of those is contradictory. There's output output tokens and There's output output tokens and There's output output tokens and reasoning output tokens, but it's the reasoning output tokens, but it's the reasoning output tokens, but it's the easiest way I can think to explain it. easiest way I can think to explain it. easiest way I can think to explain it. The reasoning tokens are the ones the The reasoning tokens are the ones the The reasoning tokens are the ones the model uses to try and improve its answer model uses to try and improve its answer model uses to try and improve its answer before giving it to you. It's the before giving it to you. It's the before giving it to you. It's the thinking tokens, the model taking time thinking tokens, the model taking time thinking tokens, the model taking time to decide what it wants to do before it to decide what it wants to do before it to decide what it wants to do before it responds. Back in my day, when I was responds. Back in my day, when I was responds. Back in my day, when I was using AI, models didn't have the ability using AI, models didn't have the ability using AI, models didn't have the ability to think. They would just start spitting to think. They would just start spitting to think. They would just start spitting out text immediately.

  8. out text immediately. out text immediately. But that means that if they caught But that means that if they caught But that means that if they caught themselves making a mistake halfway themselves making a mistake halfway themselves making a mistake halfway through, they were screwed. It means if through, they were screwed. It means if through, they were screwed. It means if they wanted to explore other options, they wanted to explore other options, they wanted to explore other options, they would have to tell you all of those they would have to tell you all of those they would have to tell you all of those and it would just make the output and it would just make the output and it would just make the output unreadable. And OpenAI back in the '01 unreadable. And OpenAI back in the '01 unreadable. And OpenAI back in the '01 days came up with this idea of letting days came up with this idea of letting days came up with this idea of letting the model talk to itself effectively to the model talk to itself effectively to the model talk to itself effectively to improve the quality of the answer before improve the quality of the answer before improve the quality of the answer before it became your problem. And that idea of it became your problem. And that idea of it became your problem. And that idea of reasoning, letting the model talk to reasoning, letting the model talk to reasoning, letting the model talk to itself before you get the response, itself before you get the response, itself before you get the response, massively increased the quality of the massively increased the quality of the massively increased the quality of the answers that they were seeing from the answers that they were seeing from the answers that they were seeing from the model, which resulted in a massive surge model, which resulted in a massive surge model, which resulted in a massive surge in the capability of these models for in the capability of these models for in the capability of these models for doing difficult thinking work from doing difficult thinking work from doing difficult thinking work from coding to science to engineering to many coding to science to engineering to many coding to science to engineering to many other things. Funny enough, even naming other things. Funny enough, even naming other things. Funny enough, even naming skateboard tricks benefited greatly from skateboard tricks benefited greatly from skateboard tricks benefited greatly from this change, letting the model talk to this change, letting the model talk to this change, letting the model talk to itself. But this had a cost, a literal itself. But this had a cost, a literal itself. But this had a cost, a literal cost. The number of tokens that you cost. The number of tokens that you cost. The number of tokens that you would get in the responses went up would get in the responses went up would get in the responses went up massively. For example, let's take massively. For example, let's take massively. For example, let's take something like Skatebench where I something like Skatebench where I something like Skatebench where I describe a skateboard trick and it names describe a skateboard trick and it names describe a skateboard trick and it names the trick. The named trick is often just the trick. The named trick is often just the trick. The named trick is often just like five or 10 tokens at most cuz it's like five or 10 tokens at most cuz it's like five or 10 tokens at most cuz it's just a couple words like switch backside just a couple words like switch backside just a couple words like switch backside kickflip is maybe six tokens. But if it kickflip is maybe six tokens. But if it kickflip is maybe six tokens. But if it had to think about that first, it could had to think about that first, it could had to think about that first, it could generate thousands if not tens of generate thousands if not tens of generate thousands if not tens of thousands of tokens as it talks to thousands of tokens as it talks to thousands of tokens as it talks to itself trying to decide what it wants to itself trying to decide what it wants to itself trying to decide what it wants to do. I'm going to give a simple example do. I'm going to give a simple example do. I'm going to give a simple example here using GLM 52 because open weight here using GLM 52 because open weight here using GLM 52 because open weight models will show you their reasoning, models will show you their reasoning, models will show you their reasoning, which makes it easy to demonstrate. I'm which makes it easy to demonstrate. I'm which makes it easy to demonstrate. I'm going to ask it to name a skate trick.

  9. going to ask it to name a skate trick. going to ask it to name a skate trick. The skater is skating switch stance. The skater is skating switch stance. The skater is skating switch stance. They pop on their tail, flip the board They pop on their tail, flip the board They pop on their tail, flip the board in the kickflip direction, and the board in the kickflip direction, and the board in the kickflip direction, and the board spins 180° backside while also flipping. spins 180° backside while also flipping. spins 180° backside while also flipping. The skater does not spin. This trick The skater does not spin. This trick The skater does not spin. This trick would be a switch varial kickflip or would be a switch varial kickflip or would be a switch varial kickflip or switch variflip. We'll see if it can switch variflip. We'll see if it can switch variflip. We'll see if it can name it properly. It's funny that our name it properly. It's funny that our name it properly. It's funny that our summary model did not name it correctly. summary model did not name it correctly. summary model did not name it correctly. This named the trick correctly. A This named the trick correctly. A This named the trick correctly. A skateboarder would call that a switch skateboarder would call that a switch skateboarder would call that a switch varial kickflip or more specifically a varial kickflip or more specifically a varial kickflip or more specifically a switch backside varial kickflip. No, we switch backside varial kickflip. No, we switch backside varial kickflip. No, we wouldn't call it that. We'd just call it wouldn't call it that. We'd just call it wouldn't call it that. We'd just call it switch varial kickflip. switch varial kickflip. switch varial kickflip. So it named this correctly. So it named this correctly. So it named this correctly. The actual response here, like this The actual response here, like this The actual response here, like this part, is probably only a few hundred part, is probably only a few hundred part, is probably only a few hundred tokens at most. I'm guessing under a tokens at most. I'm guessing under a tokens at most. I'm guessing under a hundred. But it had to reason for a bit hundred. But it had to reason for a bit hundred. But it had to reason for a bit there. Like it took its time. And you there. Like it took its time. And you there. Like it took its time. And you can see here what it did in the can see here what it did in the can see here what it did in the reasoning. reasoning. reasoning. It analyzed the request. The user wants It analyzed the request. The user wants It analyzed the request. The user wants me to name a skateboard trick the way a me to name a skateboard trick the way a me to name a skateboard trick the way a skater would. The conditions, and it skater would. The conditions, and it skater would. The conditions, and it listed all my conditions. Deconstruct listed all my conditions. Deconstruct listed all my conditions. Deconstruct the skateboard trick components. You can the skateboard trick components. You can the skateboard trick components. You can see all of the things it did here. This see all of the things it did here. This see all of the things it did here. This is a very interesting reasoning trace is a very interesting reasoning trace is a very interesting reasoning trace actually. So, if we compare this to like actually. So, if we compare this to like actually. So, if we compare this to like a dumber model, I don't know. Let's take a dumber model, I don't know. Let's take a dumber model, I don't know. Let's take any Quen model.

  10. any Quen model. any Quen model. Here the reasoning's going to be a lot Here the reasoning's going to be a lot Here the reasoning's going to be a lot more just text walls. Okay, so the user more just text walls. Okay, so the user more just text walls. Okay, so the user wants me to name a skateboard trick as a wants me to name a skateboard trick as a wants me to name a skateboard trick as a skateboarder would. Let's break down the skateboarder would. Let's break down the skateboarder would. Let's break down the description. And it got entirely wrong. Backside And it got entirely wrong. Backside kickflip 180. But it generated a kickflip 180. But it generated a kickflip 180. But it generated a shitload of tokens talking to itself shitload of tokens talking to itself shitload of tokens talking to itself trying to figure out what it would be. trying to figure out what it would be. trying to figure out what it would be. But if I did this with reasoning off, it But if I did this with reasoning off, it But if I did this with reasoning off, it would end up being way fewer tokens. would end up being way fewer tokens. would end up being way fewer tokens. This was 1,307 tokens. Same exact prompt This was 1,307 tokens. Same exact prompt This was 1,307 tokens. Same exact prompt sent on instant mode. Ends up being 745 tokens. I don't know Ends up being 745 tokens. I don't know how it ended up being that much. It how it ended up being that much. It how it ended up being that much. It should have been even less than that, should have been even less than that, should have been even less than that, but you get the idea. The reasoning is but you get the idea. The reasoning is but you get the idea. The reasoning is letting the model talk to itself as it letting the model talk to itself as it letting the model talk to itself as it names the trick. But you'll notice names the trick. But you'll notice names the trick. But you'll notice something weird if I switch over to GPT something weird if I switch over to GPT something weird if I switch over to GPT 55 and ask the same thing. First off, 55 and ask the same thing. First off, 55 and ask the same thing. First off, you'll notice the number of tokens is you'll notice the number of tokens is you'll notice the number of tokens is very small, only 342 tokens. It got the very small, only 342 tokens. It got the very small, only 342 tokens. It got the name right. It describes why the name is name right. It describes why the name is name right. It describes why the name is right. But in the reasoning trace you'll right. But in the reasoning trace you'll right. But in the reasoning trace you'll see this is not what it actually see this is not what it actually see this is not what it actually reasoned. Here's where things start to reasoned. Here's where things start to reasoned. Here's where things start to get a little mucky. The frontier labs do get a little mucky. The frontier labs do get a little mucky. The frontier labs do not actually give you the raw reasoning not actually give you the raw reasoning not actually give you the raw reasoning tokens. They give you summaries of the tokens. They give you summaries of the tokens. They give you summaries of the reasoning. The reason they do this is reasoning. The reason they do this is reasoning. The reason they do this is they don't want all of that data getting they don't want all of that data getting they don't want all of that data getting out for other companies to be able to out for other companies to be able to out for other companies to be able to copy from and create models that behave copy from and create models that behave copy from and create models that behave similarly. Because if you have all of similarly. Because if you have all of similarly. Because if you have all of the data of exactly how a GPT model came the data of exactly how a GPT model came the data of exactly how a GPT model came to the conclusion it came to, you can to the conclusion it came to, you can to the conclusion it came to, you can use that to make your own models behave use that to make your own models behave use that to make your own models behave similarly. Anthropic used to give us the

  11. similarly. Anthropic used to give us the similarly. Anthropic used to give us the full reasoning trace back in the Sonnet full reasoning trace back in the Sonnet full reasoning trace back in the Sonnet 3.5 and 3.6 days, but they have since 3.5 and 3.6 days, but they have since 3.5 and 3.6 days, but they have since stopped and also do summaries just like stopped and also do summaries just like stopped and also do summaries just like we're seeing here. This means we're kind we're seeing here. This means we're kind we're seeing here. This means we're kind of in a weird position if we try to of in a weird position if we try to of in a weird position if we try to analyze why the OpenAI models are able analyze why the OpenAI models are able analyze why the OpenAI models are able to be so efficient because we can't to be so efficient because we can't to be so efficient because we can't actually see what the reasoning traces actually see what the reasoning traces actually see what the reasoning traces are. We can't see what's going on are. We can't see what's going on are. We can't see what's going on underneath because they're not showing underneath because they're not showing underneath because they're not showing us that. They are just giving us us that. They are just giving us us that. They are just giving us summaries. Also, somebody seems to think summaries. Also, somebody seems to think summaries. Also, somebody seems to think Google gave us the raw reasoning traces. Google gave us the raw reasoning traces. Google gave us the raw reasoning traces. Google has never once gave us raw Google has never once gave us raw Google has never once gave us raw reasoning traces. They didn't even give reasoning traces. They didn't even give reasoning traces. They didn't even give us summaries until a bit after Gemini us summaries until a bit after Gemini us summaries until a bit after Gemini 2.5 Pro came out because a lot of 2.5 Pro came out because a lot of 2.5 Pro came out because a lot of people, myself included, complained so people, myself included, complained so people, myself included, complained so much about how annoying it was to not much about how annoying it was to not much about how annoying it was to not have them to show the user what the have them to show the user what the have them to show the user what the model was doing. But, there's a couple model was doing. But, there's a couple model was doing. But, there's a couple more layers we need to understand here more layers we need to understand here more layers we need to understand here before we go further when we talk about before we go further when we talk about before we go further when we talk about these reasoning traces and the things these reasoning traces and the things these reasoning traces and the things the model is doing. The summaries the the model is doing. The summaries the the model is doing. The summaries the model gives us are useful to get a rough model gives us are useful to get a rough model gives us are useful to get a rough idea of how the model got where it did, idea of how the model got where it did, idea of how the model got where it did, but they aren't the actual thinking the but they aren't the actual thinking the but they aren't the actual thinking the model did. They're not the actual model did. They're not the actual model did. They're not the actual information the model found when it was information the model found when it was information the model found when it was thinking or the other stuff it did thinking or the other stuff it did thinking or the other stuff it did throughout its time reasoning. But, the throughout its time reasoning. But, the throughout its time reasoning. But, the actual data is useful for the model when actual data is useful for the model when actual data is useful for the model when it generates the next step. So, let's it generates the next step. So, let's it generates the next step. So, let's say, going back to this example, that say, going back to this example, that say, going back to this example, that when the model was asked, "I want to add when the model was asked, "I want to add when the model was asked, "I want to add dark mode to my app." This is an example dark mode to my app." This is an example dark mode to my app." This is an example of a reasoning trace the model might do of a reasoning trace the model might do of a reasoning trace the model might do after getting this prompt.

  12. after getting this prompt. after getting this prompt. This data is useful in future steps. The This data is useful in future steps. The This data is useful in future steps. The model might be confused as to why it model might be confused as to why it model might be confused as to why it assumed the project was JS, but if this assumed the project was JS, but if this assumed the project was JS, but if this exists in the history and the model can exists in the history and the model can exists in the history and the model can ingest this, it is more likely to ingest this, it is more likely to ingest this, it is more likely to understand on the next step why it is understand on the next step why it is understand on the next step why it is where it is and what it should do next. where it is and what it should do next. where it is and what it should do next. I hope you can imagine the different I hope you can imagine the different I hope you can imagine the different things the model might think but not say things the model might think but not say things the model might think but not say that are useful in the history. Where if that are useful in the history. Where if that are useful in the history. Where if it realizes these files aren't useful, it realizes these files aren't useful, it realizes these files aren't useful, it shouldn't have to realize that every it shouldn't have to realize that every it shouldn't have to realize that every single time it does the next step. This single time it does the next step. This single time it does the next step. This information stays useful throughout the information stays useful throughout the information stays useful throughout the run, so we would want the model to have run, so we would want the model to have run, so we would want the model to have that, but it's not actually in our that, but it's not actually in our that, but it's not actually in our machine because they don't send it down machine because they don't send it down machine because they don't send it down the wire, so they have to save this on the wire, so they have to save this on the wire, so they have to save this on their end in order to populate the their end in order to populate the their end in order to populate the history properly. I've talked about this history properly. I've talked about this history properly. I've talked about this a bit in other videos about context a bit in other videos about context a bit in other videos about context management and how these histories work. management and how these histories work. management and how these histories work. I've also talked about this extensively I've also talked about this extensively I've also talked about this extensively in my video about how OpenAI uses web in my video about how OpenAI uses web in my video about how OpenAI uses web sockets now for their inference. I think sockets now for their inference. I think sockets now for their inference. I think it's one of my best technical breakdowns it's one of my best technical breakdowns it's one of my best technical breakdowns I've done recently, worth checking out I've done recently, worth checking out I've done recently, worth checking out if you want more info on all of this.

  13. if you want more info on all of this. if you want more info on all of this. The point I'm trying to make here is The point I'm trying to make here is The point I'm trying to make here is that if you want to increase the that if you want to increase the that if you want to increase the efficiency of the model, the best place efficiency of the model, the best place efficiency of the model, the best place isn't to change how short the answers isn't to change how short the answers isn't to change how short the answers are, it isn't to prompt better, although are, it isn't to prompt better, although are, it isn't to prompt better, although that does help. The best thing you can that does help. The best thing you can that does help. The best thing you can do is reduce how much thinking is going do is reduce how much thinking is going do is reduce how much thinking is going on. The fewer tokens that occur here, on. The fewer tokens that occur here, on. The fewer tokens that occur here, the fewer tokens that are used overall, the fewer tokens that are used overall, the fewer tokens that are used overall, both because it's generating fewer both because it's generating fewer both because it's generating fewer tokens in its responses, but also tokens in its responses, but also tokens in its responses, but also because it doesn't have to re-ingest because it doesn't have to re-ingest because it doesn't have to re-ingest those same tokens on every additional those same tokens on every additional those same tokens on every additional step throughout. Reducing the number of step throughout. Reducing the number of step throughout. Reducing the number of tokens in reasoning is an exponential tokens in reasoning is an exponential tokens in reasoning is an exponential decrease in the amount of tokens used decrease in the amount of tokens used decrease in the amount of tokens used overall. So, how did OpenAI do it? overall. So, how did OpenAI do it? overall. So, how did OpenAI do it? Well, again, as I mentioned before, Well, again, as I mentioned before, Well, again, as I mentioned before, other labs don't seem as focused on other labs don't seem as focused on other labs don't seem as focused on these improvements to reasoning these improvements to reasoning these improvements to reasoning efficiency. Fable is better than I would efficiency. Fable is better than I would efficiency. Fable is better than I would have expected with medium actually have expected with medium actually have expected with medium actually competing with GPT-55's reasoning competing with GPT-55's reasoning competing with GPT-55's reasoning efficiencies, but X high and max are efficiencies, but X high and max are efficiencies, but X high and max are still over 100k tokens per task for the still over 100k tokens per task for the still over 100k tokens per task for the examples that Deep Mind did. Whereas, examples that Deep Mind did. Whereas, examples that Deep Mind did. Whereas, the OpenAI models are absurdly the OpenAI models are absurdly the OpenAI models are absurdly efficient, all falling under 50k tokens efficient, all falling under 50k tokens efficient, all falling under 50k tokens on the 55 line. First and foremost, I on the 55 line. First and foremost, I on the 55 line. First and foremost, I want to be real here. The biggest reason want to be real here. The biggest reason want to be real here. The biggest reason OpenAI's models are more efficient is OpenAI's models are more efficient is OpenAI's models are more efficient is because they're more focused on because they're more focused on because they're more focused on efficiency. This is a thing they have efficiency. This is a thing they have efficiency. This is a thing they have been working on from the early days, and been working on from the early days, and been working on from the early days, and they have went out of their way to try they have went out of their way to try they have went out of their way to try and make the reasoning as efficient as and make the reasoning as efficient as and make the reasoning as efficient as possible in order to allow them to run possible in order to allow them to run possible in order to allow them to run the model more aggressively, have the the model more aggressively, have the the model more aggressively, have the models do more things for less money, models do more things for less money, models do more things for less money, use less GPUs. They want to be able to use less GPUs. They want to be able to use less GPUs. They want to be able to generate as many good answers as generate as many good answers as generate as many good answers as possible with as little compute as possible with as little compute as possible with as little compute as possible, and making it more efficient

  14. possible, and making it more efficient possible, and making it more efficient benefits them greatly as a result. And benefits them greatly as a result. And benefits them greatly as a result. And when you combine this with the actual when you combine this with the actual when you combine this with the actual cost per token, which is what people get cost per token, which is what people get cost per token, which is what people get way too fixated on, you see why this is way too fixated on, you see why this is way too fixated on, you see why this is so valuable because the costs end up so valuable because the costs end up so valuable because the costs end up being significantly lower for 5.5 than being significantly lower for 5.5 than being significantly lower for 5.5 than other models simply because they're other models simply because they're other models simply because they're doing so many fewer tokens, even though doing so many fewer tokens, even though doing so many fewer tokens, even though the price of the model is higher than it the price of the model is higher than it the price of the model is higher than it was before. If we go back to the pricing was before. If we go back to the pricing was before. If we go back to the pricing chart, they have effectively doubled the chart, they have effectively doubled the chart, they have effectively doubled the price since GPT 5.4, where before it was price since GPT 5.4, where before it was price since GPT 5.4, where before it was $2.50 per mill in and 15 per mill out, $2.50 per mill in and 15 per mill out, $2.50 per mill in and 15 per mill out, now it's 5 per mill in and 30 per mill now it's 5 per mill in and 30 per mill now it's 5 per mill in and 30 per mill out. That doubling doesn't hurt quite as out. That doubling doesn't hurt quite as out. That doubling doesn't hurt quite as bad. It is still more expensive, but bad. It is still more expensive, but bad. It is still more expensive, but it's not as bad as it could be because it's not as bad as it could be because it's not as bad as it could be because 5.5 ended up being so much more token 5.5 ended up being so much more token 5.5 ended up being so much more token efficient for a certain level of efficient for a certain level of efficient for a certain level of intelligence. So, if we look at, for intelligence. So, if we look at, for intelligence. So, if we look at, for example, 5.5 medium, it used under half example, 5.5 medium, it used under half example, 5.5 medium, it used under half as many tokens as 5.4 X high and came as many tokens as 5.4 X high and came as many tokens as 5.4 X high and came out with a higher score, which means its out with a higher score, which means its out with a higher score, which means its cost ended up also being lower at about cost ended up also being lower at about cost ended up also being lower at about half the price of 5.4 X high. If you half the price of 5.4 X high. If you half the price of 5.4 X high. If you compare the cost of 5.5 X high to 5.4 X compare the cost of 5.5 X high to 5.4 X compare the cost of 5.5 X high to 5.4 X high, it goes from 565 to 723. So, yeah, high, it goes from 565 to 723. So, yeah, high, it goes from 565 to 723. So, yeah, 5.5 at its extreme is more expensive, 5.5 at its extreme is more expensive, 5.5 at its extreme is more expensive, but if you're comparing cost relative to but if you're comparing cost relative to but if you're comparing cost relative to a certain level of intelligence, it is a certain level of intelligence, it is a certain level of intelligence, it is indeed going down, and that is largely indeed going down, and that is largely indeed going down, and that is largely due to the efficiency improvements that due to the efficiency improvements that due to the efficiency improvements that OpenAI has been working really hard to OpenAI has been working really hard to OpenAI has been working really hard to attain. We will never know the full attain. We will never know the full attain. We will never know the full details of how they have managed to do details of how they have managed to do details of how they have managed to do this, but we have had some leaks that this, but we have had some leaks that this, but we have had some leaks that show some of them. Before we can show some of them. Before we can show some of them. Before we can understand this, we need to talk a bit understand this, we need to talk a bit understand this, we need to talk a bit about Grug. Grug brain developer not so about Grug. Grug brain developer not so about Grug. Grug brain developer not so smart, but Grug brain developer program

  15. smart, but Grug brain developer program smart, but Grug brain developer program many long year and learn some things, many long year and learn some things, many long year and learn some things, although mostly still confused. Grug although mostly still confused. Grug although mostly still confused. Grug brain developer try collect learns into brain developer try collect learns into brain developer try collect learns into small, easily digestible, and funny small, easily digestible, and funny small, easily digestible, and funny page. Not only for you, young Grug, but page. Not only for you, young Grug, but page. Not only for you, young Grug, but also for him. Because as Grug brain also for him. Because as Grug brain also for him. Because as Grug brain developer get older, he forget important developer get older, he forget important developer get older, he forget important things, like what he had for breakfast things, like what he had for breakfast things, like what he had for breakfast or if put pants on. More simply put by or if put pants on. More simply put by or if put pants on. More simply put by aviator in chat here, aviator in chat here, aviator in chat here, why use many words when few word do why use many words when few word do why use many words when few word do trick? This is a silly way of writing trick? This is a silly way of writing trick? This is a silly way of writing that is intentionally written to feel that is intentionally written to feel that is intentionally written to feel dumb, but also obvious cuz that's the dumb, but also obvious cuz that's the dumb, but also obvious cuz that's the point of the grog-brained developer in point of the grog-brained developer in point of the grog-brained developer in this way of thinking pioneered by this way of thinking pioneered by this way of thinking pioneered by Carson, the writer and creator of HTMX. Carson, the writer and creator of HTMX. Carson, the writer and creator of HTMX. It is meant to show you the smart thing It is meant to show you the smart thing It is meant to show you the smart thing isn't always the thing that has the isn't always the thing that has the isn't always the thing that has the fanciest words and vocabulary and the fanciest words and vocabulary and the fanciest words and vocabulary and the most elaborate write-ups. Sometimes the most elaborate write-ups. Sometimes the most elaborate write-ups. Sometimes the simple stupid thing is the right one. simple stupid thing is the right one. simple stupid thing is the right one. And sometimes you don't need all of And sometimes you don't need all of And sometimes you don't need all of those words to communicate the value. It those words to communicate the value. It those words to communicate the value. It seems like this is one of the many seems like this is one of the many seems like this is one of the many tricks that OpenAI has been employing in tricks that OpenAI has been employing in tricks that OpenAI has been employing in order to make the models more efficient order to make the models more efficient order to make the models more efficient because if we're not seeing the because if we're not seeing the because if we're not seeing the reasoning traces, we don't care what reasoning traces, we don't care what reasoning traces, we don't care what language they're speaking, we don't care language they're speaking, we don't care language they're speaking, we don't care what vowels they're forgetting, we don't what vowels they're forgetting, we don't what vowels they're forgetting, we don't care what words they're omitting, how care what words they're omitting, how care what words they're omitting, how good their vocabulary is. We don't care good their vocabulary is. We don't care good their vocabulary is. We don't care about any of that during the reasoning about any of that during the reasoning about any of that during the reasoning step. We only care about the outputs.

  16. step. We only care about the outputs. step. We only care about the outputs. And if OpenAI has gotten to the point And if OpenAI has gotten to the point And if OpenAI has gotten to the point where they can meaningfully draw a line where they can meaningfully draw a line where they can meaningfully draw a line between the reasoning outputs and the between the reasoning outputs and the between the reasoning outputs and the actual answer output at the end, where actual answer output at the end, where actual answer output at the end, where they can make the model behave one way they can make the model behave one way they can make the model behave one way during reasoning in an entirely during reasoning in an entirely during reasoning in an entirely different way for the actual answer different way for the actual answer different way for the actual answer outputs, they would hypothetically be outputs, they would hypothetically be outputs, they would hypothetically be able to make a model that is much more able to make a model that is much more able to make a model that is much more efficient in reasoning cuz it speaks one efficient in reasoning cuz it speaks one efficient in reasoning cuz it speaks one way then and gets good answers at the way then and gets good answers at the way then and gets good answers at the end cuz it speaks a different way when end cuz it speaks a different way when end cuz it speaks a different way when it outputs them. Sadly though, we'll it outputs them. Sadly though, we'll it outputs them. Sadly though, we'll never know if OpenAI is doing this never know if OpenAI is doing this never know if OpenAI is doing this because those reasoning traces are super because those reasoning traces are super because those reasoning traces are super locked down. They've never leaked ever locked down. They've never leaked ever locked down. They've never leaked ever in history. Oh, is there a leak of a in history. Oh, is there a leak of a in history. Oh, is there a leak of a reasoning trace on my screen right now? reasoning trace on my screen right now? reasoning trace on my screen right now? Turns out there are very few things in Turns out there are very few things in Turns out there are very few things in the non-deterministic world of LLMs that the non-deterministic world of LLMs that the non-deterministic world of LLMs that are as reliable as a company like OpenAI are as reliable as a company like OpenAI are as reliable as a company like OpenAI would hope. And as a result, some of would hope. And as a result, some of would hope. And as a result, some of these reasoning traces have leaked. And these reasoning traces have leaked. And these reasoning traces have leaked. And we can see some very funny chains of we can see some very funny chains of we can see some very funny chains of thought as a result. Need agent kind thought as a result. Need agent kind thought as a result. Need agent kind maybe open hands direct okay. Need just maybe open hands direct okay. Need just maybe open hands direct okay. Need just set tools default in JQ. Need finish set tools default in JQ. Need finish set tools default in JQ. Need finish tool? Default tool name is no finish. tool? Default tool name is no finish. tool? Default tool name is no finish. Finish is tool. The list in system Finish is tool. The list in system Finish is tool. The list in system prompt for delegated had only finish.

  17. prompt for delegated had only finish. prompt for delegated had only finish. Think switch LLM invoke skill because Think switch LLM invoke skill because Think switch LLM invoke skill because tools array empty, but default agent tools array empty, but default agent tools array empty, but default agent maybe always has internal tools. Need maybe always has internal tools. Need maybe always has internal tools. Need add terminal etc. Maybe tool name's add terminal etc. Maybe tool name's add terminal etc. Maybe tool name's okay. Add tools array. Here's another okay. Add tools array. Here's another okay. Add tools array. Here's another one. Use core new nodes. Need infer. one. Use core new nodes. Need infer. one. Use core new nodes. Need infer. Note 35 maybe Note 35 maybe Note 35 maybe thing from the code base. Outputs things thing from the code base. Outputs things thing from the code base. Outputs things from code base. from code base. from code base. UI has conditioning negative. Need add UI has conditioning negative. Need add UI has conditioning negative. Need add VAE encode for images. Try. Try period. VAE encode for images. Try. Try period. VAE encode for images. Try. Try period. This is hilarious, but it's also really This is hilarious, but it's also really This is hilarious, but it's also really efficient. They have trained this model efficient. They have trained this model efficient. They have trained this model and they've R L'd it so hard that try and they've R L'd it so hard that try and they've R L'd it so hard that try period, which is probably one token. period, which is probably one token. period, which is probably one token. Oh, no. The period's another token. What Oh, no. The period's another token. What Oh, no. The period's another token. What a waste. This could be something that a waste. This could be something that a waste. This could be something that just happened as a result of how the just happened as a result of how the just happened as a result of how the model's trained. This could be something model's trained. This could be something model's trained. This could be something they actually intentionally designed for they actually intentionally designed for they actually intentionally designed for where if it didn't have the period, it where if it didn't have the period, it where if it didn't have the period, it might think it needs to keep reasoning. might think it needs to keep reasoning. might think it needs to keep reasoning. But by putting the period in, they are But by putting the period in, they are But by putting the period in, they are forcing the model to stop there and then forcing the model to stop there and then forcing the model to stop there and then start doing whatever it's going to do start doing whatever it's going to do start doing whatever it's going to do next. But we're talking about a company next. But we're talking about a company next. But we're talking about a company here that is trying their hardest to get here that is trying their hardest to get here that is trying their hardest to get this to be as simple as possible to not this to be as simple as possible to not this to be as simple as possible to not waste history in order to make the model waste history in order to make the model waste history in order to make the model generate outputs more effectively generate outputs more effectively generate outputs more effectively because that makes them faster, it makes because that makes them faster, it makes because that makes them faster, it makes them cheaper, it makes them able to them cheaper, it makes them able to them cheaper, it makes them able to reason longer and get more done in a reason longer and get more done in a reason longer and get more done in a given token budget. It's silly, but it given token budget. It's silly, but it given token budget. It's silly, but it works.

  18. works. works. Obviously, side effect one is that when Obviously, side effect one is that when Obviously, side effect one is that when it leaks, it looks really silly and you it leaks, it looks really silly and you it leaks, it looks really silly and you get funny memes on Twitter for it. And get funny memes on Twitter for it. And get funny memes on Twitter for it. And god damn, there have been a lot of funny god damn, there have been a lot of funny god damn, there have been a lot of funny examples that my chat has found for me. examples that my chat has found for me. examples that my chat has found for me. We need adjust UX. Need inspect current We need adjust UX. Need inspect current We need adjust UX. Need inspect current component perhaps parent max XL causing component perhaps parent max XL causing component perhaps parent max XL causing max 2 XL, but parent max WXL so max 2 XL, but parent max WXL so max 2 XL, but parent max WXL so ineffective. Yes, parent W full max WXL. ineffective. Yes, parent W full max WXL. ineffective. Yes, parent W full max WXL. Need baby within card. There are so many Need baby within card. There are so many Need baby within card. There are so many examples of this that they are clearly examples of this that they are clearly examples of this that they are clearly actually doing it. And in order to not actually doing it. And in order to not actually doing it. And in order to not show this to you, they hand this to show this to you, they hand this to show this to you, they hand this to another model and say, "Hey, can you another model and say, "Hey, can you another model and say, "Hey, can you summarize what we thought about here?" summarize what we thought about here?" summarize what we thought about here?" And they show that to you instead. No And they show that to you instead. No And they show that to you instead. No matter how hard they try in training matter how hard they try in training matter how hard they try in training this with between the reasoning and the this with between the reasoning and the this with between the reasoning and the actual answers that are being outputted, actual answers that are being outputted, actual answers that are being outputted, there will be leaks like this. I've seen there will be leaks like this. I've seen there will be leaks like this. I've seen them myself. Almost everyone I know has them myself. Almost everyone I know has them myself. Almost everyone I know has seen this happen like at least once at seen this happen like at least once at seen this happen like at least once at some point during their use of the some point during their use of the some point during their use of the model. That's a negative side effect model. That's a negative side effect model. That's a negative side effect that is only happening because they are that is only happening because they are that is only happening because they are reasoning this way. Versus again, if we reasoning this way. Versus again, if we reasoning this way. Versus again, if we look at the other reasoning traces that look at the other reasoning traces that look at the other reasoning traces that we got from other models earlier, like we got from other models earlier, like we got from other models earlier, like we did here with Qwen, it's writing we did here with Qwen, it's writing we did here with Qwen, it's writing plain English. It's not great English, plain English. It's not great English, plain English. It's not great English, but it's plain English. Or if we look at but it's plain English. Or if we look at but it's plain English. Or if we look at the reasoning trace that we got from the reasoning trace that we got from the reasoning trace that we got from GLM, it's doing bullet points and like GLM, it's doing bullet points and like GLM, it's doing bullet points and like lists here in order to break down the lists here in order to break down the lists here in order to break down the work that it needs to do. Different work that it needs to do. Different work that it needs to do. Different model families reason in all sorts of model families reason in all sorts of model families reason in all sorts of different ways, but OpenAI's models seem different ways, but OpenAI's models seem different ways, but OpenAI's models seem to be reasoning in a very novel way with to be reasoning in a very novel way with to be reasoning in a very novel way with this crazy token efficiency style. And this crazy token efficiency style. And this crazy token efficiency style. And now to drop some potentially hot takes now to drop some potentially hot takes now to drop some potentially hot takes about effects that I think we see as a about effects that I think we see as a about effects that I think we see as a downstream impact of these decisions.

  19. downstream impact of these decisions. downstream impact of these decisions. First and foremost, I think this is a First and foremost, I think this is a First and foremost, I think this is a meaningful part of why Claude is nicer meaningful part of why Claude is nicer meaningful part of why Claude is nicer to talk to because when Claude talks to to talk to because when Claude talks to to talk to because when Claude talks to itself, it is probably talking in plain itself, it is probably talking in plain itself, it is probably talking in plain English. Just seeing how much longer the English. Just seeing how much longer the English. Just seeing how much longer the reasoning traces are, it's probably reasoning traces are, it's probably reasoning traces are, it's probably talking in plain English. It's also talking in plain English. It's also talking in plain English. It's also worth noting that if Anthropic was to worth noting that if Anthropic was to worth noting that if Anthropic was to fix this, their income would go down fix this, their income would go down fix this, their income would go down meaningfully. If developers can do work meaningfully. If developers can do work meaningfully. If developers can do work with the models and fewer tokens are with the models and fewer tokens are with the models and fewer tokens are generated for the same work, that is generated for the same work, that is generated for the same work, that is money they're not making. So I would money they're not making. So I would money they're not making. So I would suspect other labs aren't as interested suspect other labs aren't as interested suspect other labs aren't as interested here because they're not as interested here because they're not as interested here because they're not as interested in lowering how expensive these models in lowering how expensive these models in lowering how expensive these models are for others to run. But on that note, are for others to run. But on that note, are for others to run. But on that note, I think this is also why Claude defaults I think this is also why Claude defaults I think this is also why Claude defaults to 1 mil token context windows. This to 1 mil token context windows. This to 1 mil token context windows. This isn't because it wants to fit bigger isn't because it wants to fit bigger isn't because it wants to fit bigger code bases like so many people seem to code bases like so many people seem to code bases like so many people seem to think. This is because the outputs that think. This is because the outputs that think. This is because the outputs that Claude is generating and the reasoning Claude is generating and the reasoning Claude is generating and the reasoning that Claude is doing is long as that Claude is doing is long as that Claude is doing is long as And in order for that to fit within the And in order for that to fit within the And in order for that to fit within the context window, it needs a bigger context window, it needs a bigger context window, it needs a bigger context window. They're also really bad context window. They're also really bad context window. They're also really bad at compaction, which is part of why that at compaction, which is part of why that at compaction, which is part of why that benefits them greatly. But it's also why benefits them greatly. But it's also why benefits them greatly. But it's also why when you resume old Claude threads, it when you resume old Claude threads, it when you resume old Claude threads, it tells you to not try and keep the whole tells you to not try and keep the whole tells you to not try and keep the whole context, but instead to try and compact context, but instead to try and compact context, but instead to try and compact it. Because they don't keep the it. Because they don't keep the it. Because they don't keep the reasoning trace on their side. In fact, reasoning trace on their side. In fact, reasoning trace on their side. In fact, they clear them out after like 10 or 20 they clear them out after like 10 or 20 they clear them out after like 10 or 20 minutes, if I recall. Because the minutes, if I recall. Because the minutes, if I recall. Because the reasoning tokens are just not efficient reasoning tokens are just not efficient reasoning tokens are just not efficient enough to be worth keeping around, enough to be worth keeping around, enough to be worth keeping around, because your costs would be absurd if because your costs would be absurd if because your costs would be absurd if you did. Meanwhile, the token context you did. Meanwhile, the token context you did. Meanwhile, the token context window with OpenAI models is like 200K window with OpenAI models is like 200K window with OpenAI models is like 200K or less, but they're fitting so much or less, but they're fitting so much or less, but they're fitting so much more in that because they've done all of more in that because they've done all of more in that because they've done all of the crazy grug speak in order to get the crazy grug speak in order to get the crazy grug speak in order to get there in the first place. This is also

  20. there in the first place. This is also there in the first place. This is also why, in my opinion, OpenAI models go off why, in my opinion, OpenAI models go off why, in my opinion, OpenAI models go off the rails faster. Not faster in terms of the rails faster. Not faster in terms of the rails faster. Not faster in terms of the work done or the number of steps, the work done or the number of steps, the work done or the number of steps, but faster in terms of the token context but faster in terms of the token context but faster in terms of the token context window. Because if a million token window. Because if a million token window. Because if a million token context window has all of this grug context window has all of this grug context window has all of this grug speak, keeping track of what actually speak, keeping track of what actually speak, keeping track of what actually matters gets harder than if you have matters gets harder than if you have matters gets harder than if you have properly formatted text with like properly formatted text with like properly formatted text with like headers and bold sections and lists and headers and bold sections and lists and headers and bold sections and lists and formatting to make it clear what was formatting to make it clear what was formatting to make it clear what was going on, not just the path the model going on, not just the path the model going on, not just the path the model was going down. When those tokens stop was going down. When those tokens stop was going down. When those tokens stop being paths and start being histories, being paths and start being histories, being paths and start being histories, this strategy seems to not be as this strategy seems to not be as this strategy seems to not be as effective, and it hurts the ability for effective, and it hurts the ability for effective, and it hurts the ability for these models to do things with the these models to do things with the these models to do things with the really long token contexts, which is why really long token contexts, which is why really long token contexts, which is why I also think OpenAI won't let you use I also think OpenAI won't let you use I also think OpenAI won't let you use the 1M token windows in Codex by the 1M token windows in Codex by the 1M token windows in Codex by default. I haven't tested this myself, default. I haven't tested this myself, default. I haven't tested this myself, but I've been wanting to. I actually but I've been wanting to. I actually but I've been wanting to. I actually told my team I want one of them to like told my team I want one of them to like told my team I want one of them to like go use an API key and test this for me go use an API key and test this for me go use an API key and test this for me for a bit. If anyone in chat has tried for a bit. If anyone in chat has tried for a bit. If anyone in chat has tried it by bringing your own key to Codex and it by bringing your own key to Codex and it by bringing your own key to Codex and turning on a million token context, how turning on a million token context, how turning on a million token context, how does it behave? I don't know, but my does it behave? I don't know, but my does it behave? I don't know, but my assumption is not well, because assumption is not well, because assumption is not well, because otherwise they would have just went and otherwise they would have just went and otherwise they would have just went and turned it on. That's kind of what they turned it on. That's kind of what they turned it on. That's kind of what they like to do over at OpenAI. When a thing like to do over at OpenAI. When a thing like to do over at OpenAI. When a thing is useful, it doesn't matter how many is useful, it doesn't matter how many is useful, it doesn't matter how many tokens it burns, they'll just turn it on tokens it burns, they'll just turn it on tokens it burns, they'll just turn it on for users. Even the fast mode stuff, for users. Even the fast mode stuff, for users. Even the fast mode stuff, like how the fact that fast mode is like how the fact that fast mode is like how the fact that fast mode is available within the subscription tier available within the subscription tier available within the subscription tier on CodeX, and it's not on Claude Code, on CodeX, and it's not on Claude Code, on CodeX, and it's not on Claude Code, you have to pay API prices for it, is you have to pay API prices for it, is you have to pay API prices for it, is absurd, but shows the difference in how absurd, but shows the difference in how absurd, but shows the difference in how much compute availability there is at much compute availability there is at much compute availability there is at these two labs. But, most importantly, these two labs. But, most importantly, these two labs. But, most importantly, what this means is a lot of secret what this means is a lot of secret what this means is a lot of secret sauce. When a model provider tells us a

  21. sauce. When a model provider tells us a sauce. When a model provider tells us a certain number of reasoning tokens were certain number of reasoning tokens were certain number of reasoning tokens were generated, we can't see them. We just generated, we can't see them. We just generated, we can't see them. We just have to trust them. And if the quality have to trust them. And if the quality have to trust them. And if the quality of the responses goes up high enough, of the responses goes up high enough, of the responses goes up high enough, then we'll accept the increased cost then we'll accept the increased cost then we'll accept the increased cost when the model has to reason more, even when the model has to reason more, even when the model has to reason more, even though we can't actually see what's though we can't actually see what's though we can't actually see what's going on. If we want to replicate going on. If we want to replicate going on. If we want to replicate certain behaviors or figure out why a certain behaviors or figure out why a certain behaviors or figure out why a model went down the wrong path, we model went down the wrong path, we model went down the wrong path, we can't, because all of this information can't, because all of this information can't, because all of this information is hidden from us. This also means if is hidden from us. This also means if is hidden from us. This also means if you're a competing lab trying to you're a competing lab trying to you're a competing lab trying to generate models of similar capabilities generate models of similar capabilities generate models of similar capabilities that are similarly efficient, you can't that are similarly efficient, you can't that are similarly efficient, you can't get the data on how they make it so get the data on how they make it so get the data on how they make it so efficient. You don't see the reasoning efficient. You don't see the reasoning efficient. You don't see the reasoning traces. You see the inputs, you see the traces. You see the inputs, you see the traces. You see the inputs, you see the summaries, and you see the answers. And summaries, and you see the answers. And summaries, and you see the answers. And you have to create your own methodology you have to create your own methodology you have to create your own methodology in the reasoning for your model in order in the reasoning for your model in order in the reasoning for your model in order to get similar quality answers. to get similar quality answers. to get similar quality answers. Obviously, a model like GLM-52 is Obviously, a model like GLM-52 is Obviously, a model like GLM-52 is trained on responses from models like trained on responses from models like trained on responses from models like Opus, Fable, and GPT-55, but they don't Opus, Fable, and GPT-55, but they don't Opus, Fable, and GPT-55, but they don't have the reasoning traces. So, they have have the reasoning traces. So, they have have the reasoning traces. So, they have to take the inputs and the outputs and to take the inputs and the outputs and to take the inputs and the outputs and build a system to dynamically generate build a system to dynamically generate build a system to dynamically generate what a reasoning trace might look like.

  22. what a reasoning trace might look like. what a reasoning trace might look like. And all of the labs are now trying their And all of the labs are now trying their And all of the labs are now trying their own strategies here, which is where I own strategies here, which is where I own strategies here, which is where I think things get much more interesting. think things get much more interesting. think things get much more interesting. Now that reasoning is essential to get Now that reasoning is essential to get Now that reasoning is essential to get the model to generate good outputs, the model to generate good outputs, the model to generate good outputs, we're no longer thinking of reasoning as we're no longer thinking of reasoning as we're no longer thinking of reasoning as just the model talking to itself. We're just the model talking to itself. We're just the model talking to itself. We're seeing it as like a code golf thing, seeing it as like a code golf thing, seeing it as like a code golf thing, where we can min-max in the reasoning where we can min-max in the reasoning where we can min-max in the reasoning trace both to increase the quality of trace both to increase the quality of trace both to increase the quality of the outputs, but also to increase the the outputs, but also to increase the the outputs, but also to increase the efficiency of how the model generates efficiency of how the model generates efficiency of how the model generates things, which is why GLM-52 has this new things, which is why GLM-52 has this new things, which is why GLM-52 has this new strange format for its reasoning traces. strange format for its reasoning traces. strange format for its reasoning traces. I I tried this before, but I'm actually I I tried this before, but I'm actually I I tried this before, but I'm actually curious if I use an older GLM model, curious if I use an older GLM model, curious if I use an older GLM model, will its reasoning look meaningfully will its reasoning look meaningfully will its reasoning look meaningfully different? different? different? There we go. There we go. There we go. See, I switched over from 5.2 to the See, I switched over from 5.2 to the See, I switched over from 5.2 to the standard GLM 5, and the reasoning looks standard GLM 5, and the reasoning looks standard GLM 5, and the reasoning looks entirely different. entirely different. entirely different. Also see here, let me think about this. Also see here, let me think about this. Also see here, let me think about this. Regular stance, backside 180 kickflip Regular stance, backside 180 kickflip Regular stance, backside 180 kickflip varial flip, switch stance, that would varial flip, switch stance, that would varial flip, switch stance, that would be a varial heel flip. Nope, that is be a varial heel flip. Nope, that is be a varial heel flip. Nope, that is wrong. Wait, let me reconsider. The wrong. Wait, let me reconsider. The wrong. Wait, let me reconsider. The description says flipping the kickflip description says flipping the kickflip description says flipping the kickflip direction, that means the board flips direction, that means the board flips direction, that means the board flips with the toe dragging off the heel edge. with the toe dragging off the heel edge. with the toe dragging off the heel edge. Standard kickflip. In switch stance, a Standard kickflip. In switch stance, a Standard kickflip. In switch stance, a kickflip in switch is still a switch kickflip in switch is still a switch kickflip in switch is still a switch flip. Actually, I need to think about flip. Actually, I need to think about flip. Actually, I need to think about this more carefully. See these weights this more carefully. See these weights this more carefully. See these weights and actuallys? This is how we used to and actuallys? This is how we used to and actuallys? This is how we used to think reasoning had to happen because think reasoning had to happen because think reasoning had to happen because the model was just talking to itself to the model was just talking to itself to the model was just talking to itself to get to a better answer.

  23. get to a better answer. get to a better answer. And it kept going, it kept thinking, and And it kept going, it kept thinking, and And it kept going, it kept thinking, and then eventually it got it right, switch then eventually it got it right, switch then eventually it got it right, switch varial kickflip. Even though earlier varial kickflip. Even though earlier varial kickflip. Even though earlier here it was wrong, it thought this was a here it was wrong, it thought this was a here it was wrong, it thought this was a switch varial heel flip. But it kept switch varial heel flip. But it kept switch varial heel flip. But it kept saying, "Actually, no, wait, these other saying, "Actually, no, wait, these other saying, "Actually, no, wait, these other things matter, these other things were things matter, these other things were things matter, these other things were done differently. done differently. done differently. Wait, but there's another naming Wait, but there's another naming Wait, but there's another naming convention. No, wait, let me just go convention. No, wait, let me just go convention. No, wait, let me just go with the straightforward naming. with the straightforward naming. with the straightforward naming. Actually, I think I'm overthinking Actually, I think I'm overthinking Actually, I think I'm overthinking this." Yeah, no this." Yeah, no this." Yeah, no But again, to compare this to GLM 5.2, But again, to compare this to GLM 5.2, But again, to compare this to GLM 5.2, which here we did 1,516 which here we did 1,516 which here we did 1,516 reasoning tokens tokens total. Same reasoning tokens tokens total. Same reasoning tokens tokens total. Same correct answer from GLM 5.2 did it in correct answer from GLM 5.2 did it in correct answer from GLM 5.2 did it in 600 tokens, almost a third as many 600 tokens, almost a third as many 600 tokens, almost a third as many because it thinks in an entirely because it thinks in an entirely because it thinks in an entirely different way now. different way now. different way now. I genuinely think it's really cool that I genuinely think it's really cool that I genuinely think it's really cool that you can still see this stuff in open you can still see this stuff in open you can still see this stuff in open weight models, and you can see how they weight models, and you can see how they weight models, and you can see how they have progressed as they try to get more have progressed as they try to get more have progressed as they try to get more efficient and smarter. And saying, "No, efficient and smarter. And saying, "No, efficient and smarter. And saying, "No, wait, actually" over and over again wait, actually" over and over again wait, actually" over and over again isn't the best method, even though isn't the best method, even though isn't the best method, even though that's the one that it seems like that's the one that it seems like that's the one that it seems like everyone did for a while. A lot of the everyone did for a while. A lot of the everyone did for a while. A lot of the improvements we're seeing in new models improvements we're seeing in new models improvements we're seeing in new models are the ways that the RL process, the are the ways that the RL process, the are the ways that the RL process, the reinforcement learning systems they are reinforcement learning systems they are reinforcement learning systems they are creating are resulting in new reasoning creating are resulting in new reasoning creating are resulting in new reasoning methods being invented and tested methods being invented and tested methods being invented and tested thoroughly. All that said, GLM-52 is thoroughly. All that said, GLM-52 is thoroughly. All that said, GLM-52 is still far from an efficient model, still far from an efficient model, still far from an efficient model, taking over 42,790 taking over 42,790 taking over 42,790 tokens to complete the artificial tokens to complete the artificial tokens to complete the artificial analysis uh tasks. Like, that's a per analysis uh tasks. Like, that's a per analysis uh tasks. Like, that's a per task number. It averaged 42,790 task number. It averaged 42,790 task number. It averaged 42,790 tokens per task. Whereas, something

  24. tokens per task. Whereas, something tokens per task. Whereas, something like, I don't know, uh GPT-55 medium was like, I don't know, uh GPT-55 medium was like, I don't know, uh GPT-55 medium was 5K. That is a tenth as many tokens for a 5K. That is a tenth as many tokens for a 5K. That is a tenth as many tokens for a similar level of intelligence. I think similar level of intelligence. I think similar level of intelligence. I think I've said all I have to here. I found I've said all I have to here. I found I've said all I have to here. I found this really interesting, and when I this really interesting, and when I this really interesting, and when I started to see those silly leaks of the started to see those silly leaks of the started to see those silly leaks of the grug brain style talking, it just I grug brain style talking, it just I grug brain style talking, it just I couldn't get this out of my head. I couldn't get this out of my head. I couldn't get this out of my head. I thought this is really cool, and I hope thought this is really cool, and I hope thought this is really cool, and I hope you guys agree. I know you guys miss the you guys agree. I know you guys miss the you guys agree. I know you guys miss the technical deep dives. I hope this technical deep dives. I hope this technical deep dives. I hope this qualifies as one. I found this fun as qualifies as one. I found this fun as qualifies as one. I found this fun as hell. Let me know what you guys thought hell. Let me know what you guys thought hell. Let me know what you guys thought about it. And until next time, about it. And until next time, about it. And until next time, peace, nerds.

Summary

The discussion centers on the efficiency of large language models, particularly for code generation, comparing models like Gemini and Opus 4/8 against OpenAI's offerings. While Fable is considered the most capable, OpenAI's models stand out for their impressive capabilities achieved with significantly smaller token budgets, implying a need to understand their underlying methods for more efficient AI development. The practical takeaway is that OpenAI's approach to reasoning budgets and agent browser integration offers valuable lessons for creating more effective AI.

View original episode ↗