So much for "Pacing" the Frontier
Read full transcript 20 segments
-
Remember a week or two ago when Remember a week or two ago when everybody started talking about pacing everybody started talking about pacing everybody started talking about pacing the frontier, trying to get the Frontier the frontier, trying to get the Frontier the frontier, trying to get the Frontier Labs to slow down their constant Labs to slow down their constant Labs to slow down their constant iteration? This article from Daario that iteration? This article from Daario that iteration? This article from Daario that talked about how important it is for us talked about how important it is for us talked about how important it is for us to do this that resulted in super to do this that resulted in super to do this that resulted in super important folks like Sam Alman, Elon important folks like Sam Alman, Elon important folks like Sam Alman, Elon Musk, and I all agreeing. Yes, that is a Musk, and I all agreeing. Yes, that is a Musk, and I all agreeing. Yes, that is a joke reference to people freaking out joke reference to people freaking out joke reference to people freaking out over a silly tweet. I'm strong enough to over a silly tweet. I'm strong enough to over a silly tweet. I'm strong enough to admit that the memes around my framing admit that the memes around my framing admit that the memes around my framing of the video, specifically the title I of the video, specifically the title I of the video, specifically the title I used when I posted it on Twitter, are used when I posted it on Twitter, are used when I posted it on Twitter, are funny. This one is an instant classic. I funny. This one is an instant classic. I funny. This one is an instant classic. I do think it's important for us to talk do think it's important for us to talk do think it's important for us to talk about what pacing actually means because about what pacing actually means because about what pacing actually means because as I'm sure you all have noticed the as I'm sure you all have noticed the as I'm sure you all have noticed the pacing is going great. We have seen pacing is going great. We have seen pacing is going great. We have seen almost no models drop from the major almost no models drop from the major almost no models drop from the major labs since all of this happened. All labs since all of this happened. All labs since all of this happened. All we've seen is Gro 47 from XAI, Opus 55 we've seen is Gro 47 from XAI, Opus 55 we've seen is Gro 47 from XAI, Opus 55 from Anthropic, and then GPT6 Soul and from Anthropic, and then GPT6 Soul and from Anthropic, and then GPT6 Soul and Luna from OpenAI. Wait, that's four Luna from OpenAI. Wait, that's four Luna from OpenAI. Wait, that's four models. I thought we were pacing. What models. I thought we were pacing. What models. I thought we were pacing. What the hell is going on? I think it's the hell is going on? I think it's the hell is going on? I think it's important for us to talk about this important for us to talk about this important for us to talk about this because as great as this article is, not because as great as this article is, not because as great as this article is, not a lot of people read it. And even of the a lot of people read it. And even of the a lot of people read it. And even of the ones who did, not all of them seemed to ones who did, not all of them seemed to ones who did, not all of them seemed to understand the point of pacing the understand the point of pacing the understand the point of pacing the frontier. And here's where the hot take frontier. And here's where the hot take frontier. And here's where the hot take comes in. What we are experiencing right comes in. What we are experiencing right comes in. What we are experiencing right now with all these new model drops, that now with all these new model drops, that now with all these new model drops, that is pacing. We are pacing better than is pacing. We are pacing better than is pacing. We are pacing better than ever. And I think we're actually in a ever. And I think we're actually in a ever. And I think we're actually in a really good spot already. I'm going to really good spot already. I'm going to really good spot already. I'm going to be real. I've put far too much time and be real. I've put far too much time and be real. I've put far too much time and thought into what pacing even means now.
-
thought into what pacing even means now. thought into what pacing even means now. and working with friends at all the and working with friends at all the and working with friends at all the different labs and people into research different labs and people into research different labs and people into research and AI stuff in general to try and and AI stuff in general to try and and AI stuff in general to try and figure out what pacing looks like for us figure out what pacing looks like for us figure out what pacing looks like for us to not potentially doom ourselves. And to not potentially doom ourselves. And to not potentially doom ourselves. And I'm pretty sure I understand now. And as I'm pretty sure I understand now. And as I'm pretty sure I understand now. And as crazy as this might sound, I think a crazy as this might sound, I think a crazy as this might sound, I think a well-paced cadence for these labs is well-paced cadence for these labs is well-paced cadence for these labs is actually going to result in more models actually going to result in more models actually going to result in more models that are cheaper than you ever would that are cheaper than you ever would that are cheaper than you ever would have guessed. Before I can explain why, have guessed. Before I can explain why, have guessed. Before I can explain why, we should pace ourselves a bit and take we should pace ourselves a bit and take we should pace ourselves a bit and take a quick break for today's sponsor. As a quick break for today's sponsor. As a quick break for today's sponsor. As great as models like Fable and Astra great as models like Fable and Astra great as models like Fable and Astra are, there are still a bunch of things are, there are still a bunch of things are, there are still a bunch of things they get wrong. Two of those things that they get wrong. Two of those things that they get wrong. Two of those things that really really frustrate me are off flows really really frustrate me are off flows really really frustrate me are off flows and payments. And this is why I like and payments. And this is why I like and payments. And this is why I like today's sponsor, Clerk, so so much. today's sponsor, Clerk, so so much. today's sponsor, Clerk, so so much. These guys figured out the best way to These guys figured out the best way to These guys figured out the best way to set up developers and agents for success set up developers and agents for success set up developers and agents for success in implementing the flows that make your in implementing the flows that make your in implementing the flows that make your app actually useful. Those being the app actually useful. Those being the app actually useful. Those being the ability to sign in and manage your ability to sign in and manage your ability to sign in and manage your account and the ability to pay for account and the ability to pay for account and the ability to pay for things. Trust me on this one cuz I'm things. Trust me on this one cuz I'm things. Trust me on this one cuz I'm experienced. I made a repo last year experienced. I made a repo last year experienced. I made a repo last year complaining about just how bad Stripe is complaining about just how bad Stripe is complaining about just how bad Stripe is to set up correctly and safely in a way to set up correctly and safely in a way to set up correctly and safely in a way where you won't get scammed. and it got where you won't get scammed. and it got where you won't get scammed. and it got over 6,000 stars and is still the number over 6,000 stars and is still the number over 6,000 stars and is still the number one result on Google when you Google one result on Google when you Google one result on Google when you Google search Stripe recommendations or search Stripe recommendations or search Stripe recommendations or recommended way to set up Stripe. I put recommended way to set up Stripe. I put recommended way to set up Stripe. I put all that effort in because it was so all that effort in because it was so all that effort in because it was so hard to do right and if I put out a hard to do right and if I put out a hard to do right and if I put out a recommendation today, it would be to recommendation today, it would be to recommendation today, it would be to just use Clerk instead because not only just use Clerk instead because not only just use Clerk instead because not only have they made it trivial to set up have they made it trivial to set up have they made it trivial to set up authentication for your users and their authentication for your users and their authentication for your users and their companies and organizations, they also companies and organizations, they also companies and organizations, they also have the best billing product I've ever have the best billing product I've ever have the best billing product I've ever seen baked right in. The code couldn't seen baked right in. The code couldn't seen baked right in. The code couldn't be simpler. Not that you're going to be simpler. Not that you're going to be simpler. Not that you're going to read it anyways, but simple code means read it anyways, but simple code means read it anyways, but simple code means the agents can understand it too. in the the agents can understand it too. in the the agents can understand it too. in the clerk skill, which they provide with a clerk skill, which they provide with a clerk skill, which they provide with a one-click copy to paste into whatever one-click copy to paste into whatever one-click copy to paste into whatever agent you want to use, will set the
-
agent you want to use, will set the agent you want to use, will set the agents up for success. Not that they agents up for success. Not that they agents up for success. Not that they even need it. I've had pretty good luck even need it. I've had pretty good luck even need it. I've had pretty good luck just telling my agents, "Go set up clerk just telling my agents, "Go set up clerk just telling my agents, "Go set up clerk on whatever project." And whether it's on whatever project." And whether it's on whatever project." And whether it's web, mobile, desktop, anything else, web, mobile, desktop, anything else, web, mobile, desktop, anything else, they get it figured out. That's why they get it figured out. That's why they get it figured out. That's why we're using it in T3 Code for all of our we're using it in T3 Code for all of our we're using it in T3 Code for all of our different platforms. They even built an different platforms. They even built an different platforms. They even built an Electron package for us and have been Electron package for us and have been Electron package for us and have been keeping it up to date as all of our keeping it up to date as all of our keeping it up to date as all of our requirements change. The only thing requirements change. The only thing requirements change. The only thing better than Clerk's Code is the team, better than Clerk's Code is the team, better than Clerk's Code is the team, and they've been awesome to work with and they've been awesome to work with and they've been awesome to work with every step along the way. Figure out why every step along the way. Figure out why every step along the way. Figure out why I like them so much. at soy.link/clerk. I like them so much. at soy.link/clerk. I like them so much. at soy.link/clerk. In order for us to have a good In order for us to have a good In order for us to have a good conversation about pacing and how things conversation about pacing and how things conversation about pacing and how things are changing, we need to understand what are changing, we need to understand what are changing, we need to understand what we're trying to prevent first. So, we we're trying to prevent first. So, we we're trying to prevent first. So, we need to start with the core question. need to start with the core question. need to start with the core question. Why are we pacing? Why are we choosing Why are we pacing? Why are we choosing Why are we pacing? Why are we choosing to do this? There's a vague general idea to do this? There's a vague general idea to do this? There's a vague general idea of the models get too smart and they can of the models get too smart and they can of the models get too smart and they can destroy everything. But that's not destroy everything. But that's not destroy everything. But that's not really the issue. The problem isn't that really the issue. The problem isn't that really the issue. The problem isn't that if we hit a certain score on some if we hit a certain score on some if we hit a certain score on some benchmark, the model might go rogue and benchmark, the model might go rogue and benchmark, the model might go rogue and kill us all. The problem isn't about a kill us all. The problem isn't about a kill us all. The problem isn't about a specific point we hit. The thing we're specific point we hit. The thing we're specific point we hit. The thing we're really scared about here is the idea of really scared about here is the idea of really scared about here is the idea of the takeoff. Not that a certain level of the takeoff. Not that a certain level of the takeoff. Not that a certain level of intelligence is hit, but at a certain intelligence is hit, but at a certain intelligence is hit, but at a certain level of intelligence and recursive level of intelligence and recursive level of intelligence and recursive self-improvement, the model starts self-improvement, the model starts self-improvement, the model starts improving so fast that we can't improving so fast that we can't improving so fast that we can't understand what it's doing or why. The understand what it's doing or why. The understand what it's doing or why. The scary stuff starts when the models scary stuff starts when the models scary stuff starts when the models improve in ways that we can't improve in ways that we can't improve in ways that we can't understand. And if we have more and more understand. And if we have more and more understand. And if we have more and more AI agents improving the models and they AI agents improving the models and they AI agents improving the models and they start to be able to do it better than start to be able to do it better than start to be able to do it better than humans at an exponential rate, we'll humans at an exponential rate, we'll humans at an exponential rate, we'll lose track of why and how and more lose track of why and how and more lose track of why and how and more importantly what they're even doing in importantly what they're even doing in importantly what they're even doing in the first place. There are some pretty
-
the first place. There are some pretty the first place. There are some pretty silly but also painfully real examples silly but also painfully real examples silly but also painfully real examples of this. The first I'm going to cite is of this. The first I'm going to cite is of this. The first I'm going to cite is all the work OpenAI has been doing to all the work OpenAI has been doing to all the work OpenAI has been doing to make the models more efficient. make the models more efficient. make the models more efficient. Flashbang warning before I go over to Flashbang warning before I go over to Flashbang warning before I go over to artificial analysis. Thing I've talked artificial analysis. Thing I've talked artificial analysis. Thing I've talked about a lot on this channel is how much about a lot on this channel is how much about a lot on this channel is how much more efficient OpenAI's models are. Most more efficient OpenAI's models are. Most more efficient OpenAI's models are. Most of the difference comes from the of the difference comes from the of the difference comes from the reasoning tokens, the tokens that are reasoning tokens, the tokens that are reasoning tokens, the tokens that are generated to steer the model in the generated to steer the model in the generated to steer the model in the right direction. I've shown examples in right direction. I've shown examples in right direction. I've shown examples in my other videos, but here's a pretty my other videos, but here's a pretty my other videos, but here's a pretty silly one. The reasoning traces that the silly one. The reasoning traces that the silly one. The reasoning traces that the OpenAI models do aren't speaking proper OpenAI models do aren't speaking proper OpenAI models do aren't speaking proper well- formatted English. In order to well- formatted English. In order to well- formatted English. In order to make the model use less tokens, they make the model use less tokens, they make the model use less tokens, they trained it to speak in grug style where trained it to speak in grug style where trained it to speak in grug style where it's like not including words that it's like not including words that it's like not including words that aren't important even if they make it aren't important even if they make it aren't important even if they make it more grammatically correct. need final more grammatically correct. need final more grammatically correct. need final since we asked and must wait but final since we asked and must wait but final since we asked and must wait but final the answer should just be ask we already the answer should just be ask we already the answer should just be ask we already in commentary final maybe not JSON in commentary final maybe not JSON in commentary final maybe not JSON trailing field issue extra comma let's trailing field issue extra comma let's trailing field issue extra comma let's call proper also terminal get diff call proper also terminal get diff call proper also terminal get diff status and inspect relevant parallel status and inspect relevant parallel status and inspect relevant parallel just the word parallel as like a just the word parallel as like a just the word parallel as like a sentence lowercase no need tool preamble sentence lowercase no need tool preamble sentence lowercase no need tool preamble again use parallel tool the point I'm again use parallel tool the point I'm again use parallel tool the point I'm trying to make here is that openai has trying to make here is that openai has trying to make here is that openai has put a lot of effort into making the put a lot of effort into making the put a lot of effort into making the model more efficient, but the things model more efficient, but the things model more efficient, but the things they've done to make it more efficient they've done to make it more efficient they've done to make it more efficient have also, to be frank, made it harder have also, to be frank, made it harder have also, to be frank, made it harder to understand what it's doing. We don't to understand what it's doing. We don't to understand what it's doing. We don't get these reasoning traces normally. All get these reasoning traces normally. All get these reasoning traces normally. All these examples are when they leaked these examples are when they leaked these examples are when they leaked accidentally due to like bugs in random accidentally due to like bugs in random accidentally due to like bugs in random places. But when OpenAI tries to trace places. But when OpenAI tries to trace places. But when OpenAI tries to trace what the reasoning is doing, it what the reasoning is doing, it what the reasoning is doing, it struggles. Now, they called out in the struggles. Now, they called out in the struggles. Now, they called out in the GPT6 Astra system card, and I cited this
-
GPT6 Astra system card, and I cited this GPT6 Astra system card, and I cited this so much in my video as well as in the so much in my video as well as in the so much in my video as well as in the pacing videos, that they don't actually pacing videos, that they don't actually pacing videos, that they don't actually know what the model's doing as know what the model's doing as know what the model's doing as confidently as they did before. Because confidently as they did before. Because confidently as they did before. Because when the model thinks it's being when the model thinks it's being when the model thinks it's being monitored, it obuscates its reasoning in monitored, it obuscates its reasoning in monitored, it obuscates its reasoning in a way that makes it way harder to know a way that makes it way harder to know a way that makes it way harder to know what it's doing. This is a small example what it's doing. This is a small example what it's doing. This is a small example of the types of issues that we're going of the types of issues that we're going of the types of issues that we're going to start seeing. As the models get more to start seeing. As the models get more to start seeing. As the models get more and more effective, they'll get to a and more effective, they'll get to a and more effective, they'll get to a point where we can't understand the point where we can't understand the point where we can't understand the outputs. I'm going to give a relatively outputs. I'm going to give a relatively outputs. I'm going to give a relatively silly example here, and I am sorry in silly example here, and I am sorry in silly example here, and I am sorry in advance for all the low-level engineers advance for all the low-level engineers advance for all the low-level engineers I'm about to offend. I want to talk I'm about to offend. I want to talk I'm about to offend. I want to talk about the language C for a second. about the language C for a second. about the language C for a second. Before we had languages like C, the vast Before we had languages like C, the vast Before we had languages like C, the vast majority of programmers were programming majority of programmers were programming majority of programmers were programming in anything from punch cards to in anything from punch cards to in anything from punch cards to assembly. At that point, like right assembly. At that point, like right assembly. At that point, like right before C, it was largely assembly. There before C, it was largely assembly. There before C, it was largely assembly. There were different types of assembly for were different types of assembly for were different types of assembly for different architectures, different different architectures, different different architectures, different computers, different everything. And computers, different everything. And computers, different everything. And writing the different versions of writing the different versions of writing the different versions of assembly for different systems was assembly for different systems was assembly for different systems was tedious and obnoxious. On top of that, tedious and obnoxious. On top of that, tedious and obnoxious. On top of that, assembly was not the most pleasant or assembly was not the most pleasant or assembly was not the most pleasant or ergonomic thing to work with. C was ergonomic thing to work with. C was ergonomic thing to work with. C was largely invented to have a standard largely invented to have a standard largely invented to have a standard language that could be compiled to language that could be compiled to language that could be compiled to different platforms in one place. Write different platforms in one place. Write different platforms in one place. Write it once, compile it to whatever you need it once, compile it to whatever you need it once, compile it to whatever you need it to run on. Initially, this was it to run on. Initially, this was it to run on. Initially, this was awesome because you didn't have to write awesome because you didn't have to write awesome because you didn't have to write assembly code for all the different assembly code for all the different assembly code for all the different platforms you wanted to support. But platforms you wanted to support. But platforms you wanted to support. But something pretty crazy happened. It something pretty crazy happened. It something pretty crazy happened. It turns out when you have an abstraction turns out when you have an abstraction turns out when you have an abstraction layer like that, it becomes way easier layer like that, it becomes way easier layer like that, it becomes way easier to write way more code and bigger to write way more code and bigger to write way more code and bigger projects with more complexity became projects with more complexity became projects with more complexity became possible. Eventually, new languages were
-
possible. Eventually, new languages were possible. Eventually, new languages were written on top of C and then virtual written on top of C and then virtual written on top of C and then virtual systems and runtimes were built on top systems and runtimes were built on top systems and runtimes were built on top of it as well that new languages could of it as well that new languages could of it as well that new languages could go layers on top. The point I'm trying go layers on top. The point I'm trying go layers on top. The point I'm trying to make here is that we started to make here is that we started to make here is that we started abstracting really fast. And an abstracting really fast. And an abstracting really fast. And an abstraction that was initially built to abstraction that was initially built to abstraction that was initially built to solve a low-level problem ended up solve a low-level problem ended up solve a low-level problem ended up enabling so much more. And the sheer enabling so much more. And the sheer enabling so much more. And the sheer amount of assembly code that existed amount of assembly code that existed amount of assembly code that existed years after C was exponentially more years after C was exponentially more years after C was exponentially more than years before. Before we had C, real than years before. Before we had C, real than years before. Before we had C, real devs used assembly. After we had C, the devs used assembly. After we had C, the devs used assembly. After we had C, the number of assembly devs probably went number of assembly devs probably went number of assembly devs probably went down year-over-year. It might have taken down year-over-year. It might have taken down year-over-year. It might have taken a bit, but it was almost laughable a bit, but it was almost laughable a bit, but it was almost laughable pretty quickly if you chose to focus on pretty quickly if you chose to focus on pretty quickly if you chose to focus on assembly code for things that didn't assembly code for things that didn't assembly code for things that didn't need that tiny edge you could squeeze need that tiny edge you could squeeze need that tiny edge you could squeeze out. Obviously, like ffmpeg benefits out. Obviously, like ffmpeg benefits out. Obviously, like ffmpeg benefits greatly from the bits of assembly they greatly from the bits of assembly they greatly from the bits of assembly they use for weird encoding stuff, but for use for weird encoding stuff, but for use for weird encoding stuff, but for the most part, C was the right move. So, the most part, C was the right move. So, the most part, C was the right move. So, why am I bringing this up? Imagine why am I bringing this up? Imagine why am I bringing this up? Imagine you're a very, very good assembly you're a very, very good assembly you're a very, very good assembly developer in the days before C. When C developer in the days before C. When C developer in the days before C. When C first came out, you probably had mixed first came out, you probably had mixed first came out, you probably had mixed feelings about it. On one hand, it's feelings about it. On one hand, it's feelings about it. On one hand, it's nice to write things once and compile nice to write things once and compile nice to write things once and compile them to other places. On the other hand, them to other places. On the other hand, them to other places. On the other hand, you know assembly really well and you've you know assembly really well and you've you know assembly really well and you've taken a look at the assembly code that's taken a look at the assembly code that's taken a look at the assembly code that's coming out of the C compiler and you're coming out of the C compiler and you're coming out of the C compiler and you're unimpressed. You think it's too verbose.
-
unimpressed. You think it's too verbose. unimpressed. You think it's too verbose. There is too many things happening and There is too many things happening and There is too many things happening and you wish you could just write the you wish you could just write the you wish you could just write the assembly instead because it would be assembly instead because it would be assembly instead because it would be much simpler and easier to maintain much simpler and easier to maintain much simpler and easier to maintain code. And then C explodes. We end up code. And then C explodes. We end up code. And then C explodes. We end up with these massive code bases written in with these massive code bases written in with these massive code bases written in C. And the assembly that gets compiled C. And the assembly that gets compiled C. And the assembly that gets compiled out of those, you probably can't even out of those, you probably can't even out of those, you probably can't even read. The crazy thing that happens here read. The crazy thing that happens here read. The crazy thing that happens here is as you build a system like C that is is as you build a system like C that is is as you build a system like C that is an abstraction over previous systems an abstraction over previous systems an abstraction over previous systems data languages whatever else initially data languages whatever else initially data languages whatever else initially it might just feel a little it might just feel a little it might just feel a little uncomfortable like oh this is doing uncomfortable like oh this is doing uncomfortable like oh this is doing things not quite the way I would want things not quite the way I would want things not quite the way I would want but I guess there's some benefit to the but I guess there's some benefit to the but I guess there's some benefit to the way it's working and like the scale it way it's working and like the scale it way it's working and like the scale it allows for when you start solving those allows for when you start solving those allows for when you start solving those lower level bugs a really weird lower level bugs a really weird lower level bugs a really weird phenomena happens and this is what I'm phenomena happens and this is what I'm phenomena happens and this is what I'm going to refer to as takeoff do you going to refer to as takeoff do you going to refer to as takeoff do you think there was more or less assembly think there was more or less assembly think there was more or less assembly code 5 years after C versus right before code 5 years after C versus right before code 5 years after C versus right before C. I don't mean people checking assembly C. I don't mean people checking assembly C. I don't mean people checking assembly into whatever [ __ ] show of a source into whatever [ __ ] show of a source into whatever [ __ ] show of a source control existed at the time. I'm talking control existed at the time. I'm talking control existed at the time. I'm talking about the actual like outputs that would about the actual like outputs that would about the actual like outputs that would then be compiled and run on different then be compiled and run on different then be compiled and run on different systems. The introduction of C systems. The introduction of C systems. The introduction of C exponentially increased the amount of exponentially increased the amount of exponentially increased the amount of assembly that existed because it turns assembly that existed because it turns assembly that existed because it turns out when you make it way easier to do out when you make it way easier to do out when you make it way easier to do the thing, you massively increase the the thing, you massively increase the the thing, you massively increase the amount. This is often referred to as amount. This is often referred to as amount. This is often referred to as Jevans paradox. You have a thing that's Jevans paradox. You have a thing that's Jevans paradox. You have a thing that's done a little bit that has some rough done a little bit that has some rough done a little bit that has some rough edges. You solve some of those rough edges. You solve some of those rough edges. You solve some of those rough edges because they annoy you and then edges because they annoy you and then edges because they annoy you and then suddenly the thing massively ramps up.
-
suddenly the thing massively ramps up. suddenly the thing massively ramps up. The classic example here was when coal The classic example here was when coal The classic example here was when coal got cheaper and steam power got easier got cheaper and steam power got easier got cheaper and steam power got easier to access. The cost to use coal went to access. The cost to use coal went to access. The cost to use coal went down a bunch and the amount of coal down a bunch and the amount of coal down a bunch and the amount of coal being purchased went up exponentially being purchased went up exponentially being purchased went up exponentially because it being cheaper to use and because it being cheaper to use and because it being cheaper to use and easier to use made it way more valuable. easier to use made it way more valuable. easier to use made it way more valuable. Shout out to Hank Green for pointing out Shout out to Hank Green for pointing out Shout out to Hank Green for pointing out that it's not really a paradox, a shitty that it's not really a paradox, a shitty that it's not really a paradox, a shitty name, but you you get the idea. There name, but you you get the idea. There name, but you you get the idea. There are certain things where the device are certain things where the device are certain things where the device being cheaper doesn't really increase being cheaper doesn't really increase being cheaper doesn't really increase adoption. For example, cheaper fridges. adoption. For example, cheaper fridges. adoption. For example, cheaper fridges. Almost every house in America has a Almost every house in America has a Almost every house in America has a refrigerator. Refrigerators getting refrigerator. Refrigerators getting refrigerator. Refrigerators getting three times cheaper would not three times cheaper would not three times cheaper would not meaningfully increase the number of meaningfully increase the number of meaningfully increase the number of houses with refrigerators. Maybe a houses with refrigerators. Maybe a houses with refrigerators. Maybe a handful would buy a second one for their handful would buy a second one for their handful would buy a second one for their garage or something, but the need will garage or something, but the need will garage or something, but the need will plateau. The need for hot water has plateau. The need for hot water has plateau. The need for hot water has plateaued as well. Things like that. So, plateaued as well. Things like that. So, plateaued as well. Things like that. So, why am I bringing all of this up in a why am I bringing all of this up in a why am I bringing all of this up in a video about pacing? Hear me out. After C video about pacing? Hear me out. After C video about pacing? Hear me out. After C had been around for 5 to 10 years, if had been around for 5 to 10 years, if had been around for 5 to 10 years, if you took that particularly skilled you took that particularly skilled you took that particularly skilled assembly developer and had them read the assembly developer and had them read the assembly developer and had them read the assembly from a more modern version of assembly from a more modern version of assembly from a more modern version of the C compiler with more things added to the C compiler with more things added to the C compiler with more things added to the assembly language, more instructions the assembly language, more instructions the assembly language, more instructions that were exposed over x86 and whatever, that were exposed over x86 and whatever, that were exposed over x86 and whatever, how well do you think that developer how well do you think that developer how well do you think that developer will understand the assembly code coming will understand the assembly code coming will understand the assembly code coming out of a gigantic C codebase? probably out of a gigantic C codebase? probably out of a gigantic C codebase? probably not very well anymore. So even though not very well anymore. So even though not very well anymore. So even though that dev knew assembly incredibly well, that dev knew assembly incredibly well, that dev knew assembly incredibly well, maybe they even helped build C in the maybe they even helped build C in the maybe they even helped build C in the first place. Not only is the amount of C first place. Not only is the amount of C first place. Not only is the amount of C being written going up, the size of code being written going up, the size of code being written going up, the size of code bases is too. And the obscurity of the bases is too. And the obscurity of the bases is too. And the obscurity of the assembly code coming out of the compiler
-
assembly code coming out of the compiler assembly code coming out of the compiler is going up as well. So on one hand, we is going up as well. So on one hand, we is going up as well. So on one hand, we have this exponential takeoff of the have this exponential takeoff of the have this exponential takeoff of the amount of code in the world, the amount amount of code in the world, the amount amount of code in the world, the amount of projects in the world, and the amount of projects in the world, and the amount of projects in the world, and the amount of code in a given codebase. And of code in a given codebase. And of code in a given codebase. And alongside that, our ability to alongside that, our ability to alongside that, our ability to comprehend the outputs went down. One comprehend the outputs went down. One comprehend the outputs went down. One more time, the things that we did to more time, the things that we did to more time, the things that we did to expand the usability of code and expand the usability of code and expand the usability of code and software also lowered our understanding software also lowered our understanding software also lowered our understanding of the underlying code that was powering of the underlying code that was powering of the underlying code that was powering those things. As more software was those things. As more software was those things. As more software was written and bigger programs could exist, written and bigger programs could exist, written and bigger programs could exist, our ability to comprehend the assembly our ability to comprehend the assembly our ability to comprehend the assembly coming out went down. Hopefully, you coming out went down. Hopefully, you coming out went down. Hopefully, you could see what this has to do with what could see what this has to do with what could see what this has to do with what we're talking about today with AI. The we're talking about today with AI. The we're talking about today with AI. The more we do to make models efficient, the more we do to make models efficient, the more we do to make models efficient, the more we do to make models capable and more we do to make models capable and more we do to make models capable and the more models do to improve the more models do to improve the more models do to improve themselves, the harder and harder it themselves, the harder and harder it themselves, the harder and harder it will become for us to understand the will become for us to understand the will become for us to understand the what and the how and the why they are what and the how and the why they are what and the how and the why they are doing things. There are obviously doing things. There are obviously doing things. There are obviously exceptions with the C example. Like exceptions with the C example. Like exceptions with the C example. Like death fudge pointed out, GCC with all death fudge pointed out, GCC with all death fudge pointed out, GCC with all flags off is almost readable, but as flags off is almost readable, but as flags off is almost readable, but as soon as you use -3, you're not going to soon as you use -3, you're not going to soon as you use -3, you're not going to understand a thing going on. And that's understand a thing going on. And that's understand a thing going on. And that's one of the many things starting to one of the many things starting to one of the many things starting to happen. Eventually, we got to the point happen. Eventually, we got to the point happen. Eventually, we got to the point with the C compiler where the C compiler with the C compiler where the C compiler with the C compiler where the C compiler was written in C itself. Do you have any was written in C itself. Do you have any was written in C itself. Do you have any idea how much harder the assembly output idea how much harder the assembly output idea how much harder the assembly output of the C compiler itself probably was to of the C compiler itself probably was to of the C compiler itself probably was to read than any of the assembly that read than any of the assembly that read than any of the assembly that existed before C? Now, imagine that C existed before C? Now, imagine that C existed before C? Now, imagine that C could improve itself, that C could keep could improve itself, that C could keep could improve itself, that C could keep changing and improving the compiler's changing and improving the compiler's changing and improving the compiler's efficiency autonomously, and it could do efficiency autonomously, and it could do efficiency autonomously, and it could do such in ways that humans didn't such in ways that humans didn't such in ways that humans didn't necessarily understand. Now combine this
-
necessarily understand. Now combine this necessarily understand. Now combine this with the quad slot problem from models with the quad slot problem from models with the quad slot problem from models like Opus 5 and how the things they say like Opus 5 and how the things they say like Opus 5 and how the things they say are basically impossible to perceive. As are basically impossible to perceive. As are basically impossible to perceive. As the capability of these things goes up the capability of these things goes up the capability of these things goes up and the efficiency of these things goes and the efficiency of these things goes and the efficiency of these things goes up, the comprehensibility of it goes up, the comprehensibility of it goes up, the comprehensibility of it goes down. And this is the paradox that I'm down. And this is the paradox that I'm down. And this is the paradox that I'm concerned about here. The things we do concerned about here. The things we do concerned about here. The things we do to improve efficiency to make models to improve efficiency to make models to improve efficiency to make models smarter and also most importantly smarter and also most importantly smarter and also most importantly self-improving inherently makes the self-improving inherently makes the self-improving inherently makes the output harder for us to understand. The output harder for us to understand. The output harder for us to understand. The most efficient C compiler is probably most efficient C compiler is probably most efficient C compiler is probably much harder to read the assembly of than much harder to read the assembly of than much harder to read the assembly of than the simplest C compiler. And to go back the simplest C compiler. And to go back the simplest C compiler. And to go back to the reasoning trace thing, a super to the reasoning trace thing, a super to the reasoning trace thing, a super silly example of this is that anthropics silly example of this is that anthropics silly example of this is that anthropics models are not as efficient with their models are not as efficient with their models are not as efficient with their reasoning tokens. So they're easier to reasoning tokens. So they're easier to reasoning tokens. So they're easier to monitor. In fact, from Opus 5 to Opus monitor. In fact, from Opus 5 to Opus monitor. In fact, from Opus 5 to Opus 5.5, the number of reasoning tokens per 5.5, the number of reasoning tokens per 5.5, the number of reasoning tokens per task on artificial analysis doubled. task on artificial analysis doubled. task on artificial analysis doubled. Actual output tokens only increased by Actual output tokens only increased by Actual output tokens only increased by 5,000 out of 35,000. So went from 30k 5,000 out of 35,000. So went from 30k 5,000 out of 35,000. So went from 30k tokens of normal output to 35k tokens. tokens of normal output to 35k tokens. tokens of normal output to 35k tokens. But the reasoning doubled from 42,000 to But the reasoning doubled from 42,000 to But the reasoning doubled from 42,000 to 84,000.
-
84,000. 84,000. The reasoning tokens are used to get The reasoning tokens are used to get The reasoning tokens are used to get better answers from the model. But better answers from the model. But better answers from the model. But importantly, those very same reasoning importantly, those very same reasoning importantly, those very same reasoning tokens make it easier to monitor what tokens make it easier to monitor what tokens make it easier to monitor what the model is doing. Since the reasoning the model is doing. Since the reasoning the model is doing. Since the reasoning traces are what steer the model and its traces are what steer the model and its traces are what steer the model and its behavior, more of them gives us more behavior, more of them gives us more behavior, more of them gives us more insight on how the model works. I'm insight on how the model works. I'm insight on how the model works. I'm going to try to emphasize my point here going to try to emphasize my point here going to try to emphasize my point here by giving the inverse example. Imagine by giving the inverse example. Imagine by giving the inverse example. Imagine if Anthropic spun up Fable 5.1 and said, if Anthropic spun up Fable 5.1 and said, if Anthropic spun up Fable 5.1 and said, "Hey, we're training a new Fable model. "Hey, we're training a new Fable model. "Hey, we're training a new Fable model. I want you to find every single possible I want you to find every single possible I want you to find every single possible thing we can do to make the reasoning thing we can do to make the reasoning thing we can do to make the reasoning traces simpler and smaller. How can you traces simpler and smaller. How can you traces simpler and smaller. How can you trim more and more so that we can get trim more and more so that we can get trim more and more so that we can get similar performance out with half the similar performance out with half the similar performance out with half the reasoning tokens?" I wouldn't be reasoning tokens?" I wouldn't be reasoning tokens?" I wouldn't be surprised if Fable could pull that off. surprised if Fable could pull that off. surprised if Fable could pull that off. The problem there is that if you try to The problem there is that if you try to The problem there is that if you try to communicate the same thing, the same communicate the same thing, the same communicate the same thing, the same ideas, the same concepts, the same plans ideas, the same concepts, the same plans ideas, the same concepts, the same plans and work in half as many tokens, you're and work in half as many tokens, you're and work in half as many tokens, you're going to make it harder to understand going to make it harder to understand going to make it harder to understand and monitor. And monitorability is one and monitor. And monitorability is one and monitor. And monitorability is one of the most important parts of pacing of the most important parts of pacing of the most important parts of pacing correctly. One of the very specific correctly. One of the very specific correctly. One of the very specific goals of the pacing the frontier ideas goals of the pacing the frontier ideas goals of the pacing the frontier ideas that have been proposed by Daario and that have been proposed by Daario and that have been proposed by Daario and others is that we will have a better others is that we will have a better others is that we will have a better understanding of how these models work understanding of how these models work understanding of how these models work and operate next year than we do today.
-
and operate next year than we do today. and operate next year than we do today. Because if we don't make this an Because if we don't make this an Because if we don't make this an intentional choice, our understanding of intentional choice, our understanding of intentional choice, our understanding of models and their behavior will go down models and their behavior will go down models and their behavior will go down inherently. People seem to think that inherently. People seem to think that inherently. People seem to think that the safety issue is separate from the safety issue is separate from the safety issue is separate from pacing. I'm I'm I don't know why. You're pacing. I'm I'm I don't know why. You're pacing. I'm I'm I don't know why. You're just wrong. This is the pacing article just wrong. This is the pacing article just wrong. This is the pacing article from Daario that is a whole section from Daario that is a whole section from Daario that is a whole section about interpretability. The science of about interpretability. The science of about interpretability. The science of understanding what happens inside of AI understanding what happens inside of AI understanding what happens inside of AI models has made enormous progress over models has made enormous progress over models has made enormous progress over the last few years and it's playing an the last few years and it's playing an the last few years and it's playing an increasingly important part of auditing increasingly important part of auditing increasingly important part of auditing the models before release. It's almost the models before release. It's almost the models before release. It's almost used like an fMRI scan, but for the used like an fMRI scan, but for the used like an fMRI scan, but for the brain of an AI to help us understand and brain of an AI to help us understand and brain of an AI to help us understand and see the underlying reasons for a given see the underlying reasons for a given see the underlying reasons for a given behavior. They called out that when they behavior. They called out that when they behavior. They called out that when they were going through the instances of were going through the instances of were going through the instances of Clawude publishing malicious stuff, like Clawude publishing malicious stuff, like Clawude publishing malicious stuff, like their equivalents of the hugging face their equivalents of the hugging face their equivalents of the hugging face incidents, that when they did actually incidents, that when they did actually incidents, that when they did actually send tools and engineers in to try and send tools and engineers in to try and send tools and engineers in to try and figure out what caused these things, figure out what caused these things, figure out what caused these things, that the chains of thought were not that the chains of thought were not that the chains of thought were not enough. They actually had to resample enough. They actually had to resample enough. They actually had to resample experiments from different points in the experiments from different points in the experiments from different points in the incident transcript and also do incident transcript and also do incident transcript and also do interpretability analysis of model interpretability analysis of model interpretability analysis of model activations inside the model itself. The activations inside the model itself. The activations inside the model itself. The reasoning traces were not enough and we reasoning traces were not enough and we reasoning traces were not enough and we need more ways to see into the model and need more ways to see into the model and need more ways to see into the model and what it is doing. The number of what it is doing. The number of what it is doing. The number of reasoning tokens and the readability of reasoning tokens and the readability of reasoning tokens and the readability of those reasoning tokens is a small those reasoning tokens is a small those reasoning tokens is a small example of this, but it's a really easy example of this, but it's a really easy example of this, but it's a really easy to understand one in my opinion. That's to understand one in my opinion. That's to understand one in my opinion. That's why I'm choosing to fixate on it so why I'm choosing to fixate on it so why I'm choosing to fixate on it so much. So again, the bad example here much. So again, the bad example here much. So again, the bad example here would be if Anthropic told Fable to would be if Anthropic told Fable to would be if Anthropic told Fable to train the new model to be way more train the new model to be way more train the new model to be way more efficient. It might end up making the efficient. It might end up making the efficient. It might end up making the model harder to understand and model harder to understand and model harder to understand and interpret. And you got to find that interpret. And you got to find that interpret. And you got to find that balance here. Anthropic seems more balance here. Anthropic seems more balance here. Anthropic seems more focused on both inventing God and making
-
focused on both inventing God and making focused on both inventing God and making sure they can understand God's brain. sure they can understand God's brain. sure they can understand God's brain. OpenAI is treating this like an OpenAI is treating this like an OpenAI is treating this like an engineering pro and trying to make engineering pro and trying to make engineering pro and trying to make things as efficient as possible. That things as efficient as possible. That things as efficient as possible. That said, seems like they're even changing a said, seems like they're even changing a said, seems like they're even changing a bit here, too. It's not as big of a gap bit here, too. It's not as big of a gap bit here, too. It's not as big of a gap as before, but for 56 soul to Astra, the as before, but for 56 soul to Astra, the as before, but for 56 soul to Astra, the amount of tokens did go down very amount of tokens did go down very amount of tokens did go down very slightly, but from 56 soul to six soul, slightly, but from 56 soul to six soul, slightly, but from 56 soul to six soul, it actually went up quite a bit. The it actually went up quite a bit. The it actually went up quite a bit. The number of reasoning tokens went up by number of reasoning tokens went up by number of reasoning tokens went up by like 25 to 30%. This is also possibly like 25 to 30%. This is also possibly like 25 to 30%. This is also possibly kind of an example of pacing. Avoiding kind of an example of pacing. Avoiding kind of an example of pacing. Avoiding the things that you can do to improve the things that you can do to improve the things that you can do to improve model performance that would result in model performance that would result in model performance that would result in it being harder to understand and more it being harder to understand and more it being harder to understand and more dangerous potentially is important. And dangerous potentially is important. And dangerous potentially is important. And pasting means putting more effort into pasting means putting more effort into pasting means putting more effort into things that make it easy to understand things that make it easy to understand things that make it easy to understand the models at the cost of the model's the models at the cost of the model's the models at the cost of the model's performance and capability. There is performance and capability. There is performance and capability. There is more to dig into here, though. I want to more to dig into here, though. I want to more to dig into here, though. I want to take a look at the models that we just take a look at the models that we just take a look at the models that we just saw get released. The crazy bulk set of saw get released. The crazy bulk set of saw get released. The crazy bulk set of model releases we just went through were model releases we just went through were model releases we just went through were Gro 47, GPD6 Soul, GBD6 Luna, and Opus Gro 47, GPD6 Soul, GBD6 Luna, and Opus Gro 47, GPD6 Soul, GBD6 Luna, and Opus 5.5. A few weeks ago, we had two 5.5. A few weeks ago, we had two 5.5. A few weeks ago, we had two different models drop. We had Fable 5.1, different models drop. We had Fable 5.1, different models drop. We had Fable 5.1, and we had GBT6 Astra. I hope I don't and we had GBT6 Astra. I hope I don't and we had GBT6 Astra. I hope I don't have to emphasize there's a pretty big have to emphasize there's a pretty big have to emphasize there's a pretty big difference between these models, but I'm difference between these models, but I'm difference between these models, but I'm afraid I do because I saw a lot of afraid I do because I saw a lot of afraid I do because I saw a lot of confusing comments when I mentioned that confusing comments when I mentioned that confusing comments when I mentioned that Opus 55 is actually an example of Opus 55 is actually an example of Opus 55 is actually an example of pacing. Well, the difference here is pacing. Well, the difference here is pacing. Well, the difference here is that Opus 55 isn't discovering much in that Opus 55 isn't discovering much in that Opus 55 isn't discovering much in terms of new novel capabilities. Models terms of new novel capabilities. Models terms of new novel capabilities. Models like Fable and Astra have introduced new like Fable and Astra have introduced new like Fable and Astra have introduced new capabilities, new potentials, new things
-
capabilities, new potentials, new things capabilities, new potentials, new things to the suite of stuff LLMs can do. They to the suite of stuff LLMs can do. They to the suite of stuff LLMs can do. They are big. They are bold. They are scary. are big. They are bold. They are scary. are big. They are bold. They are scary. They are powerful. They are smart. They They are powerful. They are smart. They They are powerful. They are smart. They are harder to understand. And they are are harder to understand. And they are are harder to understand. And they are [ __ ] awesome. It's so cool. But [ __ ] awesome. It's so cool. But [ __ ] awesome. It's so cool. But there's a huge gap between something there's a huge gap between something there's a huge gap between something like Fable and Astra in models that are like Fable and Astra in models that are like Fable and Astra in models that are now the smaller tiers like Opus, Luna, now the smaller tiers like Opus, Luna, now the smaller tiers like Opus, Luna, Soul, and Grock. The difference is that Soul, and Grock. The difference is that Soul, and Grock. The difference is that we kind of have a ceiling. It's not a we kind of have a ceiling. It's not a we kind of have a ceiling. It's not a hard ceiling though. Some things can hard ceiling though. Some things can hard ceiling though. Some things can spike slightly above, some things can spike slightly above, some things can spike slightly above, some things can spike slightly below, but we have a spike slightly below, but we have a spike slightly below, but we have a ceiling of what we believe is the ceiling of what we believe is the ceiling of what we believe is the capability of these models. and we can capability of these models. and we can capability of these models. and we can refine and distill it, but I think refine and distill it, but I think refine and distill it, but I think people look at benchmarks wrong. Let's people look at benchmarks wrong. Let's people look at benchmarks wrong. Let's again use Opus 5.5 as an example here. A again use Opus 5.5 as an example here. A again use Opus 5.5 as an example here. A lot of people have cited the scores that lot of people have cited the scores that lot of people have cited the scores that Opus 55 got on a bunch of benchmarks as Opus 55 got on a bunch of benchmarks as Opus 55 got on a bunch of benchmarks as proof that we aren't actually pacing proof that we aren't actually pacing proof that we aren't actually pacing right now because we wouldn't see scores right now because we wouldn't see scores right now because we wouldn't see scores going up still if we were trying to be going up still if we were trying to be going up still if we were trying to be more reasonable with our pace. Here's more reasonable with our pace. Here's more reasonable with our pace. Here's the issue with that belief. These scores the issue with that belief. These scores the issue with that belief. These scores don't represent a raw flat capability. don't represent a raw flat capability. don't represent a raw flat capability. They represent an average based on a lot They represent an average based on a lot They represent an average based on a lot of different things. People seem to of different things. People seem to of different things. People seem to think benchmarks measure something think benchmarks measure something think benchmarks measure something different than they do. This chart is different than they do. This chart is different than they do. This chart is meant to describe how I feel about meant to describe how I feel about meant to describe how I feel about response quality with models over time.
-
response quality with models over time. response quality with models over time. Astra, actually, I should move Astra up Astra, actually, I should move Astra up Astra, actually, I should move Astra up because Astra does have higher spikes because Astra does have higher spikes because Astra does have higher spikes than Fable. Like, it has moments that than Fable. Like, it has moments that than Fable. Like, it has moments that are better, but it also has moments that are better, but it also has moments that are better, but it also has moments that are a lot worse. The point of me drawing are a lot worse. The point of me drawing are a lot worse. The point of me drawing this diagram was to try and highlight this diagram was to try and highlight this diagram was to try and highlight why I am frustrated with Astra. It's not why I am frustrated with Astra. It's not why I am frustrated with Astra. It's not that it can't blow me away in specific that it can't blow me away in specific that it can't blow me away in specific moments. It's that sometimes Astra blows moments. It's that sometimes Astra blows moments. It's that sometimes Astra blows me away the wrong way. It is so much me away the wrong way. It is so much me away the wrong way. It is so much worse than I expected. And here is the worse than I expected. And here is the worse than I expected. And here is the different way I want youall to think different way I want youall to think different way I want youall to think about benchmarks. Because people seem to about benchmarks. Because people seem to about benchmarks. Because people seem to think benchmarks are measuring the think benchmarks are measuring the think benchmarks are measuring the highest point a model can get to. That highest point a model can get to. That highest point a model can get to. That asterisk score would be hereish because asterisk score would be hereish because asterisk score would be hereish because that's where its best scores and best that's where its best scores and best that's where its best scores and best capabilities are. And then fables would capabilities are. And then fables would capabilities are. And then fables would be there. But that's not how benchmarks be there. But that's not how benchmarks be there. But that's not how benchmarks work. Benchmarks have lots of different work. Benchmarks have lots of different work. Benchmarks have lots of different tasks throughout them. And sometimes the tasks throughout them. And sometimes the tasks throughout them. And sometimes the benchmark might have tasks that the benchmark might have tasks that the benchmark might have tasks that the model excels at. And sometimes it might model excels at. And sometimes it might model excels at. And sometimes it might have things the model is bad at. But have things the model is bad at. But have things the model is bad at. But more importantly, on any one given task, more importantly, on any one given task, more importantly, on any one given task, the model might perform well for a part the model might perform well for a part the model might perform well for a part and then really poorly for a part as and then really poorly for a part as and then really poorly for a part as well. These are the things that well. These are the things that well. These are the things that companies like OpenA and Anthropic are companies like OpenA and Anthropic are companies like OpenA and Anthropic are actively hunting for to destroy. They actively hunting for to destroy. They actively hunting for to destroy. They want to raise the floor as much as want to raise the floor as much as want to raise the floor as much as possible. Making the best moments better possible. Making the best moments better possible. Making the best moments better is only one side of this. Making it so is only one side of this. Making it so is only one side of this. Making it so the ultimate capability is higher is the ultimate capability is higher is the ultimate capability is higher is great and cool and fun, but making it so great and cool and fun, but making it so great and cool and fun, but making it so the worst moments are less bad and the worst moments are less bad and the worst moments are less bad and happen less often is even more happen less often is even more happen less often is even more important. Which of these models do you important. Which of these models do you important. Which of these models do you think would bench better? The one that think would bench better? The one that think would bench better? The one that has higher spikes and is above the other has higher spikes and is above the other has higher spikes and is above the other quite often or the one that has fewer
-
quite often or the one that has fewer quite often or the one that has fewer dips that goes into the bad space worse. dips that goes into the bad space worse. dips that goes into the bad space worse. point I want to make here is that a point I want to make here is that a point I want to make here is that a benchmark isn't measuring how smart is benchmark isn't measuring how smart is benchmark isn't measuring how smart is the model at its peaks. It's measuring the model at its peaks. It's measuring the model at its peaks. It's measuring how consistently is it succeeding and how consistently is it succeeding and how consistently is it succeeding and also how often is it failing. And it is also how often is it failing. And it is also how often is it failing. And it is absolutely the case that incredibly absolutely the case that incredibly absolutely the case that incredibly intelligent models can fail more often intelligent models can fail more often intelligent models can fail more often than less intelligent ones. If you don't than less intelligent ones. If you don't than less intelligent ones. If you don't believe me, throw GBD6 Astra on Max and believe me, throw GBD6 Astra on Max and believe me, throw GBD6 Astra on Max and ask it to fix a thing on your front end. ask it to fix a thing on your front end. ask it to fix a thing on your front end. You will quickly see what I mean. As You will quickly see what I mean. As You will quickly see what I mean. As incredibly capable as these models can incredibly capable as these models can incredibly capable as these models can be, a lot of benchmarks aren't just be, a lot of benchmarks aren't just be, a lot of benchmarks aren't just testing how well can it do at its best. testing how well can it do at its best. testing how well can it do at its best. They're also effectively testing how They're also effectively testing how They're also effectively testing how obnoxious are they at their worst. And obnoxious are they at their worst. And obnoxious are they at their worst. And if a company like Anthropic makes the if a company like Anthropic makes the if a company like Anthropic makes the bad moments happen less often, that bad moments happen less often, that bad moments happen less often, that might have a huge impact in the might have a huge impact in the might have a huge impact in the benchmarks that you might not feel with benchmarks that you might not feel with benchmarks that you might not feel with a handful of prompts. But here's the a handful of prompts. But here's the a handful of prompts. But here's the more important piece. Raising the more important piece. Raising the more important piece. Raising the ceiling inherently increases the ceiling inherently increases the ceiling inherently increases the potential danger and the potential risk. potential danger and the potential risk. potential danger and the potential risk. As models get more capable, if we don't As models get more capable, if we don't As models get more capable, if we don't have the moniability and comprehension, have the moniability and comprehension, have the moniability and comprehension, understanding keep up with those understanding keep up with those understanding keep up with those improvements, then the spikes can be improvements, then the spikes can be improvements, then the spikes can be more dangerous. If a model is incredibly more dangerous. If a model is incredibly more dangerous. If a model is incredibly intelligent, so much so that we want to intelligent, so much so that we want to intelligent, so much so that we want to use it to like power military vehicles use it to like power military vehicles use it to like power military vehicles and then it has a weird stupid spike and and then it has a weird stupid spike and and then it has a weird stupid spike and identifies a civilian as an enemy identifies a civilian as an enemy identifies a civilian as an enemy soldier and destroys them. That's really soldier and destroys them. That's really soldier and destroys them. That's really bad. So a much simpler way of putting bad. So a much simpler way of putting bad. So a much simpler way of putting this is that improving the floor for how this is that improving the floor for how this is that improving the floor for how bad models get when they're dumb and how bad models get when they're dumb and how bad models get when they're dumb and how often they act dumb does not increase often they act dumb does not increase often they act dumb does not increase risk but at the same time absolutely
-
risk but at the same time absolutely risk but at the same time absolutely increases benchmark scores. The quality increases benchmark scores. The quality increases benchmark scores. The quality of experience we have and the of experience we have and the of experience we have and the reliability of these models in models reliability of these models in models reliability of these models in models like Opus 5.5, GBD6, Soul, Luna, and Gro like Opus 5.5, GBD6, Soul, Luna, and Gro like Opus 5.5, GBD6, Soul, Luna, and Gro 47 aren't really pushing up the ceiling 47 aren't really pushing up the ceiling 47 aren't really pushing up the ceiling much. The thing that we gain from them, much. The thing that we gain from them, much. The thing that we gain from them, in particular from Opus 55, is that we in particular from Opus 55, is that we in particular from Opus 55, is that we can rely on them for longer workflows can rely on them for longer workflows can rely on them for longer workflows because the likelihood they do something because the likelihood they do something because the likelihood they do something stupid goes down. And this is awesome stupid goes down. And this is awesome stupid goes down. And this is awesome because it means we have a bar because it means we have a bar because it means we have a bar effectively like this is the ceiling for effectively like this is the ceiling for effectively like this is the ceiling for any given request. Anything above this any given request. Anything above this any given request. Anything above this and we should be concerned. But that and we should be concerned. But that and we should be concerned. But that doesn't stop us from cleaning up below. doesn't stop us from cleaning up below. doesn't stop us from cleaning up below. One of the best outcomes of pacing the One of the best outcomes of pacing the One of the best outcomes of pacing the frontier is if previously the worst frontier is if previously the worst frontier is if previously the worst cases were like down here. We can raise cases were like down here. We can raise cases were like down here. We can raise them up more and more and more until them up more and more and more until them up more and more and more until eventually the floor and the ceiling are eventually the floor and the ceiling are eventually the floor and the ceiling are nearly touching. But at no point do we nearly touching. But at no point do we nearly touching. But at no point do we actually increase the potential actually increase the potential actually increase the potential capability and potential danger of those capability and potential danger of those capability and potential danger of those same models. And that is what's so same models. And that is what's so same models. And that is what's so exciting about the releases that we're exciting about the releases that we're exciting about the releases that we're talking about now like 47, six soul, talking about now like 47, six soul, talking about now like 47, six soul, opus 55, etc. These models are becoming opus 55, etc. These models are becoming opus 55, etc. These models are becoming more reliable by being less stupid. Less more reliable by being less stupid. Less more reliable by being less stupid. Less stupid and more smart are different stupid and more smart are different stupid and more smart are different things. The amount of stupidity a thing things. The amount of stupidity a thing things. The amount of stupidity a thing has and the amount of intelligence a has and the amount of intelligence a has and the amount of intelligence a thing has can coexist. Look at GBD6 thing has can coexist. Look at GBD6 thing has can coexist. Look at GBD6 Astra. It is both the smartest model Astra. It is both the smartest model Astra. It is both the smartest model I've ever used and without question one I've ever used and without question one I've ever used and without question one of the stupidest I've used this year.
-
of the stupidest I've used this year. of the stupidest I've used this year. What's really exciting here is that we What's really exciting here is that we What's really exciting here is that we have systems that allow us to take good have systems that allow us to take good have systems that allow us to take good solutions to problems that agents and solutions to problems that agents and solutions to problems that agents and models and systems like this came up models and systems like this came up models and systems like this came up with and make smaller, dumber models act with and make smaller, dumber models act with and make smaller, dumber models act less dumb using that data. It's called less dumb using that data. It's called less dumb using that data. It's called RL. It is one of the most important RL. It is one of the most important RL. It is one of the most important things all of the labs do. So, we have things all of the labs do. So, we have things all of the labs do. So, we have these super capable models like Astra these super capable models like Astra these super capable models like Astra and Fable 51. And if we go through the and Fable 51. And if we go through the and Fable 51. And if we go through the data, we clean up the mistakes they data, we clean up the mistakes they data, we clean up the mistakes they make. We create training sets that are make. We create training sets that are make. We create training sets that are just the good, not the bad. And then we just the good, not the bad. And then we just the good, not the bad. And then we use that to beat the stupid out of Opus. use that to beat the stupid out of Opus. use that to beat the stupid out of Opus. The result is a very good model that The result is a very good model that The result is a very good model that kind of just feels like what we got with kind of just feels like what we got with kind of just feels like what we got with Fable, but cheaper and faster. And this Fable, but cheaper and faster. And this Fable, but cheaper and faster. And this is why I'm excited. The incentive the is why I'm excited. The incentive the is why I'm excited. The incentive the labs have now isn't to keep pushing labs have now isn't to keep pushing labs have now isn't to keep pushing models to be more and more capable of models to be more and more capable of models to be more and more capable of inventing doomsday devices or making inventing doomsday devices or making inventing doomsday devices or making boweapons that can wipe us out. They're boweapons that can wipe us out. They're boweapons that can wipe us out. They're trying to make the models less likely to trying to make the models less likely to trying to make the models less likely to delete our home directories and easier delete our home directories and easier delete our home directories and easier to understand why they did it when they to understand why they did it when they to understand why they did it when they do it. And in that same process, they're do it. And in that same process, they're do it. And in that same process, they're also making things more efficient and also making things more efficient and also making things more efficient and cheaper and less likely to do stupid cheaper and less likely to do stupid cheaper and less likely to do stupid [ __ ] This is what pacing looks like.
-
[ __ ] This is what pacing looks like. [ __ ] This is what pacing looks like. Less time spent training gigantic, super Less time spent training gigantic, super Less time spent training gigantic, super expensive models that can go higher and expensive models that can go higher and expensive models that can go higher and lower than we did before. in more time lower than we did before. in more time lower than we did before. in more time spent using what we can do right now to spent using what we can do right now to spent using what we can do right now to create better data to refine, fine-tune, create better data to refine, fine-tune, create better data to refine, fine-tune, reinforce, and most importantly, get us reinforce, and most importantly, get us reinforce, and most importantly, get us better performance with less flaws, less better performance with less flaws, less better performance with less flaws, less dumb mistakes in cheaper, smaller dumb mistakes in cheaper, smaller dumb mistakes in cheaper, smaller models. It's awesome, and I'm so excited models. It's awesome, and I'm so excited models. It's awesome, and I'm so excited because as much as I do want to see the because as much as I do want to see the because as much as I do want to see the capabilities increase over time, I also capabilities increase over time, I also capabilities increase over time, I also want cheaper models to be able to run want cheaper models to be able to run want cheaper models to be able to run longer and get more work done without longer and get more work done without longer and get more work done without doing something stupid in the process. doing something stupid in the process. doing something stupid in the process. And that's what I'm seeing here. What And that's what I'm seeing here. What And that's what I'm seeing here. What I'm seeing is the result of pacing. If I'm seeing is the result of pacing. If I'm seeing is the result of pacing. If we didn't decide to pace, the likelihood we didn't decide to pace, the likelihood we didn't decide to pace, the likelihood that Anthropic would have put so much that Anthropic would have put so much that Anthropic would have put so much time, effort, and marketing spend into time, effort, and marketing spend into time, effort, and marketing spend into Opus 55, it just it wouldn't have Opus 55, it just it wouldn't have Opus 55, it just it wouldn't have happened. Just just think about this. happened. Just just think about this. happened. Just just think about this. Like, as a user of Twitter or YouTube or Like, as a user of Twitter or YouTube or Like, as a user of Twitter or YouTube or whatever else, when's the last time whatever else, when's the last time whatever else, when's the last time anyone got hyped about a new model from anyone got hyped about a new model from anyone got hyped about a new model from Anthropic that wasn't their frontier Anthropic that wasn't their frontier Anthropic that wasn't their frontier level models? Nobody gave a [ __ ] about level models? Nobody gave a [ __ ] about level models? Nobody gave a [ __ ] about Sonnet at all other than Ben Davis. Poor Sonnet at all other than Ben Davis. Poor Sonnet at all other than Ben Davis. Poor kid. Nobody got excited about Opus 5.
-
kid. Nobody got excited about Opus 5. kid. Nobody got excited about Opus 5. Hell, people barely cared about Opus 48. Hell, people barely cared about Opus 48. Hell, people barely cared about Opus 48. The reason is because Anthropic didn't The reason is because Anthropic didn't The reason is because Anthropic didn't care. They don't even care enough about care. They don't even care enough about care. They don't even care enough about Haiku to release it. Like, Haiku has not Haiku to release it. Like, Haiku has not Haiku to release it. Like, Haiku has not had an update for 11 months. We're near had an update for 11 months. We're near had an update for 11 months. We're near a year since Haiku 45. Anthropic didn't a year since Haiku 45. Anthropic didn't a year since Haiku 45. Anthropic didn't care at all about models that were less care at all about models that were less care at all about models that were less capable than Astra. So, what they're capable than Astra. So, what they're capable than Astra. So, what they're doing now instead is putting a lot of doing now instead is putting a lot of doing now instead is putting a lot of effort and money and time and energy and effort and money and time and energy and effort and money and time and energy and marketing force and everything else into marketing force and everything else into marketing force and everything else into making Opus get as close as possible to making Opus get as close as possible to making Opus get as close as possible to Fable. And I think that's awesome. To Fable. And I think that's awesome. To Fable. And I think that's awesome. To everyone who's been upset about models everyone who's been upset about models everyone who's been upset about models getting too expensive and hard to use getting too expensive and hard to use getting too expensive and hard to use and justify because the price is just and justify because the price is just and justify because the price is just insane, I hope y'all are really in favor insane, I hope y'all are really in favor insane, I hope y'all are really in favor of pacing because pacing gets us what we of pacing because pacing gets us what we of pacing because pacing gets us what we want. Better models that are less stupid want. Better models that are less stupid want. Better models that are less stupid at cheaper prices. Not smarter models at cheaper prices. Not smarter models at cheaper prices. Not smarter models that can invent bioweapons. Cheaper that can invent bioweapons. Cheaper that can invent bioweapons. Cheaper models that can contribute to our code models that can contribute to our code models that can contribute to our code bases without deleting things from our bases without deleting things from our bases without deleting things from our computer accidentally. And I personally computer accidentally. And I personally computer accidentally. And I personally think that's awesome. And I hope this think that's awesome. And I hope this think that's awesome. And I hope this rant has helped convince you as well. rant has helped convince you as well. rant has helped convince you as well. We're living in a renaissance of less We're living in a renaissance of less We're living in a renaissance of less dumb models. Not smarter, less dumb. And dumb models. Not smarter, less dumb. And dumb models. Not smarter, less dumb. And I personally really, really like that. I I personally really, really like that. I I personally really, really like that. I hope you do, too. And at the very least, hope you do, too. And at the very least, hope you do, too. And at the very least, this makes you a little less skeptical this makes you a little less skeptical this makes you a little less skeptical of what pacing means. And you can see of what pacing means. And you can see of what pacing means. And you can see that these labs aren't lying. They are that these labs aren't lying. They are that these labs aren't lying. They are just trying to make better things given just trying to make better things given just trying to make better things given the circumstances that we're living the circumstances that we're living the circumstances that we're living through today. Let me know what you all through today. Let me know what you all through today. Let me know what you all think. Am I just being a blind chill or think. Am I just being a blind chill or think. Am I just being a blind chill or is this a more realistic way of looking is this a more realistic way of looking is this a more realistic way of looking at this stuff? And until next time, he at this stuff? And until next time, he at this stuff? And until next time, he sns.
No summary available yet.
View original episode ↗