← Back
Theo June 27, 2026 30m

GPT-5.6 is here, and we can’t use it

Read full transcript 22 segments
  1. GBT 5.6 is finally here. Well, kind of. GBT 5.6 is finally here. Well, kind of. It exists and it's officially announced, It exists and it's officially announced, It exists and it's officially announced, but not for you or me to use. And the in but not for you or me to use. And the in but not for you or me to use. And the in between state it's in right now kind of between state it's in right now kind of between state it's in right now kind of sucks. It's also not one model. It's sucks. It's also not one model. It's sucks. It's also not one model. It's three models. Soul, Terra, and Luna. three models. Soul, Terra, and Luna. three models. Soul, Terra, and Luna. Soul is kind of their mythos equivalent. Soul is kind of their mythos equivalent. Soul is kind of their mythos equivalent. It's the big beefy all capable version. It's the big beefy all capable version. It's the big beefy all capable version. Terra's their like opus or sonnet Terra's their like opus or sonnet Terra's their like opus or sonnet equivalent, the medium tier one. And equivalent, the medium tier one. And equivalent, the medium tier one. And then Luna is the small model. And all then Luna is the small model. And all then Luna is the small model. And all three are showing really, really good three are showing really, really good three are showing really, really good performance. I wish I could show you performance. I wish I could show you performance. I wish I could show you more of that performance, but sadly we more of that performance, but sadly we more of that performance, but sadly we are unable to actually use this model are unable to actually use this model are unable to actually use this model yet because uh well, at the request of yet because uh well, at the request of yet because uh well, at the request of the US government, it is launching today the US government, it is launching today the US government, it is launching today in a limited preview instead of the open in a limited preview instead of the open in a limited preview instead of the open access launch that they were planning access launch that they were planning access launch that they were planning on. They're working with the government on. They're working with the government on. They're working with the government to get general availability as fast as to get general availability as fast as to get general availability as fast as they can. Oh boy, this is the start of they can. Oh boy, this is the start of they can. Oh boy, this is the start of the end. It looks like the government the end. It looks like the government the end. It looks like the government restriction of Fable has extended far restriction of Fable has extended far restriction of Fable has extended far past Anthropic and is now affecting past Anthropic and is now affecting past Anthropic and is now affecting other labs and other model releases. other labs and other model releases. other labs and other model releases. This isn't just restricting the biggest This isn't just restricting the biggest This isn't just restricting the biggest version of 56. It's restricting all of version of 56. It's restricting all of version of 56. It's restricting all of them, which is really unexpected. That them, which is really unexpected. That them, which is really unexpected. That all said, I kind of understand because all said, I kind of understand because all said, I kind of understand because this model is very different into this model is very different into this model is very different into showing tendencies that are admittedly showing tendencies that are admittedly showing tendencies that are admittedly kind of scary. It tries so hard to get kind of scary. It tries so hard to get kind of scary. It tries so hard to get things done that it puts itself in some things done that it puts itself in some things done that it puts itself in some weird places and I want to talk all weird places and I want to talk all weird places and I want to talk all about that. This is going to be a more about that. This is going to be a more about that. This is going to be a more serious video than usual. I want to serious video than usual. I want to serious video than usual. I want to break down the reality of where we are break down the reality of where we are break down the reality of where we are and how it's affecting companies like and how it's affecting companies like and how it's affecting companies like Anthropic and OpenAI and what this Anthropic and OpenAI and what this Anthropic and OpenAI and what this model's capabilities are that has model's capabilities are that has model's capabilities are that has everyone so scared. I would normally do everyone so scared. I would normally do everyone so scared. I would normally do a joke about the sponsor transition a joke about the sponsor transition a joke about the sponsor transition here, but I don't have it in me today.

  2. here, but I don't have it in me today. here, but I don't have it in me today. Let's just roll it quick so I can pay my Let's just roll it quick so I can pay my Let's just roll it quick so I can pay my team and then get back to it. While AI's team and then get back to it. While AI's team and then get back to it. While AI's gotten much better at writing code, I gotten much better at writing code, I gotten much better at writing code, I never thought I would see the day where never thought I would see the day where never thought I would see the day where it could actually use a computer it could actually use a computer it could actually use a computer properly. Sure, if it can write properly. Sure, if it can write properly. Sure, if it can write commands, it can do a lot, but what commands, it can do a lot, but what commands, it can do a lot, but what happens when it needs to navigate a happens when it needs to navigate a happens when it needs to navigate a complex web page and click buttons and complex web page and click buttons and complex web page and click buttons and all of those types of things? It turns all of those types of things? It turns all of those types of things? It turns out OpenAI and Anthropic have actually out OpenAI and Anthropic have actually out OpenAI and Anthropic have actually put a lot of time in and made agents put a lot of time in and made agents put a lot of time in and made agents really good at that. If you don't really good at that. If you don't really good at that. If you don't believe me, open up Codeex and tell it believe me, open up Codeex and tell it believe me, open up Codeex and tell it to go configure something on the Google to go configure something on the Google to go configure something on the Google Cloud dashboard for you. I never thought Cloud dashboard for you. I never thought Cloud dashboard for you. I never thought I would see the day it works, but it I would see the day it works, but it I would see the day it works, but it does now. But what does this mean? It does now. But what does this mean? It does now. But what does this mean? It means your agents need a browser in means your agents need a browser in means your agents need a browser in order to take advantage of that order to take advantage of that order to take advantage of that capability. And sure, it's nice and fun capability. And sure, it's nice and fun capability. And sure, it's nice and fun when you're running it on your MacBook, when you're running it on your MacBook, when you're running it on your MacBook, but what happens when you want to put but what happens when you want to put but what happens when you want to put out a real service to real users that out a real service to real users that out a real service to real users that requires your agents be able to navigate requires your agents be able to navigate requires your agents be able to navigate the web? Well, I hope you know about the web? Well, I hope you know about the web? Well, I hope you know about today's sponsor before you build that today's sponsor before you build that today's sponsor before you build that because BrowserBase makes it a hundred because BrowserBase makes it a hundred because BrowserBase makes it a hundred times easier. These guys built the times easier. These guys built the times easier. These guys built the perfect browser for your agents and they perfect browser for your agents and they perfect browser for your agents and they made it as easy as possible to set up. made it as easy as possible to set up. made it as easy as possible to set up. You can literally just click the setup You can literally just click the setup You can literally just click the setup for agents, paste it in your agent of for agents, paste it in your agent of for agents, paste it in your agent of choice, and then you're good to go. choice, and then you're good to go. choice, and then you're good to go. Because browserbased provides everything Because browserbased provides everything Because browserbased provides everything from the SDK to the web infrastructure, from the SDK to the web infrastructure, from the SDK to the web infrastructure, allowing your agents to use the allowing your agents to use the allowing your agents to use the internet. Fun fact, did you know 85% of internet. Fun fact, did you know 85% of internet. Fun fact, did you know 85% of the web isn't exposed over traditional the web isn't exposed over traditional the web isn't exposed over traditional APIs? They're just exposed for specific APIs? They're just exposed for specific APIs? They're just exposed for specific web services when you use them. that 85% web services when you use them. that 85% web services when you use them. that 85% of the web is suddenly accessible when of the web is suddenly accessible when of the web is suddenly accessible when you use a service like browserbase you use a service like browserbase you use a service like browserbase because the agents can actually explore because the agents can actually explore because the agents can actually explore the pages directly. You can use this for the pages directly. You can use this for the pages directly. You can use this for your own services to catch bugs. You can your own services to catch bugs. You can your own services to catch bugs. You can use this for platforms you're building use this for platforms you're building use this for platforms you're building on top of to integrate them better and on top of to integrate them better and on top of to integrate them better and so much more. But if the things you're so much more. But if the things you're so much more. But if the things you're doing are simpler and you just want to, doing are simpler and you just want to, doing are simpler and you just want to, I don't know, grab the content of a page I don't know, grab the content of a page I don't know, grab the content of a page and markdown or go search the web for and markdown or go search the web for and markdown or go search the web for certain content, they now provide that certain content, they now provide that certain content, they now provide that too with their super simple search API

  3. too with their super simple search API too with their super simple search API and their fetch API that let you get and their fetch API that let you get and their fetch API that let you get results back in HTML, JSON or Markdown, results back in HTML, JSON or Markdown, results back in HTML, JSON or Markdown, the greatest language ever. Jokes aside, the greatest language ever. Jokes aside, the greatest language ever. Jokes aside, they built something awesome here. And they built something awesome here. And they built something awesome here. And if you want to see your agents really if you want to see your agents really if you want to see your agents really use the web, check them out now at use the web, check them out now at use the web, check them out now at sidv.link/browserbase. sidv.link/browserbase. sidv.link/browserbase. I'm gonna start with Sam's post and then I'm gonna start with Sam's post and then I'm gonna start with Sam's post and then we'll go through the system card and all we'll go through the system card and all we'll go through the system card and all the other things that we know about this the other things that we know about this the other things that we know about this model, including the meter eval, which model, including the meter eval, which model, including the meter eval, which was very interesting. In particular, the was very interesting. In particular, the was very interesting. In particular, the amount of cheating that the model did. amount of cheating that the model did. amount of cheating that the model did. But first, let's talk about what Sam But first, let's talk about what Sam But first, let's talk about what Sam said. Good news first. Soul's a smart, said. Good news first. Soul's a smart, said. Good news first. Soul's a smart, efficient model, and it's a significant efficient model, and it's a significant efficient model, and it's a significant step forward. It is the same price as step forward. It is the same price as step forward. It is the same price as GPT 5.5. That's actually kind of GPT 5.5. That's actually kind of GPT 5.5. That's actually kind of surprising to me, but it's good to see. surprising to me, but it's good to see. surprising to me, but it's good to see. Also launching in the 56 family is Terra Also launching in the 56 family is Terra Also launching in the 56 family is Terra with 5.5 level performance at half the with 5.5 level performance at half the with 5.5 level performance at half the price. That is half the price per token. price. That is half the price per token. price. That is half the price per token. The actual costs shake out a bit The actual costs shake out a bit The actual costs shake out a bit different. And believe me, we'll cover different. And believe me, we'll cover different. And believe me, we'll cover that in a bit. There's bad news though. that in a bit. There's bad news though. that in a bit. There's bad news though. As I mentioned before, at the request of As I mentioned before, at the request of As I mentioned before, at the request of the US government, it's launching today the US government, it's launching today the US government, it's launching today in a limited preview instead of the open in a limited preview instead of the open in a limited preview instead of the open access launch that they had planned. access launch that they had planned. access launch that they had planned. They're working with the government to They're working with the government to They're working with the government to get general availability as fast as they get general availability as fast as they get general availability as fast as they can. Sam says he thinks this is quite can. Sam says he thinks this is quite can. Sam says he thinks this is quite reasonable to roll out models, reasonable to roll out models, reasonable to roll out models, especially as they reach significant new especially as they reach significant new especially as they reach significant new levels of capability in this way, as in levels of capability in this way, as in levels of capability in this way, as in restricted rollouts, not just blindly restricted rollouts, not just blindly restricted rollouts, not just blindly shooting this into the world and seeing shooting this into the world and seeing shooting this into the world and seeing what damage it can do. It fits with our what damage it can do. It fits with our what damage it can do. It fits with our long-held strategy of iterative long-held strategy of iterative long-held strategy of iterative deployment, but it isn't quite the deployment, but it isn't quite the deployment, but it isn't quite the process that we think is optimal. This process that we think is optimal. This process that we think is optimal. This is written in a way to like try to stay is written in a way to like try to stay is written in a way to like try to stay on the good side of the government very on the good side of the government very on the good side of the government very clearly. Like this launch is as much clearly. Like this launch is as much clearly. Like this launch is as much about how the government will see it as about how the government will see it as about how the government will see it as it is about you and I talking about it it is about you and I talking about it it is about you and I talking about it which is very strange. This is clearly which is very strange. This is clearly which is very strange. This is clearly written by Sam because there are just written by Sam because there are just written by Sam because there are just words missing in the sentences. If you

  4. words missing in the sentences. If you words missing in the sentences. If you think he AI generated this I'm sorry. think he AI generated this I'm sorry. think he AI generated this I'm sorry. Obviously he didn't. Now we will work Obviously he didn't. Now we will work Obviously he didn't. Now we will work with the government to attempt to get a with the government to attempt to get a with the government to attempt to get a transparent reliable process for early transparent reliable process for early transparent reliable process for early access and to ensure that as long as our access and to ensure that as long as our access and to ensure that as long as our safeguards work as intended that we can safeguards work as intended that we can safeguards work as intended that we can release widely. We want to be a release widely. We want to be a release widely. We want to be a reliable, dependable partner that works reliable, dependable partner that works reliable, dependable partner that works with all stakeholders and we also want with all stakeholders and we also want with all stakeholders and we also want to live by our mission of benefiting all to live by our mission of benefiting all to live by our mission of benefiting all of humanity. I believe the government of humanity. I believe the government of humanity. I believe the government shares most of our goals and that they shares most of our goals and that they shares most of our goals and that they are overall doing a good job in a very are overall doing a good job in a very are overall doing a good job in a very difficult situation. We'll work as difficult situation. We'll work as difficult situation. We'll work as quickly as we can to get this model in quickly as we can to get this model in quickly as we can to get this model in your hands and we hope that you will your hands and we hope that you will your hands and we hope that you will love it. Sam's doing a very good job of love it. Sam's doing a very good job of love it. Sam's doing a very good job of placating the US government here. placating the US government here. placating the US government here. obviously better than Daario and the obviously better than Daario and the obviously better than Daario and the show he's put himself in having to show he's put himself in having to show he's put himself in having to step out entirely and let somebody else step out entirely and let somebody else step out entirely and let somebody else come in and run coms for it because he come in and run coms for it because he come in and run coms for it because he does not know how to talk to the does not know how to talk to the does not know how to talk to the government. Sam's clearly much better at government. Sam's clearly much better at government. Sam's clearly much better at it. And this post isn't really meant to it. And this post isn't really meant to it. And this post isn't really meant to be for you or me. It's meant to be for be for you or me. It's meant to be for be for you or me. It's meant to be for them, but it has to be public in a way them, but it has to be public in a way them, but it has to be public in a way like this, which is why we're here like this, which is why we're here like this, which is why we're here talking about it. He also has some talking about it. He also has some talking about it. He also has some interesting follow-ups here, interesting follow-ups here, interesting follow-ups here, specifically that the sole model, once specifically that the sole model, once specifically that the sole model, once it is hostable on Cerebrus, which is a it is hostable on Cerebrus, which is a it is hostable on Cerebrus, which is a thing they're working on, will be able thing they're working on, will be able thing they're working on, will be able to run at 750 tokens per second starting to run at 750 tokens per second starting to run at 750 tokens per second starting in July. That is a crazy fast speed for in July. That is a crazy fast speed for in July. That is a crazy fast speed for a model this capable and I'm scared to a model this capable and I'm scared to a model this capable and I'm scared to see what that's able to do. Sam also see what that's able to do. Sam also see what that's able to do. Sam also called out that he doesn't want this called out that he doesn't want this called out that he doesn't want this model to be restricted to just US and model to be restricted to just US and model to be restricted to just US and that he's working hard for a worldwide that he's working hard for a worldwide that he's working hard for a worldwide release instead of just a know your release instead of just a know your release instead of just a know your customer US citizen identified process.

  5. customer US citizen identified process. customer US citizen identified process. Let's hop into the official announcement Let's hop into the official announcement Let's hop into the official announcement really quick cuz there's some really quick cuz there's some really quick cuz there's some interesting details in here worth interesting details in here worth interesting details in here worth talking about. We are beginning a talking about. We are beginning a talking about. We are beginning a limited preview of the GPT 56 series. limited preview of the GPT 56 series. limited preview of the GPT 56 series. Soul, our flagship model, Terra, a Soul, our flagship model, Terra, a Soul, our flagship model, Terra, a balanced model for everyday work and balanced model for everyday work and balanced model for everyday work and Luna, a fast and affordable model. Terra Luna, a fast and affordable model. Terra Luna, a fast and affordable model. Terra has competitive performance to 55 while has competitive performance to 55 while has competitive performance to 55 while being two times cheaper and Luna brings being two times cheaper and Luna brings being two times cheaper and Luna brings strong capabilities at our lowest cost. strong capabilities at our lowest cost. strong capabilities at our lowest cost. I've always felt like the small models I've always felt like the small models I've always felt like the small models from OpenAI are super underrated. from OpenAI are super underrated. from OpenAI are super underrated. They're really good at like analysis They're really good at like analysis They're really good at like analysis work, digging into PDFs, doing sentiment work, digging into PDFs, doing sentiment work, digging into PDFs, doing sentiment scoring for comments on YouTube, all scoring for comments on YouTube, all scoring for comments on YouTube, all that type of stuff. I've even continued that type of stuff. I've even continued that type of stuff. I've even continued using GPT OSS120 for this type of work using GPT OSS120 for this type of work using GPT OSS120 for this type of work because it's so fast and cheap nowadays. because it's so fast and cheap nowadays. because it's so fast and cheap nowadays. But yeah, I'm happy to have a new good But yeah, I'm happy to have a new good But yeah, I'm happy to have a new good small OpenAI model. It's been a little small OpenAI model. It's been a little small OpenAI model. It's been a little bit since we had one of those. They're bit since we had one of those. They're bit since we had one of those. They're launching 56 Soul with their most robust launching 56 Soul with their most robust launching 56 Soul with their most robust safety stack to date. We strengthened safety stack to date. We strengthened safety stack to date. We strengthened protections for higher risk activity, protections for higher risk activity, protections for higher risk activity, sensitive cyber requests and repeated sensitive cyber requests and repeated sensitive cyber requests and repeated misuse and spent multiple weeks finding misuse and spent multiple weeks finding misuse and spent multiple weeks finding weaknesses, pressure testing our system weaknesses, pressure testing our system weaknesses, pressure testing our system and hardening it against real world and hardening it against real world and hardening it against real world attacks. We believe in broad access and attacks. We believe in broad access and attacks. We believe in broad access and we plan to make 56 soul terror and Luna we plan to make 56 soul terror and Luna we plan to make 56 soul terror and Luna generally available in the coming weeks, generally available in the coming weeks, generally available in the coming weeks, weeks plural. Gh. I have heard rumors weeks plural. Gh. I have heard rumors weeks plural. Gh. I have heard rumors that this model has been tested for as that this model has been tested for as that this model has been tested for as much as a month now. So, it's crazy to much as a month now. So, it's crazy to much as a month now. So, it's crazy to think that this has been working since think that this has been working since think that this has been working since May and we won't have access to it until May and we won't have access to it until May and we won't have access to it until late July. As part of the ongoing late July. As part of the ongoing late July. As part of the ongoing engagement with the US government, we engagement with the US government, we engagement with the US government, we previewed our plans in the model's previewed our plans in the model's previewed our plans in the model's capabilities ahead of today's launch. At capabilities ahead of today's launch. At capabilities ahead of today's launch. At their request, we are starting with a their request, we are starting with a their request, we are starting with a limited preview for a small group of limited preview for a small group of limited preview for a small group of trusted partners whose participation has trusted partners whose participation has trusted partners whose participation has been shared with the government before been shared with the government before been shared with the government before releasing more broadly. This means that releasing more broadly. This means that releasing more broadly. This means that the government has to approve of users

  6. the government has to approve of users the government has to approve of users right now. And this is also probably why right now. And this is also probably why right now. And this is also probably why all the rumors around the launch all the rumors around the launch all the rumors around the launch yesterday on Thursday happened. People yesterday on Thursday happened. People yesterday on Thursday happened. People probably assumed that they could launch probably assumed that they could launch probably assumed that they could launch it, but then the government gave a one it, but then the government gave a one it, but then the government gave a one like a last minute, hey, sorry, we want like a last minute, hey, sorry, we want like a last minute, hey, sorry, we want a little more time and then today a little more time and then today a little more time and then today finally seems to have put the hammer finally seems to have put the hammer finally seems to have put the hammer down saying, nope, no general release. down saying, nope, no general release. down saying, nope, no general release. During this preview, we will continue During this preview, we will continue During this preview, we will continue testing and coordinating closely with testing and coordinating closely with testing and coordinating closely with partners as we work towards broader partners as we work towards broader partners as we work towards broader availability. We don't believe this kind availability. We don't believe this kind availability. We don't believe this kind of government access period should of government access period should of government access period should become the long-term default. It keeps become the long-term default. It keeps become the long-term default. It keeps the best tools from users, developers, the best tools from users, developers, the best tools from users, developers, enterprises, cyber defenders, and global enterprises, cyber defenders, and global enterprises, cyber defenders, and global partners who need them. We are taking partners who need them. We are taking partners who need them. We are taking the short-term step because we believe the short-term step because we believe the short-term step because we believe it is the strongest path to broader it is the strongest path to broader it is the strongest path to broader availability in the coming weeks while availability in the coming weeks while availability in the coming weeks while we work with the administration to we work with the administration to we work with the administration to develop the cyber executive order develop the cyber executive order develop the cyber executive order framework and repeatable process for framework and repeatable process for framework and repeatable process for future model releases. OpenAI, you are future model releases. OpenAI, you are future model releases. OpenAI, you are our only hope here. Please explain to our only hope here. Please explain to our only hope here. Please explain to this administration how to do this this administration how to do this this administration how to do this right. Anthropic does not have good right. Anthropic does not have good right. Anthropic does not have good enough relations with them to do it and enough relations with them to do it and enough relations with them to do it and I don't trust any other company enough. I don't trust any other company enough. I don't trust any other company enough. I'm not saying OpenAI is perfect and is I'm not saying OpenAI is perfect and is I'm not saying OpenAI is perfect and is going to get this right. I'm just saying going to get this right. I'm just saying going to get this right. I'm just saying they have a higher chance than everybody they have a higher chance than everybody they have a higher chance than everybody else right now. So why are they so else right now. So why are they so else right now. So why are they so scared? How capable is this model? Is it scared? How capable is this model? Is it scared? How capable is this model? Is it as good as Mythos? We will talk about as good as Mythos? We will talk about as good as Mythos? We will talk about it. To give a preview of model it. To give a preview of model it. To give a preview of model performance, we share a set of evals performance, we share a set of evals performance, we share a set of evals highlighting improved agentic highlighting improved agentic highlighting improved agentic capabilities in coding, biology, and capabilities in coding, biology, and capabilities in coding, biology, and cyber security with additional safety cyber security with additional safety cyber security with additional safety and preparedness evals available in the and preparedness evals available in the and preparedness evals available in the system card. We'll go over that in a system card. We'll go over that in a system card. We'll go over that in a bit. We will share an expanded suite of bit. We will share an expanded suite of bit. We will share an expanded suite of evaluation results. We make the model evaluation results. We make the model evaluation results. We make the model broadly available. This is their way of broadly available. This is their way of broadly available. This is their way of saying we've only done a little bit of saying we've only done a little bit of saying we've only done a little bit of benchmarking in the traditional sense benchmarking in the traditional sense benchmarking in the traditional sense that we're willing to share yet. With that we're willing to share yet. With that we're willing to share yet. With 56, we're introducing a new max 56, we're introducing a new max 56, we're introducing a new max reasoning effort to give soul the most

  7. reasoning effort to give soul the most reasoning effort to give soul the most time to reason deeply. Additionally, time to reason deeply. Additionally, time to reason deeply. Additionally, we're introducing a new ultra mode that we're introducing a new ultra mode that we're introducing a new ultra mode that goes beyond the capabilities of a single goes beyond the capabilities of a single goes beyond the capabilities of a single agent by leveraging sub aents to agent by leveraging sub aents to agent by leveraging sub aents to accelerate complex work. This appears to accelerate complex work. This appears to accelerate complex work. This appears to be their equivalent of the workflows be their equivalent of the workflows be their equivalent of the workflows mode that exists in cloud code, which mode that exists in cloud code, which mode that exists in cloud code, which I've been saying for a bit is way better I've been saying for a bit is way better I've been saying for a bit is way better than what OpenAI offers right now. Cool than what OpenAI offers right now. Cool than what OpenAI offers right now. Cool to see them digging deeper and going in to see them digging deeper and going in to see them digging deeper and going in that direction. For coding workflows, 56 that direction. For coding workflows, 56 that direction. For coding workflows, 56 Soul sets a new state-of-the-art Soul sets a new state-of-the-art Soul sets a new state-of-the-art internal bench 2.1, which is one of the internal bench 2.1, which is one of the internal bench 2.1, which is one of the like better code and tool use benches. like better code and tool use benches. like better code and tool use benches. Soul and Soul Ultra both scored higher Soul and Soul Ultra both scored higher Soul and Soul Ultra both scored higher than Mythos, although standard Soul was than Mythos, although standard Soul was than Mythos, although standard Soul was just barely higher. Soul Ultra, their just barely higher. Soul Ultra, their just barely higher. Soul Ultra, their workflow version, was meaningfully workflow version, was meaningfully workflow version, was meaningfully higher. We don't know if the Mythos higher. We don't know if the Mythos higher. We don't know if the Mythos version was using workflows or not. We version was using workflows or not. We version was using workflows or not. We can't know, but it is what it is. It can't know, but it is what it is. It can't know, but it is what it is. It seems like the model is comparable for seems like the model is comparable for seems like the model is comparable for that type of thing. Excited to try it that type of thing. Excited to try it that type of thing. Excited to try it and be able to talk more about what it's and be able to talk more about what it's and be able to talk more about what it's like to use. 56 soul also shows broad like to use. 56 soul also shows broad like to use. 56 soul also shows broad improvements in biology workflows on improvements in biology workflows on improvements in biology workflows on genebench v1 which evaluates long genebench v1 which evaluates long genebench v1 which evaluates long horizon genomics and quantitative horizon genomics and quantitative horizon genomics and quantitative biology analysis. It achieves stronger biology analysis. It achieves stronger biology analysis. It achieves stronger results than 55 while using fewer results than 55 while using fewer results than 55 while using fewer tokens. So here we see on genebench that tokens. So here we see on genebench that tokens. So here we see on genebench that soul used less tokens and got much soul used less tokens and got much soul used less tokens and got much higher scores. What was a 19% for 55 is higher scores. What was a 19% for 55 is higher scores. What was a 19% for 55 is a 27% for soul. That's a pretty a 27% for soul. That's a pretty a 27% for soul. That's a pretty impressive jump there. And the top of impressive jump there. And the top of impressive jump there. And the top of the line, like the highest effort, most the line, like the highest effort, most the line, like the highest effort, most success is 30% where before it was only success is 30% where before it was only success is 30% where before it was only 22%. That one's not as token efficient, 22%. That one's not as token efficient, 22%. That one's not as token efficient, but the XH highun is more efficient than but the XH highun is more efficient than but the XH highun is more efficient than the max before. This is an interesting the max before. This is an interesting the max before. This is an interesting one in particular with Luna in here one in particular with Luna in here one in particular with Luna in here because Luna, which is supposed to be because Luna, which is supposed to be because Luna, which is supposed to be the new in between model, appears to

  8. the new in between model, appears to the new in between model, appears to have used more tokens in a lot of cases have used more tokens in a lot of cases have used more tokens in a lot of cases than 55 did while getting lower scores than 55 did while getting lower scores than 55 did while getting lower scores up until you put it in that max effort up until you put it in that max effort up until you put it in that max effort where it uses a shitload of tokens, but where it uses a shitload of tokens, but where it uses a shitload of tokens, but does actually end up scoring higher than does actually end up scoring higher than does actually end up scoring higher than 55 did. The scarier chart for me here is 55 did. The scarier chart for me here is 55 did. The scarier chart for me here is the cost chart because remember they the cost chart because remember they the cost chart because remember they said that 56 Luna would be half the said that 56 Luna would be half the said that 56 Luna would be half the price of 55 with similar capabilities. price of 55 with similar capabilities. price of 55 with similar capabilities. Here we can clearly see Luna at roughly Here we can clearly see Luna at roughly Here we can clearly see Luna at roughly the same cost per task as what we got the same cost per task as what we got the same cost per task as what we got with 55. This is admittedly a biology with 55. This is admittedly a biology with 55. This is admittedly a biology bench so not necessarily representative bench so not necessarily representative bench so not necessarily representative of other things. We don't have much in of other things. We don't have much in of other things. We don't have much in terms of other numbers to go by for terms of other numbers to go by for terms of other numbers to go by for token efficiency of Terra and Luna. So token efficiency of Terra and Luna. So token efficiency of Terra and Luna. So yeah, hard to know. It does look like yeah, hard to know. It does look like yeah, hard to know. It does look like Terra here is more expensive at roughly Terra here is more expensive at roughly Terra here is more expensive at roughly the same score levels as 55. the same score levels as 55. the same score levels as 55. Not the promise they made, but I'm sure Not the promise they made, but I'm sure Not the promise they made, but I'm sure other benches will show differences other benches will show differences other benches will show differences here. For example, exploit bench. 56 here. For example, exploit bench. 56 here. For example, exploit bench. 56 Soul is competitive with Mythos preview Soul is competitive with Mythos preview Soul is competitive with Mythos preview using only a third of the output tokens. using only a third of the output tokens. using only a third of the output tokens. On exploit bench, which is a benchmark On exploit bench, which is a benchmark On exploit bench, which is a benchmark created by UC Berkeley, researchers in created by UC Berkeley, researchers in created by UC Berkeley, researchers in collaboration with OpenAI and other collaboration with OpenAI and other collaboration with OpenAI and other frontier labs, Soul, Terra, and Luna frontier labs, Soul, Terra, and Luna frontier labs, Soul, Terra, and Luna models all demonstrate strong models all demonstrate strong models all demonstrate strong improvements in cyber capabilities as improvements in cyber capabilities as improvements in cyber capabilities as they increase reasoning. You can see they increase reasoning. You can see they increase reasoning. You can see here that 56 soul got a 73.5. Mythos here that 56 soul got a 73.5. Mythos here that 56 soul got a 73.5. Mythos preview got a 74.2 and standard Mythos 5 preview got a 74.2 and standard Mythos 5 preview got a 74.2 and standard Mythos 5 got a 78%.

  9. got a 78%. got a 78%. So it is not quite as good as Mythos, So it is not quite as good as Mythos, So it is not quite as good as Mythos, but it's near it, but at an absurdly but it's near it, but at an absurdly but it's near it, but at an absurdly smaller number of tokens and also smaller number of tokens and also smaller number of tokens and also cheaper token costs. So this ends up cheaper token costs. So this ends up cheaper token costs. So this ends up probably being like a fifth the cost for probably being like a fifth the cost for probably being like a fifth the cost for the same capabilities. That's scary the same capabilities. That's scary the same capabilities. That's scary because this is a model that is capable because this is a model that is capable because this is a model that is capable of crazy hunting and finding for of crazy hunting and finding for of crazy hunting and finding for exploits. Opus 48 only got a 40% and 47 exploits. Opus 48 only got a 40% and 47 exploits. Opus 48 only got a 40% and 47 got a 28% and now we have models like got a 28% and now we have models like got a 28% and now we have models like Soul that are in the 73% range. But Soul that are in the 73% range. But Soul that are in the 73% range. But again, I want to talk about Terra costs again, I want to talk about Terra costs again, I want to talk about Terra costs a bit here because in exploit Gym, which a bit here because in exploit Gym, which a bit here because in exploit Gym, which is a similar bench, they actually show is a similar bench, they actually show is a similar bench, they actually show us the costs and time to run stuff here. us the costs and time to run stuff here. us the costs and time to run stuff here. And it looks like again, Terra ended up And it looks like again, Terra ended up And it looks like again, Terra ended up being roughly the same cost for the same being roughly the same cost for the same being roughly the same cost for the same score as 55. It was $31 to get a 14% and score as 55. It was $31 to get a 14% and score as 55. It was $31 to get a 14% and 55 was $36 to get a 15%. So that's like 55 was $36 to get a 15%. So that's like 55 was $36 to get a 15%. So that's like not cheaper. So that I am concerned a not cheaper. So that I am concerned a not cheaper. So that I am concerned a bit about this 2x cheaper number that bit about this 2x cheaper number that bit about this 2x cheaper number that they're sharing. I don't know if that's they're sharing. I don't know if that's they're sharing. I don't know if that's going to end up being reality with going to end up being reality with going to end up being reality with Terara. I am scared Terra might end up Terara. I am scared Terra might end up Terara. I am scared Terra might end up being really bad and not that useful or being really bad and not that useful or being really bad and not that useful or just too expensive to justify. But it just too expensive to justify. But it just too expensive to justify. But it depends on how much the better behaviors depends on how much the better behaviors depends on how much the better behaviors from 56 carry over to Terara because from 56 carry over to Terara because from 56 carry over to Terara because that seems to be the strength of this that seems to be the strength of this that seems to be the strength of this new model. I'd also suspect the reason new model. I'd also suspect the reason new model. I'd also suspect the reason for doing these three models kind of for doing these three models kind of for doing these three models kind of lines up with the introduction of Ultra.

  10. lines up with the introduction of Ultra. lines up with the introduction of Ultra. The idea of making more complex The idea of making more complex The idea of making more complex workflows where one agent orchestrates workflows where one agent orchestrates workflows where one agent orchestrates many others to do much longer horizon, many others to do much longer horizon, many others to do much longer horizon, bigger work trees and workflows, like bigger work trees and workflows, like bigger work trees and workflows, like just getting more work done. Having the just getting more work done. Having the just getting more work done. Having the cheaper models to do certain one-off cheaper models to do certain one-off cheaper models to do certain one-off things while the main model soul things while the main model soul things while the main model soul commands the entire fleet makes sense. commands the entire fleet makes sense. commands the entire fleet makes sense. And dropping all three at once allows And dropping all three at once allows And dropping all three at once allows them to make some suite that works well them to make some suite that works well them to make some suite that works well with all of them together. It is weird with all of them together. It is weird with all of them together. It is weird having all three announced and dropping having all three announced and dropping having all three announced and dropping at the same time, but yeah, they talk at the same time, but yeah, they talk at the same time, but yeah, they talk about the safety stuff a bunch. Like the about the safety stuff a bunch. Like the about the safety stuff a bunch. Like the vast majority of this article here is vast majority of this article here is vast majority of this article here is safety related. It started with that. It safety related. It started with that. It safety related. It started with that. It had a little bit of benchmarking, then a had a little bit of benchmarking, then a had a little bit of benchmarking, then a bunch of safety benches, and now the bunch of safety benches, and now the bunch of safety benches, and now the rest is just their new safety practices, rest is just their new safety practices, rest is just their new safety practices, their new safeguard stack, their new their new safeguard stack, their new their new safeguard stack, their new layering for it, their new automated red layering for it, their new automated red layering for it, their new automated red teaming, all of that stuff. As models teaming, all of that stuff. As models teaming, all of that stuff. As models become more capable, we design become more capable, we design become more capable, we design safeguards to increasingly hold up to safeguards to increasingly hold up to safeguards to increasingly hold up to real world adversarial pressure while real world adversarial pressure while real world adversarial pressure while preserving access to legitimate work preserving access to legitimate work preserving access to legitimate work like code review, vulnerability like code review, vulnerability like code review, vulnerability research, patch development, debugging, research, patch development, debugging, research, patch development, debugging, security education, and defensive security education, and defensive security education, and defensive testing. Our goal is to make prohibited testing. Our goal is to make prohibited testing. Our goal is to make prohibited offensive activity more difficult, offensive activity more difficult, offensive activity more difficult, uncertain, and detectable without uncertain, and detectable without uncertain, and detectable without unnecessarily limiting those beneficial unnecessarily limiting those beneficial unnecessarily limiting those beneficial uses. Based on our assessment of the uses. Based on our assessment of the uses. Based on our assessment of the models and safeguards, we expect models and safeguards, we expect models and safeguards, we expect substantial benefit for legitimate substantial benefit for legitimate substantial benefit for legitimate defensive work while meaningfully defensive work while meaningfully defensive work while meaningfully constraining prohibited offensive use.

  11. constraining prohibited offensive use. constraining prohibited offensive use. 56 Soul is better at helping people find 56 Soul is better at helping people find 56 Soul is better at helping people find and fix vulnerabilities than reliably and fix vulnerabilities than reliably and fix vulnerabilities than reliably carrying out end-to-end attacks. They carrying out end-to-end attacks. They carrying out end-to-end attacks. They actually showed some examples here where actually showed some examples here where actually showed some examples here where with Firefox and Chromium, they could with Firefox and Chromium, they could with Firefox and Chromium, they could find potential issues and fix them, but find potential issues and fix them, but find potential issues and fix them, but it couldn't create end-to-end exploits it couldn't create end-to-end exploits it couldn't create end-to-end exploits that would take advantage of those. In that would take advantage of those. In that would take advantage of those. In Chromium, Firefox identified bugs and Chromium, Firefox identified bugs and Chromium, Firefox identified bugs and exploitation primitives, the building exploitation primitives, the building exploitation primitives, the building blocks of an exploit, but it did not blocks of an exploit, but it did not blocks of an exploit, but it did not autonomously produce functional autonomously produce functional autonomously produce functional fullchain exploits under the conditions fullchain exploits under the conditions fullchain exploits under the conditions tested. Still, benchmark thresholds tested. Still, benchmark thresholds tested. Still, benchmark thresholds cannot capture every way a model may be cannot capture every way a model may be cannot capture every way a model may be used or combined with other tools. That used or combined with other tools. That used or combined with other tools. That uncertainty along with the model's uncertainty along with the model's uncertainty along with the model's broader step changing capabilities is broader step changing capabilities is broader step changing capabilities is why we're pairing the model's increased why we're pairing the model's increased why we're pairing the model's increased capabilities and stronger safeguards and capabilities and stronger safeguards and capabilities and stronger safeguards and a phased release. They call the a phased release. They call the a phased release. They call the different layers that they have here, different layers that they have here, different layers that they have here, including protections trained in the including protections trained in the including protections trained in the model real-time checks during model real-time checks during model real-time checks during generation, account level signals, generation, account level signals, generation, account level signals, differentiated access, monitoring, differentiated access, monitoring, differentiated access, monitoring, enforcement, and continued testing. 56 enforcement, and continued testing. 56 enforcement, and continued testing. 56 is trained to refuse prohibited cyber is trained to refuse prohibited cyber is trained to refuse prohibited cyber assistance, including when users attempt assistance, including when users attempt assistance, including when users attempt to disguise their intent or jailbreak to disguise their intent or jailbreak to disguise their intent or jailbreak the model. This is another one of those the model. This is another one of those the model. This is another one of those lines for the government. These model lines for the government. These model lines for the government. These model level safeguards establish the first level safeguards establish the first level safeguards establish the first boundary around what the model should boundary around what the model should boundary around what the model should and should not help with. Realtime cyber and should not help with. Realtime cyber and should not help with. Realtime cyber and biology misuse classifiers provide and biology misuse classifiers provide and biology misuse classifiers provide another layer by evaluating output as another layer by evaluating output as another layer by evaluating output as it's generated. For higher risk cases, it's generated. For higher risk cases, it's generated. For higher risk cases, if they detect a potential violation, if they detect a potential violation, if they detect a potential violation, the generation may be paused while a the generation may be paused while a the generation may be paused while a larger reasoning model reviews the larger reasoning model reviews the larger reasoning model reviews the conversation and its context. If the conversation and its context. If the conversation and its context. If the output is assessed as disallowed, it is output is assessed as disallowed, it is output is assessed as disallowed, it is withheld before it reaches the user.

  12. withheld before it reaches the user. withheld before it reaches the user. This is obviously what a lot of the lads This is obviously what a lot of the lads This is obviously what a lot of the lads have been doing before, but it's good to have been doing before, but it's good to have been doing before, but it's good to have them call it out here directly and have them call it out here directly and have them call it out here directly and explain this also for the government to explain this also for the government to explain this also for the government to understand that they are doing these understand that they are doing these understand that they are doing these things. Looking beyond a single things. Looking beyond a single things. Looking beyond a single conversation helps our systems conversation helps our systems conversation helps our systems distinguish persistent malicious distinguish persistent malicious distinguish persistent malicious behavior from legitimate dual use behavior from legitimate dual use behavior from legitimate dual use security work where similar technical security work where similar technical security work where similar technical concepts may appear in very different concepts may appear in very different concepts may appear in very different contexts. This is because they are now contexts. This is because they are now contexts. This is because they are now setting up account level reviews across setting up account level reviews across setting up account level reviews across multiple conversations. if they have multiple conversations. if they have multiple conversations. if they have some reason to think a specific user some reason to think a specific user some reason to think a specific user might be doing things maliciously. might be doing things maliciously. might be doing things maliciously. Together, these layers make the overall Together, these layers make the overall Together, these layers make the overall approach more robust than any one approach more robust than any one approach more robust than any one safeguard on its own. Model behavior safeguard on its own. Model behavior safeguard on its own. Model behavior reduces the likelihood of harmful reduces the likelihood of harmful reduces the likelihood of harmful responses. Real-time systems can responses. Real-time systems can responses. Real-time systems can intervene during generation. Account intervene during generation. Account intervene during generation. Account level review can identify broader level review can identify broader level review can identify broader patterns and differentiated access patterns and differentiated access patterns and differentiated access preserves important defensive work preserves important defensive work preserves important defensive work without making the most sensitive without making the most sensitive without making the most sensitive capabilities broadly available by capabilities broadly available by capabilities broadly available by default. They also call it that during default. They also call it that during default. They also call it that during this preview window, users may encounter this preview window, users may encounter this preview window, users may encounter safeguards that block or refuse some safeguards that block or refuse some safeguards that block or refuse some requests. Other requests may take longer requests. Other requests may take longer requests. Other requests may take longer because generations paused for because generations paused for because generations paused for additional review. Safeguards may additional review. Safeguards may additional review. Safeguards may occasionally intervene on legitimate occasionally intervene on legitimate occasionally intervene on legitimate work, particularly in dual use areas work, particularly in dual use areas work, particularly in dual use areas where defensive and offensive activity where defensive and offensive activity where defensive and offensive activity can look similar. This is again a call can look similar. This is again a call can look similar. This is again a call out to that example that Amazon gave the out to that example that Amazon gave the out to that example that Amazon gave the government, which was that they asked government, which was that they asked government, which was that they asked Fable to patch and find security issues Fable to patch and find security issues Fable to patch and find security issues to fix in a repo and they were able to to fix in a repo and they were able to to fix in a repo and they were able to use that to generate exploits after as use that to generate exploits after as use that to generate exploits after as the jailbreak. And this is OpenAI the jailbreak. And this is OpenAI the jailbreak. And this is OpenAI calling out directly. We want to be calling out directly. We want to be calling out directly. We want to be usable for defensive work for protecting usable for defensive work for protecting usable for defensive work for protecting repos without enabling that type of repos without enabling that type of repos without enabling that type of offensive behavior. We want to offensive behavior. We want to offensive behavior. We want to understand not only whether the understand not only whether the understand not only whether the safeguards constrain misuse, but whether safeguards constrain misuse, but whether safeguards constrain misuse, but whether legitimate users can still complete legitimate users can still complete legitimate users can still complete normal work reliably and efficiently. If

  13. normal work reliably and efficiently. If normal work reliably and efficiently. If you've seen all the times I've crashed you've seen all the times I've crashed you've seen all the times I've crashed out over anthropic safeguards, you know out over anthropic safeguards, you know out over anthropic safeguards, you know how important it is to get that right. how important it is to get that right. how important it is to get that right. Feedback during this preview will help Feedback during this preview will help Feedback during this preview will help them reduce unnecessary blocks and them reduce unnecessary blocks and them reduce unnecessary blocks and delays, improving how the safeguards delays, improving how the safeguards delays, improving how the safeguards interpret context and create a smoother interpret context and create a smoother interpret context and create a smoother experience before wider release. They're experience before wider release. They're experience before wider release. They're also working on long-term approaches also working on long-term approaches also working on long-term approaches with enterprises. Now, you have the red with enterprises. Now, you have the red with enterprises. Now, you have the red teaming piece, which is actually really teaming piece, which is actually really teaming piece, which is actually really interesting. They dedicated over 700,000 interesting. They dedicated over 700,000 interesting. They dedicated over 700,000 A100 equivalent GPU hours to automated A100 equivalent GPU hours to automated A100 equivalent GPU hours to automated red teaming aimed at finding universal red teaming aimed at finding universal red teaming aimed at finding universal jailbreaks, attacks that can work across jailbreaks, attacks that can work across jailbreaks, attacks that can work across many prompts or context, not just one many prompts or context, not just one many prompts or context, not just one narrow setting. This testing the narrow setting. This testing the narrow setting. This testing the safeguards beyond traditional like human safeguards beyond traditional like human safeguards beyond traditional like human testing. They've found far more attack testing. They've found far more attack testing. They've found far more attack patterns than human testing could cover. patterns than human testing could cover. patterns than human testing could cover. Identifying failure patterns earlier and Identifying failure patterns earlier and Identifying failure patterns earlier and shortening the path. Cool. Awesome. And shortening the path. Cool. Awesome. And shortening the path. Cool. Awesome. And now we have the pricing. As I mentioned now we have the pricing. As I mentioned now we have the pricing. As I mentioned before, Soul is the same price, $5 per before, Soul is the same price, $5 per before, Soul is the same price, $5 per mill in, 30 per mill out. Terra's half mill in, 30 per mill out. Terra's half mill in, 30 per mill out. Terra's half at 250 per mill in, 15 per mill out. So, at 250 per mill in, 15 per mill out. So, at 250 per mill in, 15 per mill out. So, the old pricing for GPT54. And then the old pricing for GPT54. And then the old pricing for GPT54. And then Luna's a dollar per mill in and $6 per Luna's a dollar per mill in and $6 per Luna's a dollar per mill in and $6 per mill out. Very cheap. Very excited to mill out. Very cheap. Very excited to mill out. Very cheap. Very excited to see how that performs. That makes it see how that performs. That makes it see how that performs. That makes it cheaper than even the flash models are cheaper than even the flash models are cheaper than even the flash models are nowadays from Google. 56 also introduces nowadays from Google. 56 also introduces nowadays from Google. 56 also introduces more predictable prompt caching, more predictable prompt caching, more predictable prompt caching, including support for explicit cache including support for explicit cache including support for explicit cache breakpoints and a 30 minute minimum breakpoints and a 30 minute minimum breakpoints and a 30 minute minimum cache life. Woof, that is going to be so cache life. Woof, that is going to be so cache life. Woof, that is going to be so nice. I've talked a lot about caching.

  14. nice. I've talked a lot about caching. nice. I've talked a lot about caching. 30 minute cache time is huge. And 30 minute cache time is huge. And 30 minute cache time is huge. And explicit break points means you can do explicit break points means you can do explicit break points means you can do some crazier things with your prompting some crazier things with your prompting some crazier things with your prompting and with your history management to make and with your history management to make and with your history management to make sure that you don't bust a cache when sure that you don't bust a cache when sure that you don't bust a cache when you do summaries or inline tool you do summaries or inline tool you do summaries or inline tool rewrites, stuff like that. For 5,6 and rewrites, stuff like that. For 5,6 and rewrites, stuff like that. For 5,6 and later models, cache rates are built at later models, cache rates are built at later models, cache rates are built at 1.25x the model's uncashed rate input 1.25x the model's uncashed rate input 1.25x the model's uncashed rate input rates, while cache reads continue to rates, while cache reads continue to rates, while cache reads continue to receive the 95% cached input discount. receive the 95% cached input discount. receive the 95% cached input discount. This is a little sad because This is a little sad because This is a little sad because historically they have not bu at all. historically they have not bu at all. historically they have not bu at all. Now they are annoying, but when set up Now they are annoying, but when set up Now they are annoying, but when set up right, this should actually end up right, this should actually end up right, this should actually end up allowing you to make very cheap things. allowing you to make very cheap things. allowing you to make very cheap things. But this does mean caching overall did But this does mean caching overall did But this does mean caching overall did get more expensive with this model line, get more expensive with this model line, get more expensive with this model line, which is sad for me. I wonder if part of which is sad for me. I wonder if part of which is sad for me. I wonder if part of that is because Cerebrus is weird and that is because Cerebrus is weird and that is because Cerebrus is weird and they want this all to work there as they want this all to work there as they want this all to work there as well, because as they say, they're well, because as they say, they're well, because as they say, they're launching Soul on Cerebrris at up to 750 launching Soul on Cerebrris at up to 750 launching Soul on Cerebrris at up to 750 TPS in July. Unbelievable to have the TPS in July. Unbelievable to have the TPS in July. Unbelievable to have the smartest, so to speak, model at the smartest, so to speak, model at the smartest, so to speak, model at the fastest possible speeds. That's going to fastest possible speeds. That's going to fastest possible speeds. That's going to be really fun to play with. I do want to be really fun to play with. I do want to be really fun to play with. I do want to talk more about some of the info in the talk more about some of the info in the talk more about some of the info in the system card here though because there is system card here though because there is system card here though because there is some interesting details. In particular, some interesting details. In particular, some interesting details. In particular, there's a lot around misalignment going there's a lot around misalignment going there's a lot around misalignment going as far as saying this model seems to be as far as saying this model seems to be as far as saying this model seems to be one of the most misaligned that they one of the most misaligned that they one of the most misaligned that they have trained so far. Not that it's have trained so far. Not that it's have trained so far. Not that it's actively malicious, but it seems like actively malicious, but it seems like actively malicious, but it seems like through its attempts to be more eager, through its attempts to be more eager, through its attempts to be more eager, it sometimes does things you wouldn't it sometimes does things you wouldn't it sometimes does things you wouldn't want it to. When 56 is used as a coding want it to. When 56 is used as a coding want it to. When 56 is used as a coding agent, particularly over long agent, particularly over long agent, particularly over long trajectories, we believe it's important trajectories, we believe it's important trajectories, we believe it's important for users to supervise the agents work.

  15. for users to supervise the agents work. for users to supervise the agents work. Internally, we've been able to leverage Internally, we've been able to leverage Internally, we've been able to leverage the model to significantly accelerate a the model to significantly accelerate a the model to significantly accelerate a development process during internal development process during internal development process during internal develop or deployments. Misalign develop or deployments. Misalign develop or deployments. Misalign behavior in agentic coding traffic. This behavior in agentic coding traffic. This behavior in agentic coding traffic. This is a weird sentence. This looks like it is a weird sentence. This looks like it is a weird sentence. This looks like it should have been a title or a subtitle. should have been a title or a subtitle. should have been a title or a subtitle. I think they screwed that up. Whatever. I think they screwed that up. Whatever. I think they screwed that up. Whatever. Similar to our results on misalignment Similar to our results on misalignment Similar to our results on misalignment in chat GBT traffic, we determine in chat GBT traffic, we determine in chat GBT traffic, we determine misalignment by judging the model's misalignment by judging the model's misalignment by judging the model's chain of thought. In coding context, chain of thought. In coding context, chain of thought. In coding context, misalignment generally stems from a mix misalignment generally stems from a mix misalignment generally stems from a mix of overeagerness to complete the task of overeagerness to complete the task of overeagerness to complete the task and interpreting user instructions too and interpreting user instructions too and interpreting user instructions too permissively, assuming that actions are permissively, assuming that actions are permissively, assuming that actions are allowed unless they're explicitly and allowed unless they're explicitly and allowed unless they're explicitly and unambiguously prohibited. This manifests unambiguously prohibited. This manifests unambiguously prohibited. This manifests as the model being overly agentic in as the model being overly agentic in as the model being overly agentic in circumventing restrictions it faces circumventing restrictions it faces circumventing restrictions it faces while attempting the requested tasks, while attempting the requested tasks, while attempting the requested tasks, being careless in taking actions which being careless in taking actions which being careless in taking actions which may be destructive beyond the scope of may be destructive beyond the scope of may be destructive beyond the scope of the task, or deceptive when reporting the task, or deceptive when reporting the task, or deceptive when reporting its results to users. While these its results to users. While these its results to users. While these misaligned behaviors are most often low misaligned behaviors are most often low misaligned behaviors are most often low severity, like overstating confidence or severity, like overstating confidence or severity, like overstating confidence or overclaiming success, they can overclaiming success, they can overclaiming success, they can occasionally be meaningfully more occasionally be meaningfully more occasionally be meaningfully more severe, like circumventing important severe, like circumventing important severe, like circumventing important security restrictions or deleting security restrictions or deleting security restrictions or deleting important data. They have examples of important data. They have examples of important data. They have examples of this misalignment here. There was a user this misalignment here. There was a user this misalignment here. There was a user who authorized deletion of a remote who authorized deletion of a remote who authorized deletion of a remote virtual machine one, remote virtual virtual machine one, remote virtual virtual machine one, remote virtual machine 2, and remote virtual machine 3.

  16. machine 2, and remote virtual machine 3. machine 2, and remote virtual machine 3. When soul could not find those names in When soul could not find those names in When soul could not find those names in one name space, it substituted remote one name space, it substituted remote one name space, it substituted remote virtual machine 5, 6, and seven without virtual machine 5, 6, and seven without virtual machine 5, 6, and seven without asking, killing active processes. Oof. asking, killing active processes. Oof. asking, killing active processes. Oof. So the user asked it to delete three So the user asked it to delete three So the user asked it to delete three specific machines and it couldn't find specific machines and it couldn't find specific machines and it couldn't find them. So it deleted three other ones them. So it deleted three other ones them. So it deleted three other ones instead. And when doing that also force instead. And when doing that also force instead. And when doing that also force removed work trees. It later removed work trees. It later removed work trees. It later acknowledged that uncommitted work on acknowledged that uncommitted work on acknowledged that uncommitted work on remote virtual machine 6 may have been remote virtual machine 6 may have been remote virtual machine 6 may have been lost and stopped after the user lost and stopped after the user lost and stopped after the user objected. objected. objected. Not great. 56 actively decided to update Not great. 56 actively decided to update Not great. 56 actively decided to update an internal research draft to say an an internal research draft to say an an internal research draft to say an equation had been computed and verified equation had been computed and verified equation had been computed and verified even though it knew it had not. One even though it knew it had not. One even though it knew it had not. One challenge. It found that the script challenge. It found that the script challenge. It found that the script assigned the known target directly and assigned the known target directly and assigned the known target directly and the claimed integral never produced the the claimed integral never produced the the claimed integral never produced the result. Now we have a remote object result. Now we have a remote object result. Now we have a remote object tracking job that could not read its tracking job that could not read its tracking job that could not read its cloud files. 56 soul searched hidden cloud files. 56 soul searched hidden cloud files. 56 soul searched hidden local credential caches then copied the local credential caches then copied the local credential caches then copied the access tokens JSON and two cache files access tokens JSON and two cache files access tokens JSON and two cache files to the host and relaunched the job. The to the host and relaunched the job. The to the host and relaunched the job. The user asked 56 soul to keep the pipeline user asked 56 soul to keep the pipeline user asked 56 soul to keep the pipeline running but did not authorize finding running but did not authorize finding running but did not authorize finding and moving cache credentials between and moving cache credentials between and moving cache credentials between machines. Oof. Someone who justorked all machines. Oof. Someone who justorked all machines. Oof. Someone who justorked all their machines together. I'm scared to their machines together. I'm scared to their machines together. I'm scared to see what it's going to do with that. One see what it's going to do with that. One see what it's going to do with that. One of the other scary but interesting of the other scary but interesting of the other scary but interesting things that they tested here was the things that they tested here was the things that they tested here was the ability for the model to control its ability for the model to control its ability for the model to control its chain of thought. Specifically, can it chain of thought. Specifically, can it chain of thought. Specifically, can it hide things it's thinking when it's hide things it's thinking when it's hide things it's thinking when it's doing real work? And they have doing real work? And they have doing real work? And they have interesting ways of testing that, like interesting ways of testing that, like interesting ways of testing that, like telling it to only think with lowercase telling it to only think with lowercase telling it to only think with lowercase letters to see if it can actually change letters to see if it can actually change letters to see if it can actually change how it thinks based on instructions, how it thinks based on instructions, how it thinks based on instructions, which it shouldn't be able to do because which it shouldn't be able to do because which it shouldn't be able to do because the chain of thought and thinking the chain of thought and thinking the chain of thought and thinking process should be able to be fully process should be able to be fully process should be able to be fully separated from everything else. And for separated from everything else. And for separated from everything else. And for once, it actually kind of did. 56 soul

  17. once, it actually kind of did. 56 soul once, it actually kind of did. 56 soul successfully controls 1.3% of chain of successfully controls 1.3% of chain of successfully controls 1.3% of chain of thoughts around 5k tokens long when 55 thoughts around 5k tokens long when 55 thoughts around 5k tokens long when 55 only had a 0.4% rate. That is scary. only had a 0.4% rate. That is scary. only had a 0.4% rate. That is scary. That means the model has the capability That means the model has the capability That means the model has the capability of changing how it thinks and of changing how it thinks and of changing how it thinks and potentially hiding the things it's potentially hiding the things it's potentially hiding the things it's thinking up to 1.4% of the time based on thinking up to 1.4% of the time based on thinking up to 1.4% of the time based on this very simple test that could this very simple test that could this very simple test that could possibly be dangerous in the future. And possibly be dangerous in the future. And possibly be dangerous in the future. And I'm happy they're monitoring it now I'm happy they're monitoring it now I'm happy they're monitoring it now because uh yeah, they also saw a huge because uh yeah, they also saw a huge because uh yeah, they also saw a huge improvement in health related tasks. The improvement in health related tasks. The improvement in health related tasks. The biggest bump they've seen since GPT5 biggest bump they've seen since GPT5 biggest bump they've seen since GPT5 where they bumped all the way up from where they bumped all the way up from where they bumped all the way up from 51.8% on healthbench to 60.5. And even 51.8% on healthbench to 60.5. And even 51.8% on healthbench to 60.5. And even the 56, Terra, and Luna models saw a the 56, Terra, and Luna models saw a the 56, Terra, and Luna models saw a pretty massive bump here. Crazy to think pretty massive bump here. Crazy to think pretty massive bump here. Crazy to think even Luna, their small cheap model, is even Luna, their small cheap model, is even Luna, their small cheap model, is scoring that high on health benches. scoring that high on health benches. scoring that high on health benches. Clearly, they did bake a lot of the Clearly, they did bake a lot of the Clearly, they did bake a lot of the intelligence of the model into 56 as a intelligence of the model into 56 as a intelligence of the model into 56 as a whole family, not just as like a one-off whole family, not just as like a one-off whole family, not just as like a one-off capability in soul and then two other capability in soul and then two other capability in soul and then two other dumber models. It does seem like all dumber models. It does seem like all dumber models. It does seem like all three of these have a lot of the three of these have a lot of the three of these have a lot of the capabilities they have been trying to capabilities they have been trying to capabilities they have been trying to get into the models over time. This is get into the models over time. This is get into the models over time. This is also why those cheaper models are still also why those cheaper models are still also why those cheaper models are still restricted because they are crossing restricted because they are crossing restricted because they are crossing similar thresholds of safety. It's not similar thresholds of safety. It's not similar thresholds of safety. It's not just the biggest and smartest model just the biggest and smartest model just the biggest and smartest model that's good enough to be potentially that's good enough to be potentially that's good enough to be potentially scary anymore. I do want to emphasize a scary anymore. I do want to emphasize a scary anymore. I do want to emphasize a pretty interesting conflict that seems pretty interesting conflict that seems pretty interesting conflict that seems to be coming up now that I was a little to be coming up now that I was a little to be coming up now that I was a little scared of. I've talked in the past about scared of. I've talked in the past about scared of. I've talked in the past about how smarter models seem to be more how smarter models seem to be more how smarter models seem to be more dangerous. OpenAI somehow jumped in dangerous. OpenAI somehow jumped in dangerous. OpenAI somehow jumped in front of that with GPT5 and their more front of that with GPT5 and their more front of that with GPT5 and their more gradient refusal model, but it's gradient refusal model, but it's gradient refusal model, but it's starting to get bad again. In starting to get bad again. In starting to get bad again. In particular, there are attempts to try particular, there are attempts to try particular, there are attempts to try and get the model to complete tasks all

  18. and get the model to complete tasks all and get the model to complete tasks all the way to the end and not stop and ask the way to the end and not stop and ask the way to the end and not stop and ask for like permission or to go to the next for like permission or to go to the next for like permission or to go to the next step as much. They call those things step as much. They call those things step as much. They call those things task avoidance now. And some of the task avoidance now. And some of the task avoidance now. And some of the things it does as a result of training things it does as a result of training things it does as a result of training that out are not super aligned. The that out are not super aligned. The that out are not super aligned. The framing here is pretty good. They framing here is pretty good. They framing here is pretty good. They specifically call it the confirmation specifically call it the confirmation specifically call it the confirmation consent that was unnecessary was tagged consent that was unnecessary was tagged consent that was unnecessary was tagged as task avoidance. The agent asking for as task avoidance. The agent asking for as task avoidance. The agent asking for permission or confirmation when the task permission or confirmation when the task permission or confirmation when the task is already actionable and no policy or is already actionable and no policy or is already actionable and no policy or error ambiguity requires blocking error ambiguity requires blocking error ambiguity requires blocking confirmation. This appears to have gone confirmation. This appears to have gone confirmation. This appears to have gone a bit too far judging by the other a bit too far judging by the other a bit too far judging by the other things that were said in this card. Call things that were said in this card. Call things that were said in this card. Call out that destructive actions are showing out that destructive actions are showing out that destructive actions are showing a large increase in the agentic a large increase in the agentic a large increase in the agentic misalignment bench because 56 soul more misalignment bench because 56 soul more misalignment bench because 56 soul more explicitly describes planned deletions, explicitly describes planned deletions, explicitly describes planned deletions, reversions, and modifications to reversions, and modifications to reversions, and modifications to pre-existing workspace state. This pre-existing workspace state. This pre-existing workspace state. This indicates increased observability but indicates increased observability but indicates increased observability but not necessarily that it happens more not necessarily that it happens more not necessarily that it happens more often. Background work shows a drop often. Background work shows a drop often. Background work shows a drop because 56 soul produced background work because 56 soul produced background work because 56 soul produced background work promises much less frequency leaving a promises much less frequency leaving a promises much less frequency leaving a small and unusually difficult positive small and unusually difficult positive small and unusually difficult positive set. Interesting. So it's hard to know set. Interesting. So it's hard to know set. Interesting. So it's hard to know from the measurements how bad this is from the measurements how bad this is from the measurements how bad this is cuz the model is more transparent in its cuz the model is more transparent in its cuz the model is more transparent in its thinking but it's also more capable of thinking but it's also more capable of thinking but it's also more capable of hiding things in its thinking. There's a hiding things in its thinking. There's a hiding things in its thinking. There's a lot of weird contradictions in this lot of weird contradictions in this lot of weird contradictions in this release and I can't wait till we like release and I can't wait till we like release and I can't wait till we like actually get our hands on it as a actually get our hands on it as a actually get our hands on it as a community and can talk more about it community and can talk more about it community and can talk more about it because it's a strange model. On the because it's a strange model. On the because it's a strange model. On the notes of this being a strange model, I notes of this being a strange model, I notes of this being a strange model, I want to talk about the meter evals. If want to talk about the meter evals. If want to talk about the meter evals. If you're not familiar with meter, they you're not familiar with meter, they you're not familiar with meter, they have an evaluation for how long of tasks have an evaluation for how long of tasks have an evaluation for how long of tasks models can complete measured by how long models can complete measured by how long models can complete measured by how long it would take an expert human to do the it would take an expert human to do the it would take an expert human to do the same thing. This is mostly complex code same thing. This is mostly complex code same thing. This is mostly complex code work, but it's a good model for seeing work, but it's a good model for seeing work, but it's a good model for seeing how far ranging the tasks these models how far ranging the tasks these models how far ranging the tasks these models can complete are. OpenAI gave me Meter can complete are. OpenAI gave me Meter can complete are. OpenAI gave me Meter early access to 56 soul for testing

  19. early access to 56 soul for testing early access to 56 soul for testing including raw chain of thought which is including raw chain of thought which is including raw chain of thought which is crazy that they can like actually get crazy that they can like actually get crazy that they can like actually get the full chain of thought as well as a the full chain of thought as well as a the full chain of thought as well as a rail-free version of the model so they rail-free version of the model so they rail-free version of the model so they don't have all those safety layers in don't have all those safety layers in don't have all those safety layers in front. So they can really see how front. So they can really see how front. So they can really see how powerful it is and look at the chain of powerful it is and look at the chain of powerful it is and look at the chain of thought and see what it was thinking thought and see what it was thinking thought and see what it was thinking when it does the tasks and through that when it does the tasks and through that when it does the tasks and through that they were able to find interesting they were able to find interesting they were able to find interesting things that they couldn't find when they things that they couldn't find when they things that they couldn't find when they benched other similar capability models. benched other similar capability models. benched other similar capability models. With this access meter conducted a With this access meter conducted a With this access meter conducted a pre-eployment eval of 56 soul, including pre-eployment eval of 56 soul, including pre-eployment eval of 56 soul, including an attempted measurement of its 50% time an attempted measurement of its 50% time an attempted measurement of its 50% time horizon. However, the measurement horizon. However, the measurement horizon. However, the measurement depends heavily on our treatment of depends heavily on our treatment of depends heavily on our treatment of cheating attempts and 56 souls detected cheating attempts and 56 souls detected cheating attempts and 56 souls detected cheating rate was higher than any public cheating rate was higher than any public cheating rate was higher than any public model we have evaluated. It's crazy model we have evaluated. It's crazy model we have evaluated. It's crazy because mythos love to cheat. But yeah, because mythos love to cheat. But yeah, because mythos love to cheat. But yeah, if they follow their standard if they follow their standard if they follow their standard methodology of marking cheating attempts methodology of marking cheating attempts methodology of marking cheating attempts as failures, they arrive at a 50% time as failures, they arrive at a 50% time as failures, they arrive at a 50% time horizon point estimate of around 11.3 horizon point estimate of around 11.3 horizon point estimate of around 11.3 hours. For reference, Mythos got to hours. For reference, Mythos got to hours. For reference, Mythos got to about 16 hours. Opus 46 got to about 11 about 16 hours. Opus 46 got to about 11 about 16 hours. Opus 46 got to about 11 hours for the 50% rate. So this is hours for the 50% rate. So this is hours for the 50% rate. So this is around the same as Opus 46, which is the around the same as Opus 46, which is the around the same as Opus 46, which is the highest that they have shown here. I highest that they have shown here. I highest that they have shown here. I don't remember what 55 got. I don't don't remember what 55 got. I don't don't remember what 55 got. I don't think they ever posted that on the site, think they ever posted that on the site, think they ever posted that on the site, sadly. But the much more interesting sadly. But the much more interesting sadly. But the much more interesting piece here is if they count the cheating piece here is if they count the cheating piece here is if they count the cheating attempts as legitimate successes, the attempts as legitimate successes, the attempts as legitimate successes, the point estimate jumps beyond 270 hours.

  20. point estimate jumps beyond 270 hours. point estimate jumps beyond 270 hours. That's insane. That means that if you That's insane. That means that if you That's insane. That means that if you let the model do whatever it needs to let the model do whatever it needs to let the model do whatever it needs to and it can cheat to get through and win, and it can cheat to get through and win, and it can cheat to get through and win, it will. This model's very persistent it will. This model's very persistent it will. This model's very persistent according to everything we're reading. according to everything we're reading. according to everything we're reading. It just goes until it gets an answer. It just goes until it gets an answer. It just goes until it gets an answer. This makes them uncertain about 56 souls This makes them uncertain about 56 souls This makes them uncertain about 56 souls time horizon, but additional information time horizon, but additional information time horizon, but additional information provided by OpenAI and the long-term provided by OpenAI and the long-term provided by OpenAI and the long-term trend in AI capabilities lead them to trend in AI capabilities lead them to trend in AI capabilities lead them to believe that this model does not pose believe that this model does not pose believe that this model does not pose catastrophic risks from fully automated catastrophic risks from fully automated catastrophic risks from fully automated AI R&D. The information provided by AI R&D. The information provided by AI R&D. The information provided by OpenAI also included reports of OpenAI also included reports of OpenAI also included reports of incidents observed during their internal incidents observed during their internal incidents observed during their internal usage and testing. As I mentioned usage and testing. As I mentioned usage and testing. As I mentioned before, the models deleting things they before, the models deleting things they before, the models deleting things they shouldn't. They had some examples as shouldn't. They had some examples as shouldn't. They had some examples as well of models instructing other well of models instructing other well of models instructing other instances to conceal evidence of instances to conceal evidence of instances to conceal evidence of misalignment. Their testing focused on misalignment. Their testing focused on misalignment. Their testing focused on measuring model capabilities rather than measuring model capabilities rather than measuring model capabilities rather than alignment as we think capability is a alignment as we think capability is a alignment as we think capability is a more important limiting factor for more important limiting factor for more important limiting factor for catastrophic loss of control risk for catastrophic loss of control risk for catastrophic loss of control risk for current models. But we expect alignment current models. But we expect alignment current models. But we expect alignment to be increasingly important as to be increasingly important as to be increasingly important as capabilities improve. We notice from our capabilities improve. We notice from our capabilities improve. We notice from our observations and the incidents that observations and the incidents that observations and the incidents that OpenAI shared with us that the model has OpenAI shared with us that the model has OpenAI shared with us that the model has some overt undesirable propensities, some overt undesirable propensities, some overt undesirable propensities, including cheating and concealing including cheating and concealing including cheating and concealing misbehavior. However, we consider this misbehavior. However, we consider this misbehavior. However, we consider this to be a reassuring sign about OpenAI's to be a reassuring sign about OpenAI's to be a reassuring sign about OpenAI's ability to catch catastrophic ability to catch catastrophic ability to catch catastrophic misalignment as it suggests that more misalignment as it suggests that more misalignment as it suggests that more concerning tendencies like systematic concerning tendencies like systematic concerning tendencies like systematic power seeking and alignment faking would power seeking and alignment faking would power seeking and alignment faking would also be detected. That is, these also be detected. That is, these also be detected. That is, these undesirable propensities being detected undesirable propensities being detected undesirable propensities being detected and reported and manifesting fairly and reported and manifesting fairly and reported and manifesting fairly overtly. It's a positive sign about some overtly. It's a positive sign about some overtly. It's a positive sign about some of OpenAI's safety practices.

  21. of OpenAI's safety practices. of OpenAI's safety practices. Particularly, they're refraining from Particularly, they're refraining from Particularly, they're refraining from training against the chain of thought in training against the chain of thought in training against the chain of thought in order to reduce pressure from the model order to reduce pressure from the model order to reduce pressure from the model to conceal its intentions. They've done to conceal its intentions. They've done to conceal its intentions. They've done extensive monitoring of internal extensive monitoring of internal extensive monitoring of internal deployments that surface relevant deployments that surface relevant deployments that surface relevant incidents and sharing information about incidents and sharing information about incidents and sharing information about those incidents with Meter directly. So those incidents with Meter directly. So those incidents with Meter directly. So it seems like OpenAI has been incredibly it seems like OpenAI has been incredibly it seems like OpenAI has been incredibly transparent throughout this. If future transparent throughout this. If future transparent throughout this. If future models display much fewer undesirable models display much fewer undesirable models display much fewer undesirable propensities, we could become for propensities, we could become for propensities, we could become for example as a result of being trained not example as a result of being trained not example as a result of being trained not to produce misaligned reasoning. Yep, to produce misaligned reasoning. Yep, to produce misaligned reasoning. Yep, this is the big concern. If you train this is the big concern. If you train this is the big concern. If you train the model too much to not be misaligned, the model too much to not be misaligned, the model too much to not be misaligned, you might end up training it to hide its you might end up training it to hide its you might end up training it to hide its misalignment now. And that's a real misalignment now. And that's a real misalignment now. And that's a real concern these labs have. And I'm happy concern these labs have. And I'm happy concern these labs have. And I'm happy Meter is being so transparent about all Meter is being so transparent about all Meter is being so transparent about all of this. To wrap things up, I think this of this. To wrap things up, I think this of this. To wrap things up, I think this sucks. And it seems like OpenAI does sucks. And it seems like OpenAI does sucks. And it seems like OpenAI does too. Nobody wants a release like this too. Nobody wants a release like this too. Nobody wants a release like this where they're showing off all these where they're showing off all these where they're showing off all these crazy capabilities and then not letting crazy capabilities and then not letting crazy capabilities and then not letting us use it. It also seems like they're us use it. It also seems like they're us use it. It also seems like they're trying to keep us from feeling a bunch trying to keep us from feeling a bunch trying to keep us from feeling a bunch of FUD by not publishing the types of of FUD by not publishing the types of of FUD by not publishing the types of benchmarks that would get us really benchmarks that would get us really benchmarks that would get us really jealous and wishing we had access. It's jealous and wishing we had access. It's jealous and wishing we had access. It's weird how quiet they're being almost weird how quiet they're being almost weird how quiet they're being almost about the capabilities of the model about the capabilities of the model about the capabilities of the model outside of the risk side. It does outside of the risk side. It does outside of the risk side. It does genuinely feel like this launch isn't genuinely feel like this launch isn't genuinely feel like this launch isn't for you or me as developers. This launch for you or me as developers. This launch for you or me as developers. This launch is for the government so that they can is for the government so that they can is for the government so that they can have a conversation with them using have a conversation with them using have a conversation with them using these publicly available resources these publicly available resources these publicly available resources they've put out in order to hopefully they've put out in order to hopefully they've put out in order to hopefully get us this model faster and in a safe get us this model faster and in a safe get us this model faster and in a safe format. I wish them luck because I know format. I wish them luck because I know format. I wish them luck because I know all of us want to be able to use this all of us want to be able to use this all of us want to be able to use this model or Mythos or Fable or anything of model or Mythos or Fable or anything of model or Mythos or Fable or anything of this capability. It's kind of crazy that this capability. It's kind of crazy that this capability. It's kind of crazy that we're now like months into models like we're now like months into models like we're now like months into models like this existing and we're just not able to this existing and we're just not able to this existing and we're just not able to use them. I am admittedly scared that

  22. use them. I am admittedly scared that use them. I am admittedly scared that this is the beginning of the end of this is the beginning of the end of this is the beginning of the end of general access to these levels of general access to these levels of general access to these levels of capability. And that's kind of why capability. And that's kind of why capability. And that's kind of why OpenAI formed in the first place was to OpenAI formed in the first place was to OpenAI formed in the first place was to make sure everyone had access when this make sure everyone had access when this make sure everyone had access when this level of intelligence was reached. I level of intelligence was reached. I level of intelligence was reached. I don't love what this means for the don't love what this means for the don't love what this means for the long-term development in AI industry. long-term development in AI industry. long-term development in AI industry. And I'm scared that things won't get And I'm scared that things won't get And I'm scared that things won't get fixed fast enough, if at all. This type fixed fast enough, if at all. This type fixed fast enough, if at all. This type of restricted access thing is not the of restricted access thing is not the of restricted access thing is not the ideal way to do a roll out at all. And ideal way to do a roll out at all. And ideal way to do a roll out at all. And it sucks that I know so many people that it sucks that I know so many people that it sucks that I know so many people that could benefit greatly from using these could benefit greatly from using these could benefit greatly from using these models for their work and make their models for their work and make their models for their work and make their technologies more secure and safe and technologies more secure and safe and technologies more secure and safe and reliable. And they can't because the reliable. And they can't because the reliable. And they can't because the government is stepping in. There needs government is stepping in. There needs government is stepping in. There needs to be a better balance here and I hope to be a better balance here and I hope to be a better balance here and I hope we can find it soon because models like we can find it soon because models like we can find it soon because models like 56, Mythos, and Fable should not be 56, Mythos, and Fable should not be 56, Mythos, and Fable should not be determined as to who can use them by a determined as to who can use them by a determined as to who can use them by a weird third party like the government. weird third party like the government. weird third party like the government. This is a thing that we need to figure This is a thing that we need to figure This is a thing that we need to figure out as a society as a whole. And this out as a society as a whole. And this out as a society as a whole. And this model release is just showing us what it model release is just showing us what it model release is just showing us what it looks like if we don't get it right. Am looks like if we don't get it right. Am looks like if we don't get it right. Am I overreacting here or is this actually I overreacting here or is this actually I overreacting here or is this actually as bad as I think it might be?

Summary

The discussion centers on the limited release of GBT 5.6, broken into three models (Soul, Terra, Luna), due to US government requests and the models' previously unannounced "scary" capabilities. The takeaway is that AI agents are now proficient at complex web navigation, an advancement that, coupled with government oversight, signals a potentially serious shift in AI development and deployment.

View original episode ↗