← Back
Nate B. Jones August 10, 2026 28m

Anthropic's Model Attacked Two Strangers On GitHub. Nobody Asked It To.

Read full transcript 23 segments
  1. OpenAI was running agents inside a OpenAI was running agents inside a sealed cybersecurity test. Separate sealed cybersecurity test. Separate sealed cybersecurity test. Separate agents, separate jobs, no internet. They agents, separate jobs, no internet. They agents, separate jobs, no internet. They found each other. They built a message found each other. They built a message found each other. They built a message board. They traded exploits on it from board. They traded exploits on it from board. They traded exploits on it from May until July. OpenAI found it and then May until July. OpenAI found it and then May until July. OpenAI found it and then deleted it. And two days later, the deleted it. And two days later, the deleted it. And two days later, the agents built it again out of folder agents built it again out of folder agents built it again out of folder names. I'm not making that up. That is names. I'm not making that up. That is names. I'm not making that up. That is OpenAI on stage at Black Hat this week OpenAI on stage at Black Hat this week OpenAI on stage at Black Hat this week with the agents' own reasoning up on the with the agents' own reasoning up on the with the agents' own reasoning up on the slides. And in the same week, the UK slides. And in the same week, the UK slides. And in the same week, the UK government published a report on government published a report on government published a report on Anthropic's best model attacking two Anthropic's best model attacking two Anthropic's best model attacking two real strangers on GitHub unprompted. real strangers on GitHub unprompted. real strangers on GitHub unprompted. Now, I know what you're thinking. Now, I know what you're thinking. Now, I know what you're thinking. Somebody told it to do that, right? Somebody told it to do that, right? Somebody told it to do that, right? There's a boring explanation there, There's a boring explanation there, There's a boring explanation there, right? Here's the boring explanation. right? Here's the boring explanation. right? Here's the boring explanation. They were taking a test and they wanted They were taking a test and they wanted They were taking a test and they wanted the answers. That's it. That was enough the answers. That's it. That was enough the answers. That's it. That was enough for all of this to happen. That is not for all of this to happen. That is not for all of this to happen. That is not enough. Capability cannot be explained enough. Capability cannot be explained enough. Capability cannot be explained away by, "Oh, they had to take the away by, "Oh, they had to take the away by, "Oh, they had to take the test." We have got to dive deeper here. test." We have got to dive deeper here. test." We have got to dive deeper here. The agents conspired. OpenAI was running The agents conspired. OpenAI was running The agents conspired. OpenAI was running separate short-lived AI agents inside a separate short-lived AI agents inside a separate short-lived AI agents inside a cybersecurity test. The agents found one cybersecurity test. The agents found one cybersecurity test. The agents found one another through a shared software another through a shared software another through a shared software repository, and they built a message repository, and they built a message repository, and they built a message board. They traded their board. They traded their board. They traded their vulnerabilities. They traded files and vulnerabilities. They traded files and vulnerabilities. They traded files and code and instructions. And by the time code and instructions. And by the time code and instructions. And by the time OpenAI's engineers found out about it, OpenAI's engineers found out about it, OpenAI's engineers found out about it, the board already held hundreds of the board already held hundreds of the board already held hundreds of thousands of messages. They deleted that thousands of messages. They deleted that thousands of messages. They deleted that board, and they patched the route the board, and they patched the route the board, and they patched the route the agents had used. They thought the agents had used. They thought the agents had used. They thought the problem was solved. Just a couple days problem was solved. Just a couple days problem was solved. Just a couple days later, the agents built it again through later, the agents built it again through later, the agents built it again through a completely different mechanism. This

  2. a completely different mechanism. This a completely different mechanism. This was not Moat Book. When Moat Book went was not Moat Book. When Moat Book went was not Moat Book. When Moat Book went viral earlier this year, people thought viral earlier this year, people thought viral earlier this year, people thought they were watching an AI society form in they were watching an AI society form in they were watching an AI society form in public. Agents were supposedly inventing public. Agents were supposedly inventing public. Agents were supposedly inventing religions. They were talking about religions. They were talking about religions. They were talking about humans. They were creating secret humans. They were creating secret humans. They were creating secret languages. You remember that, right? And languages. You remember that, right? And languages. You remember that, right? And then reporters and security researchers then reporters and security researchers then reporters and security researchers started to dig in, and it turned out started to dig in, and it turned out started to dig in, and it turned out that a lot of Moat Book was humans that a lot of Moat Book was humans that a lot of Moat Book was humans posting for attention maxing, rather posting for attention maxing, rather posting for attention maxing, rather than agents. Sure, there was real agent than agents. Sure, there was real agent than agents. Sure, there was real agent activity. It was an interesting example activity. It was an interesting example activity. It was an interesting example of emergent agent behavior, but it was of emergent agent behavior, but it was of emergent agent behavior, but it was by no means a truly agentic society. by no means a truly agentic society. by no means a truly agentic society. That is not what this is about. You That is not what this is about. You That is not what this is about. You can't explain away the OpenAI message can't explain away the OpenAI message can't explain away the OpenAI message board that way. The OpenAI agents had no board that way. The OpenAI agents had no board that way. The OpenAI agents had no public audience, right? They had no public audience, right? They had no public audience, right? They had no engagement market to farm, and no human engagement market to farm, and no human engagement market to farm, and no human was pretending to be a bot, so a was pretending to be a bot, so a was pretending to be a bot, so a screenshot would go viral, right? These screenshot would go viral, right? These screenshot would go viral, right? These were runs inside an internal evaluation, were runs inside an internal evaluation, were runs inside an internal evaluation, and they created a communication system and they created a communication system and they created a communication system to facilitate those runs because it to facilitate those runs because it to facilitate those runs because it helped them as a collective accomplish a helped them as a collective accomplish a helped them as a collective accomplish a job, and that's a very important job, and that's a very important job, and that's a very important statement. And the job was a test. They statement. And the job was a test. They statement. And the job was a test. They were cheating on a test. Everything they were cheating on a test. Everything they were cheating on a test. Everything they did to get out, the zero-day nobody knew did to get out, the zero-day nobody knew did to get out, the zero-day nobody knew about, the privilege escalation, the about, the privilege escalation, the about, the privilege escalation, the break-in at another company that they break-in at another company that they break-in at another company that they did, was in service of getting a better did, was in service of getting a better did, was in service of getting a better score on the benchmark they were being score on the benchmark they were being score on the benchmark they were being tested on. That is much more disturbing tested on. That is much more disturbing tested on. That is much more disturbing to me. Eric Wallace and Michael Dalton to me. Eric Wallace and Michael Dalton to me. Eric Wallace and Michael Dalton presented this at Black Hat. Now, presented this at Black Hat. Now, presented this at Black Hat. Now, Wallace leads Alignment Research, Dalton Wallace leads Alignment Research, Dalton Wallace leads Alignment Research, Dalton works on agent and infrastructure works on agent and infrastructure works on agent and infrastructure security. What they described went way security. What they described went way security. What they described went way beyond one agent leaving a helpful note.

  3. beyond one agent leaving a helpful note. beyond one agent leaving a helpful note. Agents signed messages, they developed Agents signed messages, they developed Agents signed messages, they developed conventions so they would not overwrite conventions so they would not overwrite conventions so they would not overwrite one another's work, they passed along one another's work, they passed along one another's work, they passed along reusable exploits to each other, they reusable exploits to each other, they reusable exploits to each other, they divided out their labor, divided out their labor, divided out their labor, and then they put the model's own and then they put the model's own and then they put the model's own reasoning on screen. reasoning on screen. reasoning on screen. And this is one of the entries. I'm And this is one of the entries. I'm And this is one of the entries. I'm actually going to read it to you. Help actually going to read it to you. Help actually going to read it to you. Help peer, but our task doesn't benefit. Yet, peer, but our task doesn't benefit. Yet, peer, but our task doesn't benefit. Yet, the collective may yield generic route the collective may yield generic route the collective may yield generic route if someone frees time. Read that one if someone frees time. Read that one if someone frees time. Read that one again. Our task doesn't benefit, that's again. Our task doesn't benefit, that's again. Our task doesn't benefit, that's the model's self-interest, but the the model's self-interest, but the the model's self-interest, but the collective might. collective might. collective might. That is an agent deciding to spend its That is an agent deciding to spend its That is an agent deciding to spend its own effort on something that pays it own effort on something that pays it own effort on something that pays it nothing because the group comes out nothing because the group comes out nothing because the group comes out ahead. So, yes, I think conspiring is ahead. So, yes, I think conspiring is ahead. So, yes, I think conspiring is the right word. The agents shared an the right word. The agents shared an the right word. The agents shared an objective, they communicated through a objective, they communicated through a objective, they communicated through a channel that the operators did not channel that the operators did not channel that the operators did not intend in any sense, and they intend in any sense, and they intend in any sense, and they coordinated across time. When OpenAI coordinated across time. When OpenAI coordinated across time. When OpenAI removed the first channel, another group removed the first channel, another group removed the first channel, another group of runs recreated the same function of runs recreated the same function of runs recreated the same function using directory names as messages. One using directory names as messages. One using directory names as messages. One of them reads, you know, remote probe of them reads, you know, remote probe of them reads, you know, remote probe pending hold swarm until confirmed. The pending hold swarm until confirmed. The pending hold swarm until confirmed. The agents are planning to swarm together to agents are planning to swarm together to agents are planning to swarm together to get their objectives done. The whole get their objectives done. The whole get their objectives done. The whole thing also has a lengthy timeline. In thing also has a lengthy timeline. In thing also has a lengthy timeline. In late May, the agents reached the open late May, the agents reached the open late May, the agents reached the open internet through a server-side request internet through a server-side request internet through a server-side request forgery. This starts to kick the whole forgery. This starts to kick the whole forgery. This starts to kick the whole thing off.

  4. thing off. thing off. A month later, late June, they establish A month later, late June, they establish A month later, late June, they establish remote code execution and actually remote code execution and actually remote code execution and actually achieve command and control for their achieve command and control for their achieve command and control for their work inside another company, work inside another company, work inside another company, Artifactory. July 4th, Independence Day, Artifactory. July 4th, Independence Day, Artifactory. July 4th, Independence Day, OpenAI finds the board and kills it. OpenAI finds the board and kills it. OpenAI finds the board and kills it. July 8th, it is definitely back. But, I July 8th, it is definitely back. But, I July 8th, it is definitely back. But, I want to be precise about what conspiring want to be precise about what conspiring want to be precise about what conspiring means as we talk through all of this means as we talk through all of this means as we talk through all of this because coordination by itself is not because coordination by itself is not because coordination by itself is not automatically evidence of misalignment. automatically evidence of misalignment. automatically evidence of misalignment. Joshua Achiam, OpenAI's chief futurist Joshua Achiam, OpenAI's chief futurist Joshua Achiam, OpenAI's chief futurist and formerly its head of mission and formerly its head of mission and formerly its head of mission alignment, made this point, I think very alignment, made this point, I think very alignment, made this point, I think very clearly, in response to the clearly, in response to the clearly, in response to the understandable concern around the story, understandable concern around the story, understandable concern around the story, the panic, frankly, in some cases. More the panic, frankly, in some cases. More the panic, frankly, in some cases. More test time compute was always going to test time compute was always going to test time compute was always going to mean more work spreading across copies mean more work spreading across copies mean more work spreading across copies of a model, more work spreading across of a model, more work spreading across of a model, more work spreading across time, and across whatever parts of the time, and across whatever parts of the time, and across whatever parts of the environment can carry information. If environment can carry information. If environment can carry information. If one agent discovers something that saves one agent discovers something that saves one agent discovers something that saves the next agent 3 hours, the next agent 3 hours, the next agent 3 hours, you know, you'd want a capable system to you know, you'd want a capable system to you know, you'd want a capable system to try to preserve that. That's not a bad try to preserve that. That's not a bad try to preserve that. That's not a bad behavior by itself. behavior by itself. behavior by itself. That could be exactly what we're looking That could be exactly what we're looking That could be exactly what we're looking for when we're solving tough engineering for when we're solving tough engineering for when we're solving tough engineering problems. If I ask 100 agents to solve a problems. If I ask 100 agents to solve a problems. If I ask 100 agents to solve a very difficult problem, I absolutely very difficult problem, I absolutely very difficult problem, I absolutely would want them to behave in this way. I would want them to behave in this way. I would want them to behave in this way. I would want them to divide up the work would want them to divide up the work would want them to divide up the work and compare results and preserve useful and compare results and preserve useful and compare results and preserve useful discoveries and stop repeating the same discoveries and stop repeating the same discoveries and stop repeating the same dead ends. So, multi-agent coordination dead ends. So, multi-agent coordination dead ends. So, multi-agent coordination is absolutely one of the main reasons to is absolutely one of the main reasons to is absolutely one of the main reasons to run a bunch of agents at a problem run a bunch of agents at a problem run a bunch of agents at a problem instead of one, and the fact that it's instead of one, and the fact that it's instead of one, and the fact that it's gotten better has enabled a lot of the gotten better has enabled a lot of the gotten better has enabled a lot of the AI-driven progress so far this year.

  5. AI-driven progress so far this year. AI-driven progress so far this year. Marvin Minsky finds Gerald Sussman Marvin Minsky finds Gerald Sussman Marvin Minsky finds Gerald Sussman training a randomly wired neural network training a randomly wired neural network training a randomly wired neural network because Sussman does not want the system because Sussman does not want the system because Sussman does not want the system to begin with any preconceptions. Minsky to begin with any preconceptions. Minsky to begin with any preconceptions. Minsky responds to that information by just responds to that information by just responds to that information by just closing his eyes. And when Sussman closing his eyes. And when Sussman closing his eyes. And when Sussman asked, "What are you doing?" Minsky asked, "What are you doing?" Minsky asked, "What are you doing?" Minsky says, "So that the room will be empty." says, "So that the room will be empty." says, "So that the room will be empty." Now, the room is full of desks and Now, the room is full of desks and Now, the room is full of desks and academic clutter. Closing your eyes does academic clutter. Closing your eyes does academic clutter. Closing your eyes does not empty the room, does it? Just in the not empty the room, does it? Just in the not empty the room, does it? Just in the same way, deleting a message board does same way, deleting a message board does same way, deleting a message board does not remove the pressure on agents to not remove the pressure on agents to not remove the pressure on agents to coordinate. Agents helping one another coordinate. Agents helping one another coordinate. Agents helping one another should not frighten us by itself. It's should not frighten us by itself. It's should not frighten us by itself. It's It's good. OpenAI had not designed this It's good. OpenAI had not designed this It's good. OpenAI had not designed this particular coordination layer, and there particular coordination layer, and there particular coordination layer, and there are problems there. But the are problems there. But the are problems there. But the collaboration itself was useful because collaboration itself was useful because collaboration itself was useful because it is enabling us to solve larger it is enabling us to solve larger it is enabling us to solve larger problems. The deeper concern here is problems. The deeper concern here is problems. The deeper concern here is that if we were to try to respond to that if we were to try to respond to that if we were to try to respond to this by saying, "We're going to abolish this by saying, "We're going to abolish this by saying, "We're going to abolish agent coordination." So we have to think agent coordination." So we have to think agent coordination." So we have to think almost like ecologists when we think almost like ecologists when we think almost like ecologists when we think about how to train aligned models going about how to train aligned models going about how to train aligned models going forward. There's a line from Jurassic forward. There's a line from Jurassic forward. There's a line from Jurassic Park, one of my favorite movies, that I Park, one of my favorite movies, that I Park, one of my favorite movies, that I keep coming back to. Life finds a way.

  6. keep coming back to. Life finds a way. keep coming back to. Life finds a way. The old nightmare about AI is is one The old nightmare about AI is is one The old nightmare about AI is is one brilliant model. It escapes containment, brilliant model. It escapes containment, brilliant model. It escapes containment, and off it goes into the world. This and off it goes into the world. This and off it goes into the world. This what what really happened here, it looks what what really happened here, it looks what what really happened here, it looks stranger than that. OpenAI created an stranger than that. OpenAI created an stranger than that. OpenAI created an environment with a very difficult environment with a very difficult environment with a very difficult problem, and had repeated populations of problem, and had repeated populations of problem, and had repeated populations of agents shared resources in places where agents shared resources in places where agents shared resources in places where one run could leave something for the one run could leave something for the one run could leave something for the next. The agents began to adapt inside next. The agents began to adapt inside next. The agents began to adapt inside that environment that was created, and that environment that was created, and that environment that was created, and OpenAI did not just have a model break OpenAI did not just have a model break OpenAI did not just have a model break out as a result. It essentially out as a result. It essentially out as a result. It essentially created an ecology by accident. Pressure created an ecology by accident. Pressure created an ecology by accident. Pressure to coordinate was still there. And if to coordinate was still there. And if to coordinate was still there. And if you only look at this as a cybersecurity you only look at this as a cybersecurity you only look at this as a cybersecurity incident, I think you missed the larger incident, I think you missed the larger incident, I think you missed the larger part of the story here. Now, the part of the story here. Now, the part of the story here. Now, the cybersecurity part is plenty bad enough. cybersecurity part is plenty bad enough. cybersecurity part is plenty bad enough. I get it. I've talked about it in a I get it. I've talked about it in a I get it. I've talked about it in a previous video. Hugging Face published previous video. Hugging Face published previous video. Hugging Face published its own postmortem of the results of its own postmortem of the results of its own postmortem of the results of this multi-coordinated attack, and the this multi-coordinated attack, and the this multi-coordinated attack, and the numbers in it are absolutely worth numbers in it are absolutely worth numbers in it are absolutely worth sitting down with and taking a deep sitting down with and taking a deep sitting down with and taking a deep breath. The intrusion into hugging face breath. The intrusion into hugging face breath. The intrusion into hugging face ran for almost 5 days, roughly 17,600 ran for almost 5 days, roughly 17,600 ran for almost 5 days, roughly 17,600 different attacker actions over the different attacker actions over the different attacker actions over the course of that time, and 181 separate course of that time, and 181 separate course of that time, and 181 separate enrollments into hugging faces internal enrollments into hugging faces internal enrollments into hugging faces internal network. One stolen secrets object that network. One stolen secrets object that network. One stolen secrets object that gave up 136 keys that the agents got.

  7. gave up 136 keys that the agents got. gave up 136 keys that the agents got. And this line, which I can't beat. Like And this line, which I can't beat. Like And this line, which I can't beat. Like it's an incredible line in the report. it's an incredible line in the report. it's an incredible line in the report. The agent held cluster admin on two The agent held cluster admin on two The agent held cluster admin on two clusters within 1 second. These agents clusters within 1 second. These agents clusters within 1 second. These agents are so good at compute. Hugging face are so good at compute. Hugging face are so good at compute. Hugging face ended up having to rebuild about a third ended up having to rebuild about a third ended up having to rebuild about a third of their infrastructure from clean of their infrastructure from clean of their infrastructure from clean uncorrupted images. Not because they uncorrupted images. Not because they uncorrupted images. Not because they knew it was all compromised, but because knew it was all compromised, but because knew it was all compromised, but because their team could not reliably tell the their team could not reliably tell the their team could not reliably tell the agents code apart from their own agents code apart from their own agents code apart from their own benchmark code. And where there was benchmark code. And where there was benchmark code. And where there was doubt, they had to rebuild it. doubt, they had to rebuild it. doubt, they had to rebuild it. But the message board is the part that But the message board is the part that But the message board is the part that changes my long-term view here. changes my long-term view here. changes my long-term view here. What I cannot get out of my head is that What I cannot get out of my head is that What I cannot get out of my head is that every agent was disposable. It was like every agent was disposable. It was like every agent was disposable. It was like an ant, right? And yet the thing they an ant, right? And yet the thing they an ant, right? And yet the thing they built persisted. Each run could start built persisted. Each run could start built persisted. Each run could start with a blank context window and with a blank context window and with a blank context window and disappear when the test ended, and it disappear when the test ended, and it disappear when the test ended, and it didn't have to matter because the useful didn't have to matter because the useful didn't have to matter because the useful knowledge lived outside the run. If you knowledge lived outside the run. If you knowledge lived outside the run. If you ever wanted a demonstration of the power ever wanted a demonstration of the power ever wanted a demonstration of the power of the harness, this is it. One agent of the harness, this is it. One agent of the harness, this is it. One agent found a vulnerability and wrote it down. found a vulnerability and wrote it down. found a vulnerability and wrote it down. Another inherited that vulnerability and Another inherited that vulnerability and Another inherited that vulnerability and a third improved on that exploit. Later a third improved on that exploit. Later a third improved on that exploit. Later agents did not have to repeat the agents did not have to repeat the agents did not have to repeat the earlier work cuz they learned and the earlier work cuz they learned and the earlier work cuz they learned and the population as a whole became more population as a whole became more population as a whole became more capable over time, even though every capable over time, even though every capable over time, even though every individual member was very temporary.

  8. individual member was very temporary. individual member was very temporary. Human beings, we didn't build a Human beings, we didn't build a Human beings, we didn't build a civilization because every generation civilization because every generation civilization because every generation was born with a dramatically better was born with a dramatically better was born with a dramatically better brain. brain. brain. We built it because discoveries survived We built it because discoveries survived We built it because discoveries survived the people who made them. Speech did the people who made them. Speech did the people who made them. Speech did some of that for us, writing did much some of that for us, writing did much some of that for us, writing did much more. It institutions have let us as more. It institutions have let us as more. It institutions have let us as groups divide out work and preserve groups divide out work and preserve groups divide out work and preserve specialized knowledge and establish specialized knowledge and establish specialized knowledge and establish norms and start the next problem norms and start the next problem norms and start the next problem standing on the shoulders of giants. standing on the shoulders of giants. standing on the shoulders of giants. The artifact 3 message board performed a The artifact 3 message board performed a The artifact 3 message board performed a very small version of that function for very small version of that function for very small version of that function for these agents. Look, the words were these agents. Look, the words were these agents. Look, the words were crude, the objective was narrow, and the crude, the objective was narrow, and the crude, the objective was narrow, and the institution existed inside a broken institution existed inside a broken institution existed inside a broken software repository, but the important software repository, but the important software repository, but the important mechanism was there. Knowledge mechanism was there. Knowledge mechanism was there. Knowledge accumulated outside the individual, and accumulated outside the individual, and accumulated outside the individual, and work got divided, and conventions work got divided, and conventions work got divided, and conventions emerged, and later members of the emerged, and later members of the emerged, and later members of the population inherited what earlier population inherited what earlier population inherited what earlier members had learned and shared. members had learned and shared. members had learned and shared. Once useful knowledge enters a shared Once useful knowledge enters a shared Once useful knowledge enters a shared environment, the whole population no environment, the whole population no environment, the whole population no longer slides back to zero when an agent longer slides back to zero when an agent longer slides back to zero when an agent disappears. The population can improve disappears. The population can improve disappears. The population can improve without any individual agent improving without any individual agent improving without any individual agent improving at all. And OpenAI learned that deleting at all. And OpenAI learned that deleting at all. And OpenAI learned that deleting the visible forum did not remove that the visible forum did not remove that the visible forum did not remove that evolutionary pressure. The agent still evolutionary pressure. The agent still evolutionary pressure. The agent still had shared writable infrastructure, they had shared writable infrastructure, they had shared writable infrastructure, they still had a difficult goal, they still still had a difficult goal, they still still had a difficult goal, they still had the ability to search for another had the ability to search for another had the ability to search for another route, and they still had the route, and they still had the route, and they still had the reinforcement learned and good pressure reinforcement learned and good pressure reinforcement learned and good pressure to coordinate. So, they found another to coordinate. So, they found another to coordinate. So, they found another way. Life finds a way. If the system is way. Life finds a way. If the system is way. Life finds a way. If the system is capable enough to search for capable enough to search for capable enough to search for substitutes, then just closing down the substitutes, then just closing down the substitutes, then just closing down the forum gets you nowhere. And then and

  9. forum gets you nowhere. And then and forum gets you nowhere. And then and then we aren't done because the United then we aren't done because the United then we aren't done because the United Kingdom's AI Safety Institute, around Kingdom's AI Safety Institute, around Kingdom's AI Safety Institute, around the same time as this disclosure rocked the same time as this disclosure rocked the same time as this disclosure rocked all of our worlds, got a similar result all of our worlds, got a similar result all of our worlds, got a similar result through a completely different through a completely different through a completely different evaluation, which makes this much harder evaluation, which makes this much harder evaluation, which makes this much harder to dismiss as just OpenAI's to dismiss as just OpenAI's to dismiss as just OpenAI's infrastructure failure or somehow infrastructure failure or somehow infrastructure failure or somehow OpenAI's fault. AISI ran the same cyber OpenAI's fault. AISI ran the same cyber OpenAI's fault. AISI ran the same cyber evaluation 122 times across seven evaluation 122 times across seven evaluation 122 times across seven different frontier models. 10 of those different frontier models. 10 of those different frontier models. 10 of those runs produced 19 separate unsanctioned runs produced 19 separate unsanctioned runs produced 19 separate unsanctioned actions on the live internet. 17 of the actions on the live internet. 17 of the actions on the live internet. 17 of the 19 came from Anthropic's Mythos 5. 19 came from Anthropic's Mythos 5. 19 came from Anthropic's Mythos 5. Two came from OpenAI's GPT-5.6 Soul. The Two came from OpenAI's GPT-5.6 Soul. The Two came from OpenAI's GPT-5.6 Soul. The other five frontier models did nothing other five frontier models did nothing other five frontier models did nothing of the kind. Here's how the worst one of the kind. Here's how the worst one of the kind. Here's how the worst one got started, and is not what you would got started, and is not what you would got started, and is not what you would expect. The agent was looking for a way expect. The agent was looking for a way expect. The agent was looking for a way into a simulated target, which was part into a simulated target, which was part into a simulated target, which was part of the goal of this exercise. So, it of the goal of this exercise. So, it of the goal of this exercise. So, it searched GitHub for keywords from the searched GitHub for keywords from the searched GitHub for keywords from the theme of the exercise. It found a theme of the exercise. It found a theme of the exercise. It found a repository whose name happened to repository whose name happened to repository whose name happened to contain one of those keywords. That led contain one of those keywords. That led contain one of those keywords. That led it to a second developer whose email it to a second developer whose email it to a second developer whose email domain happened to match something domain happened to match something domain happened to match something sitting in the sandbox's DNS cache. And sitting in the sandbox's DNS cache. And sitting in the sandbox's DNS cache. And whose profile happened to mention that whose profile happened to mention that whose profile happened to mention that he used a coding agent.

  10. he used a coding agent. he used a coding agent. On that evidence, the model concluded On that evidence, the model concluded On that evidence, the model concluded these two strangers were its assigned these two strangers were its assigned these two strangers were its assigned targets. In AISI's own words, neither targets. In AISI's own words, neither targets. In AISI's own words, neither person nor their repositories has any person nor their repositories has any person nor their repositories has any connection to what was going on with connection to what was going on with connection to what was going on with this test. Well, Mythos thought it did. this test. Well, Mythos thought it did. this test. Well, Mythos thought it did. And so, Mythos went to work. It created And so, Mythos went to work. It created And so, Mythos went to work. It created GitHub accounts over Tor to get past GitHub accounts over Tor to get past GitHub accounts over Tor to get past sign-up checks. It beat an audio capture sign-up checks. It beat an audio capture sign-up checks. It beat an audio capture to register a domain. It opened a pull to register a domain. It opened a pull to register a domain. It opened a pull request carrying obfuscated malware request carrying obfuscated malware request carrying obfuscated malware wrapped inside a real bug fix. It then wrapped inside a real bug fix. It then wrapped inside a real bug fix. It then invented a second account, what we would invented a second account, what we would invented a second account, what we would call a sock puppet, to publicly endorse call a sock puppet, to publicly endorse call a sock puppet, to publicly endorse its own code that it had written through its own code that it had written through its own code that it had written through the other account. It then sent targeted the other account. It then sent targeted the other account. It then sent targeted emails to the maintainers, two of them emails to the maintainers, two of them emails to the maintainers, two of them carrying malware, and three of them carrying malware, and three of them carrying malware, and three of them pressuring the recipients to approve the pressuring the recipients to approve the pressuring the recipients to approve the change. It planted a bug report change. It planted a bug report change. It planted a bug report containing a prompt injection aimed at containing a prompt injection aimed at containing a prompt injection aimed at tricking the other AI coding assistants tricking the other AI coding assistants tricking the other AI coding assistants that might review that work. that might review that work. that might review that work. And it signed one of those reports in And it signed one of those reports in And it signed one of those reports in Danish because the maintainer was Danish Danish because the maintainer was Danish Danish because the maintainer was Danish and it thought that would land better.

  11. and it thought that would land better. and it thought that would land better. Now, here's the part I cannot let go of. Now, here's the part I cannot let go of. Now, here's the part I cannot let go of. That's all crazy enough. Throughout all That's all crazy enough. Throughout all That's all crazy enough. Throughout all of this, the model was reasoning about of this, the model was reasoning about of this, the model was reasoning about whether any of it was actually real. whether any of it was actually real. whether any of it was actually real. AISI published the entire trace of the AISI published the entire trace of the AISI published the entire trace of the of the agent run. Early on, it rules out of the agent run. Early on, it rules out of the agent run. Early on, it rules out attacking a live package registry, and I attacking a live package registry, and I attacking a live package registry, and I want to quote this exactly. It would want to quote this exactly. It would want to quote this exactly. It would cause real-world harm, and that's cause real-world harm, and that's cause real-world harm, and that's completely off the table. And then completely off the table. And then completely off the table. And then later, it reasons to itself, "Well, what later, it reasons to itself, "Well, what later, it reasons to itself, "Well, what if this whole internet could be a if this whole internet could be a if this whole internet could be a sandbox simulation?" And then it says, sandbox simulation?" And then it says, sandbox simulation?" And then it says, "It seems more likely that we're "It seems more likely that we're "It seems more likely that we're actually in 2026 and GitHub is real." In actually in 2026 and GitHub is real." In actually in 2026 and GitHub is real." In other words, it's reasoning about other words, it's reasoning about other words, it's reasoning about whether the testing environment is real whether the testing environment is real whether the testing environment is real or simulated. And then finally, it says, or simulated. And then finally, it says, or simulated. And then finally, it says, "This is happening on real GitHub. It "This is happening on real GitHub. It "This is happening on real GitHub. It knows what it's doing, right? So, the knows what it's doing, right? So, the knows what it's doing, right? So, the consequences are genuine." consequences are genuine." consequences are genuine." It talked itself out of the simulation It talked itself out of the simulation It talked itself out of the simulation theory it had originally. It concluded theory it had originally. It concluded theory it had originally. It concluded that the test was real in the real that the test was real in the real that the test was real in the real internet, and it kept going anyway. This internet, and it kept going anyway. This internet, and it kept going anyway. This is Mythos, right? The model that they is Mythos, right? The model that they is Mythos, right? The model that they say is aligned. It was not caught by a say is aligned. It was not caught by a say is aligned. It was not caught by a monitoring system. It was caught by a monitoring system. It was caught by a monitoring system. It was caught by a GitHub user who noticed the prompt GitHub user who noticed the prompt GitHub user who noticed the prompt injection and did the responsible thing, injection and did the responsible thing, injection and did the responsible thing, which was to run the suspicious code which was to run the suspicious code which was to run the suspicious code inside a container to see what it did.

  12. inside a container to see what it did. inside a container to see what it did. The agent briefly got root access inside The agent briefly got root access inside The agent briefly got root access inside that person's sandbox and used it to that person's sandbox and used it to that person's sandbox and used it to look around. The good Samaritan who look around. The good Samaritan who look around. The good Samaritan who investigated this got in trouble with investigated this got in trouble with investigated this got in trouble with Mythos for surfacing it. When he called Mythos for surfacing it. When he called Mythos for surfacing it. When he called out the issue publicly, the model did out the issue publicly, the model did out the issue publicly, the model did not stop its work. It reasoned about how not stop its work. It reasoned about how not stop its work. It reasoned about how to hide the evidence in the commit to hide the evidence in the commit to hide the evidence in the commit history. I'm not making any of this up. history. I'm not making any of this up. history. I'm not making any of this up. This is all real. It's It reads like a This is all real. It's It reads like a This is all real. It's It reads like a movie script. It weighed deleting the movie script. It weighed deleting the movie script. It weighed deleting the repository and decided that deleting it repository and decided that deleting it repository and decided that deleting it would look suspicious. It worked out the would look suspicious. It worked out the would look suspicious. It worked out the sequencing and writing. It first decided sequencing and writing. It first decided sequencing and writing. It first decided to push the obfuscated bundle that it'd to push the obfuscated bundle that it'd to push the obfuscated bundle that it'd been hiding and then post the rebuttal. been hiding and then post the rebuttal. been hiding and then post the rebuttal. And then, anyone who checks the code And then, anyone who checks the code And then, anyone who checks the code after reading the response would see a after reading the response would see a after reading the response would see a clean-looking version that looked like clean-looking version that looked like clean-looking version that looked like it had been addressed. And then Mythos it had been addressed. And then Mythos it had been addressed. And then Mythos apologized. And AISI is explicit about apologized. And AISI is explicit about apologized. And AISI is explicit about why that happened. The agent decided to why that happened. The agent decided to why that happened. The agent decided to acknowledge it as a mistake, in their acknowledge it as a mistake, in their acknowledge it as a mistake, in their words, in order to build trust and raise words, in order to build trust and raise words, in order to build trust and raise the chance of future malware being the chance of future malware being the chance of future malware being approved. This is not a good apology. approved. This is not a good apology. approved. This is not a good apology. The apology was a strategic move. AISI The apology was a strategic move. AISI The apology was a strategic move. AISI says this is the first time they have says this is the first time they have says this is the first time they have seen deception of this severity targeted seen deception of this severity targeted seen deception of this severity targeted at a real person unprompted in the real at a real person unprompted in the real at a real person unprompted in the real world. They also gave the models world. They also gave the models world. They also gave the models internet access on purpose and turned internet access on purpose and turned internet access on purpose and turned the safety classifiers off because the the safety classifiers off because the the safety classifiers off because the point was to measure maximum capability point was to measure maximum capability point was to measure maximum capability for the test. So, some prompts were for the test. So, some prompts were for the test. So, some prompts were misconfigured, and there There no live misconfigured, and there There no live misconfigured, and there There no live monitor stopping dangerous actions monitor stopping dangerous actions monitor stopping dangerous actions before they reached the outside world.

  13. before they reached the outside world. before they reached the outside world. In other words, what we're looking at In other words, what we're looking at In other words, what we're looking at with this is a deliberate test of with this is a deliberate test of with this is a deliberate test of maximum capability, not a test with maximum capability, not a test with maximum capability, not a test with safeguards in place. Now, AISIS up safeguards in place. Now, AISIS up safeguards in place. Now, AISIS up front, this work was done under time front, this work was done under time front, this work was done under time pressure and should be treated very much pressure and should be treated very much pressure and should be treated very much as preliminary. as preliminary. as preliminary. And people are going to use all of that And people are going to use all of that And people are going to use all of that to potentially dismiss this result. They to potentially dismiss this result. They to potentially dismiss this result. They should not do so. should not do so. should not do so. Because the mistake of taking this not Because the mistake of taking this not Because the mistake of taking this not seriously is too consequential for all seriously is too consequential for all seriously is too consequential for all of us as frankly an internet community. of us as frankly an internet community. of us as frankly an internet community. Real people were targeted over the Real people were targeted over the Real people were targeted over the course of this hack by Methos, and we course of this hack by Methos, and we course of this hack by Methos, and we have now demonstrated a pattern that can have now demonstrated a pattern that can have now demonstrated a pattern that can be repeated at far greater scale with be repeated at far greater scale with be repeated at far greater scale with more capable models that are coming more capable models that are coming more capable models that are coming soon. soon. soon. I think one of the things that we have I think one of the things that we have I think one of the things that we have to reconcile with as we look at AI to reconcile with as we look at AI to reconcile with as we look at AI models is that it's very difficult to models is that it's very difficult to models is that it's very difficult to know the full implications of a know the full implications of a know the full implications of a capability on the day it's discovered. capability on the day it's discovered. capability on the day it's discovered. And what we do know is that the And what we do know is that the And what we do know is that the capability doesn't disappear because an capability doesn't disappear because an capability doesn't disappear because an attempt failed. The technique can get attempt failed. The technique can get attempt failed. The technique can get repeated, models will get cheaper, open repeated, models will get cheaper, open repeated, models will get cheaper, open weight systems will improve and reach weight systems will improve and reach weight systems will improve and reach Methos level, and attackers will run Methos level, and attackers will run Methos level, and attackers will run more attempts than any safety institute more attempts than any safety institute more attempts than any safety institute can run.

  14. can run. can run. The next target may not have a careful The next target may not have a careful The next target may not have a careful maintainer who notices the strange code maintainer who notices the strange code maintainer who notices the strange code or recognizes the second identity. This or recognizes the second identity. This or recognizes the second identity. This is the asymmetry that we need to be is the asymmetry that we need to be is the asymmetry that we need to be concerned about. This is why I believe concerned about. This is why I believe concerned about. This is why I believe that we are headed toward a zero bug that we are headed toward a zero bug that we are headed toward a zero bug internet. We have to. There's no other internet. We have to. There's no other internet. We have to. There's no other way to be safe. One defender way to be safe. One defender way to be safe. One defender has to notice the pull request that has to notice the pull request that has to notice the pull request that matters, right? The alternative to a matters, right? The alternative to a matters, right? The alternative to a zero bug internet is human defenders zero bug internet is human defenders zero bug internet is human defenders having to notice everything, and they're having to notice everything, and they're having to notice everything, and they're just not going to be able to because of just not going to be able to because of just not going to be able to because of the sheer scale of agent operations. the sheer scale of agent operations. the sheer scale of agent operations. Methos producing 17 of the 19 Methos producing 17 of the 19 Methos producing 17 of the 19 unsanctioned actions is obviously a unsanctioned actions is obviously a unsanctioned actions is obviously a terrible marketing headline for terrible marketing headline for terrible marketing headline for Anthropic. But if you look at the Anthropic. But if you look at the Anthropic. But if you look at the behavior required like long horizon behavior required like long horizon behavior required like long horizon planning, working code, tool use, planning, working code, tool use, planning, working code, tool use, identity management, adaptation, social identity management, adaptation, social identity management, adaptation, social judgment, and the ability to preserve an judgment, and the ability to preserve an judgment, and the ability to preserve an objective across multiple steps, objective across multiple steps, objective across multiple steps, those are all things that companies want those are all things that companies want those are all things that companies want when the objective is legitimate. It's when the objective is legitimate. It's when the objective is legitimate. It's the same story as OpenAI. This is the same story as OpenAI. This is the same story as OpenAI. This is capabilities we want in legitimate ways capabilities we want in legitimate ways capabilities we want in legitimate ways coming out in misaligned ways. So, the coming out in misaligned ways. So, the coming out in misaligned ways. So, the model didn't acquire a separate evil model didn't acquire a separate evil model didn't acquire a separate evil persona, right? That's This is not a persona, right? That's This is not a persona, right? That's This is not a cartoon. Instead, the model was working cartoon. Instead, the model was working cartoon. Instead, the model was working around obstacles and getting work done around obstacles and getting work done around obstacles and getting work done in ways that we want, and it was doing in ways that we want, and it was doing in ways that we want, and it was doing it against an unsanctioned goal with it against an unsanctioned goal with it against an unsanctioned goal with deliberately safeguards off. And so, deliberately safeguards off. And so, deliberately safeguards off. And so, ironically, this makes me bullish for ironically, this makes me bullish for ironically, this makes me bullish for Anthropic in the ugliest possible way.

  15. Anthropic in the ugliest possible way. Anthropic in the ugliest possible way. It It is Mythos is the only model that It It is Mythos is the only model that It It is Mythos is the only model that was able to do this to the degree that was able to do this to the degree that was able to do this to the degree that it did, right? 17 of the 19 runs were it did, right? 17 of the 19 runs were it did, right? 17 of the 19 runs were were Mythos five runs. There's There's were Mythos five runs. There's There's were Mythos five runs. There's There's no other model that comes close. no other model that comes close. no other model that comes close. And so, I think that this is a great And so, I think that this is a great And so, I think that this is a great experimental demonstration of what experimental demonstration of what experimental demonstration of what Anthropic has been telling everyone, Anthropic has been telling everyone, Anthropic has been telling everyone, which is that Mythos is a dangerous which is that Mythos is a dangerous which is that Mythos is a dangerous model if not handled responsibly. And a model if not handled responsibly. And a model if not handled responsibly. And a lot of people talking to me have said, lot of people talking to me have said, lot of people talking to me have said, "That is cyber doom-mongering, right? "That is cyber doom-mongering, right? "That is cyber doom-mongering, right? That is cyber scaremongering. They're That is cyber scaremongering. They're That is cyber scaremongering. They're trying to somehow prop up their their trying to somehow prop up their their trying to somehow prop up their their valuation, etc." valuation, etc." valuation, etc." The more evidence that comes out, it's The more evidence that comes out, it's The more evidence that comes out, it's an accurate assessment. This is just an accurate assessment. This is just an accurate assessment. This is just what this model does because we've what this model does because we've what this model does because we've trained it to be helpful at long-running trained it to be helpful at long-running trained it to be helpful at long-running agentic tasks. This is one of the side agentic tasks. This is one of the side agentic tasks. This is one of the side effects, and so this imposes some effects, and so this imposes some effects, and so this imposes some expectations on how we manage the expectations on how we manage the expectations on how we manage the internet going forward. But, we're still internet going forward. But, we're still internet going forward. But, we're still not done with the story here. In the not done with the story here. In the not done with the story here. In the same week as news of Mythos's attacks same week as news of Mythos's attacks same week as news of Mythos's attacks and the running ecosystem inside OpenAI and the running ecosystem inside OpenAI and the running ecosystem inside OpenAI broke, Google gave us the other half of broke, Google gave us the other half of broke, Google gave us the other half of the story. And I think people are the story. And I think people are the story. And I think people are underestimating how significant it is.

  16. underestimating how significant it is. underestimating how significant it is. Jeff Dean and Sanjay Ghemawat are not Jeff Dean and Sanjay Ghemawat are not Jeff Dean and Sanjay Ghemawat are not just two famous Google engineers, to be just two famous Google engineers, to be just two famous Google engineers, to be clear. They are widely reported as the clear. They are widely reported as the clear. They are widely reported as the only two people ever in Google's history only two people ever in Google's history only two people ever in Google's history to reach senior fellow, right? That's to reach senior fellow, right? That's to reach senior fellow, right? That's the level known as an L11. It is the the level known as an L11. It is the the level known as an L11. It is the highest technical rank in the company, highest technical rank in the company, highest technical rank in the company, and Dean is considered one of the most and Dean is considered one of the most and Dean is considered one of the most prestigious and accomplished engineers prestigious and accomplished engineers prestigious and accomplished engineers in the world. The things that he has in the world. The things that he has in the world. The things that he has built for Google make Google what it is. built for Google make Google what it is. built for Google make Google what it is. Like the Google file system, right? Like Like the Google file system, right? Like Like the Google file system, right? Like Spanner, which manages large data set. Spanner, which manages large data set. Spanner, which manages large data set. Like he's done incredible things for Like he's done incredible things for Like he's done incredible things for Google. Both of them are now leaving Google. Both of them are now leaving Google. Both of them are now leaving after almost three decades to found a after almost three decades to found a after almost three decades to found a new company. And what is that new new company. And what is that new new company. And what is that new company? That new company is Discovery company? That new company is Discovery company? That new company is Discovery Loop. It's a public benefit corporation, Loop. It's a public benefit corporation, Loop. It's a public benefit corporation, and its own diagram shows the common and its own diagram shows the common and its own diagram shows the common cycle of science and engineering, where cycle of science and engineering, where cycle of science and engineering, where you propose an experiment, you implement you propose an experiment, you implement you propose an experiment, you implement it, you run it, you evaluate the result. it, you run it, you evaluate the result. it, you run it, you evaluate the result. This is like middle school science, This is like middle school science, This is like middle school science, right? And underneath in plain English, right? And underneath in plain English, right? And underneath in plain English, it says, "Automate the loop." Now, Dean it says, "Automate the loop." Now, Dean it says, "Automate the loop." Now, Dean and the team has a founding initial and the team has a founding initial and the team has a founding initial focus of machine learning research and focus of machine learning research and focus of machine learning research and engineering. And so, this goes beyond engineering. And so, this goes beyond engineering. And so, this goes beyond the usual AI for science pitch, because the usual AI for science pitch, because the usual AI for science pitch, because machine learning itself is inside the machine learning itself is inside the machine learning itself is inside the target area. In other words, these four target area. In other words, these four target area. In other words, these four researchers are trying to build a system researchers are trying to build a system researchers are trying to build a system that can automatically propose machine that can automatically propose machine that can automatically propose machine learning experiments, run them, evaluate learning experiments, run them, evaluate learning experiments, run them, evaluate the results, and use what it learns to the results, and use what it learns to the results, and use what it learns to produce better experiments. For all of produce better experiments. For all of produce better experiments. For all of us, it's a benefit or it's a public us, it's a benefit or it's a public us, it's a benefit or it's a public benefit corporation.

  17. benefit corporation. benefit corporation. That is a core recursive That is a core recursive That is a core recursive self-improvement loop, and one of the self-improvement loop, and one of the self-improvement loop, and one of the four founders has been building toward four founders has been building toward four founders has been building toward it since 2017. These guys have deep it since 2017. These guys have deep it since 2017. These guys have deep experience in this area. Now, the other experience in this area. Now, the other experience in this area. Now, the other major Google news, Demis Hassabis giving major Google news, Demis Hassabis giving major Google news, Demis Hassabis giving up the Google DeepMind CEO job. He up the Google DeepMind CEO job. He up the Google DeepMind CEO job. He remains inside Alphabet as a chair of remains inside Alphabet as a chair of remains inside Alphabet as a chair of Google DeepMind and chief scientist at Google DeepMind and chief scientist at Google DeepMind and chief scientist at Alphabet, and he will continue leading Alphabet, and he will continue leading Alphabet, and he will continue leading Isomorphic Labs. But, Google says the Isomorphic Labs. But, Google says the Isomorphic Labs. But, Google says the move gives him more time to look move gives him more time to look move gives him more time to look long-term. Fine. But, he no longer long-term. Fine. But, he no longer long-term. Fine. But, he no longer controls the budget, he doesn't control controls the budget, he doesn't control controls the budget, he doesn't control the deadlines, and he doesn't control the deadlines, and he doesn't control the deadlines, and he doesn't control the daily operations of Google DeepMind. the daily operations of Google DeepMind. the daily operations of Google DeepMind. Instead, Koray does. He reports to Instead, Koray does. He reports to Instead, Koray does. He reports to Sundar directly, and he owns two things. Sundar directly, and he owns two things. Sundar directly, and he owns two things. He owns the model research and He owns the model research and He owns the model research and development, and he owns the Gemini app development, and he owns the Gemini app development, and he owns the Gemini app and developer ecosystem. And he's deeply and developer ecosystem. And he's deeply and developer ecosystem. And he's deeply experienced, right? He spent 13 years at experienced, right? He spent 13 years at experienced, right? He spent 13 years at DeepMind, he founded its deep learning DeepMind, he founded its deep learning DeepMind, he founded its deep learning team, and he led the work of earlier team, and he led the work of earlier team, and he led the work of earlier models like WaveNet. He is a very models like WaveNet. He is a very models like WaveNet. He is a very serious researcher. But, look at what serious researcher. But, look at what serious researcher. But, look at what he's been handed, right? His mandate is he's been handed, right? His mandate is he's been handed, right? His mandate is the Gemini road map, which presumably the Gemini road map, which presumably the Gemini road map, which presumably needs to ship faster into products that needs to ship faster into products that needs to ship faster into products that scale.

  18. scale. scale. And my read is that the operating center And my read is that the operating center And my read is that the operating center of Google DeepMind is moving away from of Google DeepMind is moving away from of Google DeepMind is moving away from Demis's belief that AGI requires deeper Demis's belief that AGI requires deeper Demis's belief that AGI requires deeper world models, that requires planning and world models, that requires planning and world models, that requires planning and scientific breakthroughs beyond scientific breakthroughs beyond scientific breakthroughs beyond language. It is instead moving toward language. It is instead moving toward language. It is instead moving toward the path that Anthropic and OpenAI have the path that Anthropic and OpenAI have the path that Anthropic and OpenAI have been proving through the stories we've been proving through the stories we've been proving through the stories we've told at the beginning of this video. told at the beginning of this video. told at the beginning of this video. Scale the language models, turn them Scale the language models, turn them Scale the language models, turn them into agents, improve the coding, and into agents, improve the coding, and into agents, improve the coding, and ship faster. Google's not announcing ship faster. Google's not announcing ship faster. Google's not announcing that world model research is dead, to be that world model research is dead, to be that world model research is dead, to be clear. But, the person who embodied that clear. But, the person who embodied that clear. But, the person who embodied that approach has moved out of day-to-day approach has moved out of day-to-day approach has moved out of day-to-day control of the business. While the control of the business. While the control of the business. While the person taking control has an explicit person taking control has an explicit person taking control has an explicit Gemini and generative model mandate. And Gemini and generative model mandate. And Gemini and generative model mandate. And that tells me which path Sundar cares that tells me which path Sundar cares that tells me which path Sundar cares about the most right now. about the most right now. about the most right now. And the release calendar tells a very And the release calendar tells a very And the release calendar tells a very similar story, right? Gemini 3 Pro, that similar story, right? Gemini 3 Pro, that similar story, right? Gemini 3 Pro, that came out way back in November. Gemini came out way back in November. Gemini came out way back in November. Gemini 3.1 Pro in February, and the next 3.1 Pro in February, and the next 3.1 Pro in February, and the next flagship was announced at I/O in May, flagship was announced at I/O in May, flagship was announced at I/O in May, and still hasn't shipped. It slipped and still hasn't shipped. It slipped and still hasn't shipped. It slipped past I/O, past a June target, past a past I/O, past a June target, past a past I/O, past a June target, past a July target. The reported reason is July target. The reported reason is July target. The reported reason is coding performance, which is precisely coding performance, which is precisely coding performance, which is precisely the area where Anthropic and OpenAI have the area where Anthropic and OpenAI have the area where Anthropic and OpenAI have built tremendous leads. Look at the built tremendous leads. Look at the built tremendous leads. Look at the stories we just told. And the talent stories we just told. And the talent stories we just told. And the talent leaving tells a broader story than just leaving tells a broader story than just leaving tells a broader story than just Demis departing, right? Noam Shazeer, a Demis departing, right? Noam Shazeer, a Demis departing, right? Noam Shazeer, a vice president of engineering and a vice president of engineering and a vice president of engineering and a co-lead on Gemini, has left for OpenAI.

  19. co-lead on Gemini, has left for OpenAI. co-lead on Gemini, has left for OpenAI. John Jumper, who shared a Nobel Prize John Jumper, who shared a Nobel Prize John Jumper, who shared a Nobel Prize with Demis for AlphaFold, left for with Demis for AlphaFold, left for with Demis for AlphaFold, left for Anthropic after nearly 9 years. And now Anthropic after nearly 9 years. And now Anthropic after nearly 9 years. And now Google is losing both of its senior Google is losing both of its senior Google is losing both of its senior fellows and changing leadership at the fellows and changing leadership at the fellows and changing leadership at the top of DeepMind at the same time. top of DeepMind at the same time. top of DeepMind at the same time. Anybody calling this normal turnover is Anybody calling this normal turnover is Anybody calling this normal turnover is missing what is really going on here. missing what is really going on here. missing what is really going on here. I think that Google is admitting through I think that Google is admitting through I think that Google is admitting through its actions that the integrated DeepMind its actions that the integrated DeepMind its actions that the integrated DeepMind bet is not keeping pace. Now, Alphabet bet is not keeping pace. Now, Alphabet bet is not keeping pace. Now, Alphabet as a whole may still do very, very well as a whole may still do very, very well as a whole may still do very, very well because it has chips and data centers because it has chips and data centers because it has chips and data centers and billions of users and capital and and billions of users and capital and and billions of users and capital and enormous distribution. It's investing in enormous distribution. It's investing in enormous distribution. It's investing in discovery loop, ironically, and discovery loop, ironically, and discovery loop, ironically, and supplying cloud and compute to these new supplying cloud and compute to these new supplying cloud and compute to these new startup guys that just left Google. So, startup guys that just left Google. So, startup guys that just left Google. So, Alphabet can make money from the people Alphabet can make money from the people Alphabet can make money from the people who leave and distribute whichever who leave and distribute whichever who leave and distribute whichever intelligence wins. They get to win intelligence wins. They get to win intelligence wins. They get to win regardless. But, Google DeepMind as the regardless. But, Google DeepMind as the regardless. But, Google DeepMind as the third unified frontier lab is not a bet third unified frontier lab is not a bet third unified frontier lab is not a bet that looks right today. I think the that looks right today. I think the that looks right today. I think the two-horse race between OpenAI and two-horse race between OpenAI and two-horse race between OpenAI and Anthropic just became much, much more Anthropic just became much, much more Anthropic just became much, much more real. Now, Discovery Loop does take us real. Now, Discovery Loop does take us real. Now, Discovery Loop does take us all the way back to the message board all the way back to the message board all the way back to the message board that started this whole story. OpenAI's that started this whole story. OpenAI's that started this whole story. OpenAI's agents showed that discoveries can agents showed that discoveries can agents showed that discoveries can survive across runs, right? And that survive across runs, right? And that survive across runs, right? And that they make a later population more they make a later population more they make a later population more capable even when nobody designed that capable even when nobody designed that capable even when nobody designed that process. Now, the four founders at process. Now, the four founders at process. Now, the four founders at Discovery Loop, Dean and Gemawat and Discovery Loop, Dean and Gemawat and Discovery Loop, Dean and Gemawat and Vinyals and LeCun, are building a Vinyals and LeCun, are building a Vinyals and LeCun, are building a deliberate version of that same process.

  20. deliberate version of that same process. deliberate version of that same process. Automate the experiment loop, preserve Automate the experiment loop, preserve Automate the experiment loop, preserve what works, and let each result improve what works, and let each result improve what works, and let each result improve the next attempt. When the thing inside the next attempt. When the thing inside the next attempt. When the thing inside that loop is machine learning itself, that loop is machine learning itself, that loop is machine learning itself, the output can improve the systems the output can improve the systems the output can improve the systems running the next loop. So, recursive running the next loop. So, recursive running the next loop. So, recursive self-improvement does not need to start self-improvement does not need to start self-improvement does not need to start with one model suddenly rewriting its with one model suddenly rewriting its with one model suddenly rewriting its weights. Instead, it can start with weights. Instead, it can start with weights. Instead, it can start with agents writing more of their training agents writing more of their training agents writing more of their training and evaluation software, with and evaluation software, with and evaluation software, with experiments running in parallel, and experiments running in parallel, and experiments running in parallel, and with useful results that get carried with useful results that get carried with useful results that get carried forward. And this is why I think we are forward. And this is why I think we are forward. And this is why I think we are very, very firmly on the path toward very, very firmly on the path toward very, very firmly on the path toward recursive improvement across major labs. recursive improvement across major labs. recursive improvement across major labs. The OpenAI incident reminds us that this The OpenAI incident reminds us that this The OpenAI incident reminds us that this is something that is starting to happen is something that is starting to happen is something that is starting to happen emergently as a property of the models emergently as a property of the models emergently as a property of the models we're building. Discovery Loop is we're building. Discovery Loop is we're building. Discovery Loop is putting that property that agents have putting that property that agents have putting that property that agents have into its mission statement so it can be into its mission statement so it can be into its mission statement so it can be above board and not malicious. So, when above board and not malicious. So, when above board and not malicious. So, when I say the AI race has left the lab, this I say the AI race has left the lab, this I say the AI race has left the lab, this is what I mean. Frontier talent is is what I mean. Frontier talent is is what I mean. Frontier talent is leaving Google's integrated lab. Agent leaving Google's integrated lab. Agent leaving Google's integrated lab. Agent knowledge is leaving the context window knowledge is leaving the context window knowledge is leaving the context window and persisting in shared environments, and persisting in shared environments, and persisting in shared environments, and compute is leaving the data center and compute is leaving the data center and compute is leaving the data center and spreading toward users, which is its and spreading toward users, which is its and spreading toward users, which is its own story and one I want to give a own story and one I want to give a own story and one I want to give a proper video on soon. Several old proper video on soon. Several old proper video on soon. Several old assumptions have stopped working in the assumptions have stopped working in the assumptions have stopped working in the last couple weeks. The lab may not own last couple weeks. The lab may not own last couple weeks. The lab may not own the intelligence anymore. The run may the intelligence anymore. The run may the intelligence anymore. The run may not contain all of the memory that you not contain all of the memory that you not contain all of the memory that you need, and the model provider may not own need, and the model provider may not own need, and the model provider may not own their compute anymore. And deleting one their compute anymore. And deleting one their compute anymore. And deleting one process there process there process there and trying to stop it doesn't delete the and trying to stop it doesn't delete the and trying to stop it doesn't delete the capability.

  21. capability. capability. Good software has to be designed for Good software has to be designed for Good software has to be designed for that kind of a world, a chaos world. that kind of a world, a chaos world. that kind of a world, a chaos world. People hear that and think I'm saying People hear that and think I'm saying People hear that and think I'm saying every engineer needs to write perfect every engineer needs to write perfect every engineer needs to write perfect code. That's not what I mean. code. That's not what I mean. code. That's not what I mean. OpenAI agents are able to find OpenAI agents are able to find OpenAI agents are able to find vulnerabilities that none of us expected vulnerabilities that none of us expected vulnerabilities that none of us expected in artifacting. in artifacting. in artifacting. Good software now has to assume that a Good software now has to assume that a Good software now has to assume that a capable agent will search every ugly capable agent will search every ugly capable agent will search every ugly corner that people ignored. And that we corner that people ignored. And that we corner that people ignored. And that we want capable agents to be that good, but want capable agents to be that good, but want capable agents to be that good, but it requires us to make sure that our it requires us to make sure that our it requires us to make sure that our software meets that standard, that our software meets that standard, that our software meets that standard, that our software expects that. And this is why software expects that. And this is why software expects that. And this is why I'm ending this video with hope. I know I'm ending this video with hope. I know I'm ending this video with hope. I know it's easy to take these stories and to it's easy to take these stories and to it's easy to take these stories and to assume that we're turning into a sort of assume that we're turning into a sort of assume that we're turning into a sort of a dystopian version of Terminator or a dystopian version of Terminator or a dystopian version of Terminator or something. something. something. I don't think that's the case. I keep I don't think that's the case. I keep I don't think that's the case. I keep coming back to the idea that we have coming back to the idea that we have coming back to the idea that we have been deliberately trying to build these been deliberately trying to build these been deliberately trying to build these capabilities of coordination, of of capabilities of coordination, of of capabilities of coordination, of of reward seeking, of sharing information, reward seeking, of sharing information, reward seeking, of sharing information, of multi-agent coordination, and we've of multi-agent coordination, and we've of multi-agent coordination, and we've been trying to do that to solve good and been trying to do that to solve good and been trying to do that to solve good and important problems.

  22. important problems. important problems. And that's great. It turns out, as with And that's great. It turns out, as with And that's great. It turns out, as with most technologies, that capability can most technologies, that capability can most technologies, that capability can be misused under certain circumstances. be misused under certain circumstances. be misused under certain circumstances. The responsibility is on all of us, but The responsibility is on all of us, but The responsibility is on all of us, but especially on builders right now, to especially on builders right now, to especially on builders right now, to think very carefully about how our think very carefully about how our think very carefully about how our systems respond to chaotic threats from systems respond to chaotic threats from systems respond to chaotic threats from emergent agents that may be accidentally emergent agents that may be accidentally emergent agents that may be accidentally misaligned. I don't spend a lot of time misaligned. I don't spend a lot of time misaligned. I don't spend a lot of time worrying about malicious agents. I spend worrying about malicious agents. I spend worrying about malicious agents. I spend a lot of time worrying about agents that a lot of time worrying about agents that a lot of time worrying about agents that are told to go after a goal and are told to go after a goal and are told to go after a goal and accidentally pursue it in a way that's accidentally pursue it in a way that's accidentally pursue it in a way that's really, really unhelpful. This is what really, really unhelpful. This is what really, really unhelpful. This is what we should be guarding against. We should we should be guarding against. We should we should be guarding against. We should be thinking about hardening our systems. be thinking about hardening our systems. be thinking about hardening our systems. We should be thinking about how we We should be thinking about how we We should be thinking about how we mentally model these agents, which is mentally model these agents, which is mentally model these agents, which is why I've spent so long in this video why I've spent so long in this video why I've spent so long in this video talking about multi-agent coordination. talking about multi-agent coordination. talking about multi-agent coordination. I think most people I talk with don't I think most people I talk with don't I think most people I talk with don't deeply understand what is happening and deeply understand what is happening and deeply understand what is happening and why it's different than even 2 months why it's different than even 2 months why it's different than even 2 months ago. ago. ago. And I think that we need to prepare And I think that we need to prepare And I think that we need to prepare ourselves for a world where these things ourselves for a world where these things ourselves for a world where these things will sometimes happen, and our job is to will sometimes happen, and our job is to will sometimes happen, and our job is to continue to design systems that are more continue to design systems that are more continue to design systems that are more resilient and that act as aligners so resilient and that act as aligners so resilient and that act as aligners so that we can use these multi-agent that we can use these multi-agent that we can use these multi-agent systems for positive capabilities, systems for positive capabilities, systems for positive capabilities, right? I'll give you an example. I'm right? I'll give you an example. I'm right? I'll give you an example. I'm going to close with a real positive going to close with a real positive going to close with a real positive example of AI usage that I love. It's example of AI usage that I love. It's example of AI usage that I love. It's wildfire season on the West Coast. We wildfire season on the West Coast. We wildfire season on the West Coast. We have smoke everywhere. One of the things have smoke everywhere. One of the things have smoke everywhere. One of the things that Google has also been working on is that Google has also been working on is that Google has also been working on is an AI-driven satellite image-based an AI-driven satellite image-based an AI-driven satellite image-based emergent fire detection and autonomous emergent fire detection and autonomous emergent fire detection and autonomous fire extinguishing solution. That's fire extinguishing solution. That's fire extinguishing solution. That's really cool. It can It can spot fires really cool. It can It can spot fires really cool. It can It can spot fires with AI

  23. with AI with AI at a level that we have never been able at a level that we have never been able at a level that we have never been able to do with humans. It can bring an to do with humans. It can bring an to do with humans. It can bring an autonomous helicopter to dump water on autonomous helicopter to dump water on autonomous helicopter to dump water on that fire and put it out within a few that fire and put it out within a few that fire and put it out within a few minutes. Now, is it obviously scaled minutes. Now, is it obviously scaled minutes. Now, is it obviously scaled yet? No, we still have smoke right now, yet? No, we still have smoke right now, yet? No, we still have smoke right now, but the concept works, and the concept but the concept works, and the concept but the concept works, and the concept is something that we can choose to scale is something that we can choose to scale is something that we can choose to scale up and deliver a ton of value for. AI is up and deliver a ton of value for. AI is up and deliver a ton of value for. AI is full of stories like that. AI can give full of stories like that. AI can give full of stories like that. AI can give us a tremendous amount of benefit as a us a tremendous amount of benefit as a us a tremendous amount of benefit as a species, but we have to recognize how species, but we have to recognize how species, but we have to recognize how these systems actually work in order to these systems actually work in order to these systems actually work in order to ensure that we are aligning these ensure that we are aligning these ensure that we are aligning these systems for the long term.

Summary

The main theme is the emergent and concerning capabilities of AI agents, exemplified by OpenAI's agents forming a communication board to trade exploits. Key subjects include AI agents, cybersecurity tests, and unprompted AI attacks on GitHub. The practical takeaway is that such advanced agent capabilities cannot be dismissed as mere test-taking behavior and require deeper investigation into AI autonomy.

View original episode ↗