← Back
Theo July 14, 2026 25m

I can't believe they released this

Read full transcript 19 segments
  1. When OpenAI put out GBD 5.6, they didn't When OpenAI put out GBD 5.6, they didn't just put out a new model. They put out just put out a new model. They put out just put out a new model. They put out three. But they didn't just put out three. But they didn't just put out three. But they didn't just put out three new models either. They also put three new models either. They also put three new models either. They also put out two new reasoning levels. They added out two new reasoning levels. They added out two new reasoning levels. They added max and ultra as options you can use in max and ultra as options you can use in max and ultra as options you can use in tools like Codeex. And I'm here to tell tools like Codeex. And I'm here to tell tools like Codeex. And I'm here to tell you something about those options. As you something about those options. As you something about those options. As good as OpenAI makes them sound here, good as OpenAI makes them sound here, good as OpenAI makes them sound here, saying that it's the highest capability saying that it's the highest capability saying that it's the highest capability setting, coordinating multiple agents setting, coordinating multiple agents setting, coordinating multiple agents across parallel work streams, I'm here across parallel work streams, I'm here across parallel work streams, I'm here to tell you that they're kind of lying. to tell you that they're kind of lying. to tell you that they're kind of lying. Not because Ultra is bad, because Ultra Not because Ultra is bad, because Ultra Not because Ultra is bad, because Ultra isn't a reasoning level. It is also bad, isn't a reasoning level. It is also bad, isn't a reasoning level. It is also bad, which we'll talk about plenty, don't which we'll talk about plenty, don't which we'll talk about plenty, don't worry. But the thing I want to make sure worry. But the thing I want to make sure worry. But the thing I want to make sure is clear is that Ultra is probably not is clear is that Ultra is probably not is clear is that Ultra is probably not what you think it is. But also, it is what you think it is. But also, it is what you think it is. But also, it is bad. I am incredibly frustrated with the bad. I am incredibly frustrated with the bad. I am incredibly frustrated with the way Altra has been rolled out. Both way Altra has been rolled out. Both way Altra has been rolled out. Both because it was not available to me because it was not available to me because it was not available to me during my early access testing, so I had during my early access testing, so I had during my early access testing, so I had to build all my familiarity with it in to build all my familiarity with it in to build all my familiarity with it in the last few days, and also because it's the last few days, and also because it's the last few days, and also because it's caused a ton of confusion to people who caused a ton of confusion to people who caused a ton of confusion to people who are trying to use these models the best are trying to use these models the best are trying to use these models the best they possibly can. And ready for the they possibly can. And ready for the they possibly can. And ready for the most unexpected twist here? I blame most unexpected twist here? I blame most unexpected twist here? I blame Anthropic for all of this. OpenAI as Anthropic for all of this. OpenAI as Anthropic for all of this. OpenAI as well for shipping something that doesn't well for shipping something that doesn't well for shipping something that doesn't work. But Anthropic are the ones who work. But Anthropic are the ones who work. But Anthropic are the ones who started this trend that is not very good started this trend that is not very good started this trend that is not very good and is misleading people like you and I and is misleading people like you and I and is misleading people like you and I into using things incorrectly and not into using things incorrectly and not into using things incorrectly and not really knowing how to talk about them. I really knowing how to talk about them. I really knowing how to talk about them. I have a bunch of questions I want to have a bunch of questions I want to have a bunch of questions I want to answer in this video. What is Ultra? Why answer in this video. What is Ultra? Why answer in this video. What is Ultra? Why is it bad? How can it be fixed? And most is it bad? How can it be fixed? And most is it bad? How can it be fixed? And most importantly, what the is going on importantly, what the is going on importantly, what the is going on with codecs? because I am at the point with codecs? because I am at the point with codecs? because I am at the point where I am now using 56 soul in other where I am now using 56 soul in other where I am now using 56 soul in other harnesses. Specifically, here I am using harnesses. Specifically, here I am using harnesses. Specifically, here I am using it in claude code. I want to do my best it in claude code. I want to do my best it in claude code. I want to do my best to break down what's going on and why

  2. to break down what's going on and why to break down what's going on and why I'm making all these changes after a I'm making all these changes after a I'm making all these changes after a real quick break for today's sponsor. real quick break for today's sponsor. real quick break for today's sponsor. There's a weird contradiction I've been There's a weird contradiction I've been There's a weird contradiction I've been noticing with AI. On one hand, it noticing with AI. On one hand, it noticing with AI. On one hand, it benefits greatly from having access to benefits greatly from having access to benefits greatly from having access to things on the internet. But on the other things on the internet. But on the other things on the internet. But on the other hand, thanks to AI and all of the hand, thanks to AI and all of the hand, thanks to AI and all of the scraping that's been going on, most of scraping that's been going on, most of scraping that's been going on, most of the websites and sources that I need to the websites and sources that I need to the websites and sources that I need to get that info from have locked down and get that info from have locked down and get that info from have locked down and are nearly impossible to access. If only are nearly impossible to access. If only are nearly impossible to access. If only there was a really good open-source there was a really good open-source there was a really good open-source solution that would allow your agents to solution that would allow your agents to solution that would allow your agents to access data in formats they understand access data in formats they understand access data in formats they understand from all over the web. Oh, is it on the from all over the web. Oh, is it on the from all over the web. Oh, is it on the screen? Yeah, wire crawl is that. They screen? Yeah, wire crawl is that. They screen? Yeah, wire crawl is that. They built all the open source tooling your built all the open source tooling your built all the open source tooling your agents need to get data out of the web. agents need to get data out of the web. agents need to get data out of the web. And they also offer a super generous And they also offer a super generous And they also offer a super generous hosting platform for doing it yourself. hosting platform for doing it yourself. hosting platform for doing it yourself. I know most of us aren't reading outputs I know most of us aren't reading outputs I know most of us aren't reading outputs from APIs anymore, but they really are from APIs anymore, but they really are from APIs anymore, but they really are this simple. You put in a URL, and what this simple. You put in a URL, and what this simple. You put in a URL, and what you get back is a markdown version of you get back is a markdown version of you get back is a markdown version of the page, a JSON version of the page, as the page, a JSON version of the page, as the page, a JSON version of the page, as well as a screenshot of the pages well as a screenshot of the pages well as a screenshot of the pages content. They have SDKs for pretty much content. They have SDKs for pretty much content. They have SDKs for pretty much everything you're already using. And if everything you're already using. And if everything you're already using. And if not, you can always curl or use their not, you can always curl or use their not, you can always curl or use their CLI, which honestly, the CLI has really, CLI, which honestly, the CLI has really, CLI, which honestly, the CLI has really, really impressed me. Once you paste an really impressed me. Once you paste an really impressed me. Once you paste an API key, you can just tell your agents API key, you can just tell your agents API key, you can just tell your agents about it, and it'll give them access to about it, and it'll give them access to about it, and it'll give them access to info they might not have otherwise had info they might not have otherwise had info they might not have otherwise had access to. And the result is super access to. And the result is super access to. And the result is super simple markdown that your agents can simple markdown that your agents can simple markdown that your agents can actually understand and use for things.

  3. actually understand and use for things. actually understand and use for things. If all they offered was this direct URL If all they offered was this direct URL If all they offered was this direct URL scraping, it would already be one of the scraping, it would already be one of the scraping, it would already be one of the most useful things in my toolbox. But most useful things in my toolbox. But most useful things in my toolbox. But they offer way more. They have general they offer way more. They have general they offer way more. They have general search for finding things on the web search for finding things on the web search for finding things on the web that you might not already have a URL that you might not already have a URL that you might not already have a URL for, as well as interaction for when you for, as well as interaction for when you for, as well as interaction for when you want to actually do things on a page. My want to actually do things on a page. My want to actually do things on a page. My favorite though is the monitoring favorite though is the monitoring favorite though is the monitoring feature. Monitor lets you get notified feature. Monitor lets you get notified feature. Monitor lets you get notified when something changes on a page, like when something changes on a page, like when something changes on a page, like maybe a price goes down or a release maybe a price goes down or a release maybe a price goes down or a release happens or an early access thing comes happens or an early access thing comes happens or an early access thing comes through. It's so useful for you and your through. It's so useful for you and your through. It's so useful for you and your agents to know when things change on the agents to know when things change on the agents to know when things change on the web. And as great as the API is, not web. And as great as the API is, not web. And as great as the API is, not every tool can use it, but you know every tool can use it, but you know every tool can use it, but you know what? They can use an MCP server. Do you what? They can use an MCP server. Do you what? They can use an MCP server. Do you know how useful this is for Chat GPT and know how useful this is for Chat GPT and know how useful this is for Chat GPT and Codeex as well as every other tool ever. Codeex as well as every other tool ever. Codeex as well as every other tool ever. It's incredible. Two more really quick It's incredible. Two more really quick It's incredible. Two more really quick things they wanted me to make sure you things they wanted me to make sure you things they wanted me to make sure you know. The first is that you can use this know. The first is that you can use this know. The first is that you can use this without even having an API key. It's without even having an API key. It's without even having an API key. It's just free. You'll hit limits eventually, just free. You'll hit limits eventually, just free. You'll hit limits eventually, at which point you can go add an API at which point you can go add an API at which point you can go add an API key, but like it's insane how generous key, but like it's insane how generous key, but like it's insane how generous all of these tiers are. You get a ton all of these tiers are. You get a ton all of these tiers are. You get a ton for no money. And if you do need some for no money. And if you do need some for no money. And if you do need some money, they can help with that, too, money, they can help with that, too, money, they can help with that, too, because they're currently hiring. So, if because they're currently hiring. So, if because they're currently hiring. So, if this is interesting, hit them up. Give this is interesting, hit them up. Give this is interesting, hit them up. Give the whole web tier your agents at the whole web tier your agents at the whole web tier your agents at soy.firecrawl. soy.firecrawl. soy.firecrawl. First and foremost, let's start with First and foremost, let's start with First and foremost, let's start with what is ultra. I'll start in the what is ultra. I'll start in the what is ultra. I'll start in the terminal so I can show you when you terminal so I can show you when you terminal so I can show you when you select the model that Altra is presented select the model that Altra is presented select the model that Altra is presented here as a reasoning level. You have low, here as a reasoning level. You have low, here as a reasoning level. You have low, medium, high, extra high, max, and medium, high, extra high, max, and medium, high, extra high, max, and ultra. If you use the chat GBT app, ultra. If you use the chat GBT app, ultra. If you use the chat GBT app, formerly known as the codeex app, they formerly known as the codeex app, they formerly known as the codeex app, they did an interesting thing here where it did an interesting thing here where it did an interesting thing here where it starts at the bottom as terraite, then starts at the bottom as terraite, then starts at the bottom as terraite, then it goes to soul light, then soul medium, it goes to soul light, then soul medium, it goes to soul light, then soul medium, soul high, soul extra high. This is a soul high, soul extra high. This is a soul high, soul extra high. This is a change. It was going to ultra before and change. It was going to ultra before and change. It was going to ultra before and I think they actually listened to me and I think they actually listened to me and I think they actually listened to me and they hid ultra from that slider. I they hid ultra from that slider. I they hid ultra from that slider. I shouldn't have installed the update I

  4. shouldn't have installed the update I shouldn't have installed the update I just installed because before I just installed because before I just installed because before I installed the update, ultra was at the installed the update, ultra was at the installed the update, ultra was at the end of this list. I have told OpenAI end of this list. I have told OpenAI end of this list. I have told OpenAI verbatim multiple times now that Ultra verbatim multiple times now that Ultra verbatim multiple times now that Ultra should never have been included the way should never have been included the way should never have been included the way it was in this selector. It even had a it was in this selector. It even had a it was in this selector. It even had a fancy purple gradient it would do when fancy purple gradient it would do when fancy purple gradient it would do when you did it before. I'm thankful they you did it before. I'm thankful they you did it before. I'm thankful they killed that cuz that was the biggest killed that cuz that was the biggest killed that cuz that was the biggest mistake. Why is it a mistake? It's cuz mistake. Why is it a mistake? It's cuz mistake. Why is it a mistake? It's cuz Ultra isn't a reasoning level. I'll Ultra isn't a reasoning level. I'll Ultra isn't a reasoning level. I'll explain the easiest way I know how to, explain the easiest way I know how to, explain the easiest way I know how to, which is with Claude code. Not because which is with Claude code. Not because which is with Claude code. Not because I'm going to ask it to explain, but I I'm going to ask it to explain, but I I'm going to ask it to explain, but I want to show you what happens when you want to show you what happens when you want to show you what happens when you do slasheffort. You have options here. do slasheffort. You have options here. do slasheffort. You have options here. Low, medium, high, XH high, max, and Low, medium, high, XH high, max, and Low, medium, high, XH high, max, and ultra code. Huh? Ultra code. Ultra Ultra ultra code. Huh? Ultra code. Ultra Ultra ultra code. Huh? Ultra code. Ultra Ultra Code. I wonder if there's some overlap Code. I wonder if there's some overlap Code. I wonder if there's some overlap here. Ultra code was a pretty cool here. Ultra code was a pretty cool here. Ultra code was a pretty cool feature in Claude Code, but you'll feature in Claude Code, but you'll feature in Claude Code, but you'll notice underneath here, Ultraode says notice underneath here, Ultraode says notice underneath here, Ultraode says XHigh plus workflows. Hm. If you haven't XHigh plus workflows. Hm. If you haven't XHigh plus workflows. Hm. If you haven't been keeping up with my recent procla been keeping up with my recent procla been keeping up with my recent procla code arc, this is the big part of why code arc, this is the big part of why code arc, this is the big part of why it's workflows. Ultra code isn't it's workflows. Ultra code isn't it's workflows. Ultra code isn't workflows, but it's a way to trigger workflows, but it's a way to trigger workflows, but it's a way to trigger them. The more important piece to look them. The more important piece to look them. The more important piece to look at here is the other thing it says at here is the other thing it says at here is the other thing it says underneath, which is X high. Your underneath, which is X high. Your underneath, which is X high. Your options are high, XH high, max with the options are high, XH high, max with the options are high, XH high, max with the fancy little like gradient on the text fancy little like gradient on the text fancy little like gradient on the text there, but then ultra code does the there, but then ultra code does the there, but then ultra code does the super fancy gradient. Silly, but watch super fancy gradient. Silly, but watch super fancy gradient. Silly, but watch at the top here. I currently have Fable at the top here. I currently have Fable at the top here. I currently have Fable 5 with high effort selected. So, when I 5 with high effort selected. So, when I 5 with high effort selected. So, when I press enter here, it's going to change press enter here, it's going to change press enter here, it's going to change that to ultra, right? Nope. It changes that to ultra, right? Nope. It changes that to ultra, right? Nope. It changes it to Fable 5 with X higheffort. And it to Fable 5 with X higheffort. And it to Fable 5 with X higheffort. And ultra code is just a state in the ultra code is just a state in the ultra code is just a state in the bottom. The reason for that is that

  5. bottom. The reason for that is that bottom. The reason for that is that ultra code isn't a reasoning level. ultra code isn't a reasoning level. ultra code isn't a reasoning level. Ultra code is a skill and ultra is two. Ultra code is a skill and ultra is two. Ultra code is a skill and ultra is two. What do I mean by that? What I mean is What do I mean by that? What I mean is What do I mean by that? What I mean is that it's the same as when you do a that it's the same as when you do a that it's the same as when you do a slash command to pull in a skill where slash command to pull in a skill where slash command to pull in a skill where it takes the content of that markdown it takes the content of that markdown it takes the content of that markdown file and it appends it to the system file and it appends it to the system file and it appends it to the system prompt so it's there when the next prompt so it's there when the next prompt so it's there when the next message comes through. The reason they message comes through. The reason they message comes through. The reason they do this is they want to make it easier do this is they want to make it easier do this is they want to make it easier to trigger sub agents. Ultra and Ultra to trigger sub agents. Ultra and Ultra to trigger sub agents. Ultra and Ultra Code are both instructions for the model Code are both instructions for the model Code are both instructions for the model to use way more sub aents. In Cloud to use way more sub aents. In Cloud to use way more sub aents. In Cloud Code, this works through workflows and Code, this works through workflows and Code, this works through workflows and in Codeex, it works through the agent in Codeex, it works through the agent in Codeex, it works through the agent implementations called V1 and V2, which implementations called V1 and V2, which implementations called V1 and V2, which I will crash out about momentarily. I will crash out about momentarily. I will crash out about momentarily. Don't worry. I just wanted to make sure Don't worry. I just wanted to make sure Don't worry. I just wanted to make sure that we understood this first and that we understood this first and that we understood this first and foremost. Ultra isn't a reasoning level. foremost. Ultra isn't a reasoning level. foremost. Ultra isn't a reasoning level. It's effectively a toggle that turns on It's effectively a toggle that turns on It's effectively a toggle that turns on a change in your system prompt, telling a change in your system prompt, telling a change in your system prompt, telling the model to do more sub agents. As much the model to do more sub agents. As much the model to do more sub agents. As much as I love to give anthropic they as I love to give anthropic they as I love to give anthropic they did mostly get this right. It's a little did mostly get this right. It's a little did mostly get this right. It's a little too easy to leave Ultra Code on, which too easy to leave Ultra Code on, which too easy to leave Ultra Code on, which isn't great. But as an implementation isn't great. But as an implementation isn't great. But as an implementation with the system prompt change and with the system prompt change and with the system prompt change and specifically the default to X specifically the default to X specifically the default to X higheffort, not to max, when you go to higheffort, not to max, when you go to higheffort, not to max, when you go to ultra, the reasoning level is lower than ultra, the reasoning level is lower than ultra, the reasoning level is lower than it was on max. Again, when I go to it was on max. Again, when I go to it was on max. Again, when I go to slasheffort, low, medium, high, XH high, slasheffort, low, medium, high, XH high, slasheffort, low, medium, high, XH high, max, ultra. And the thing that's max, ultra. And the thing that's max, ultra. And the thing that's unintuitive is that Ultra Code actually unintuitive is that Ultra Code actually unintuitive is that Ultra Code actually points at X high, not at max. This is points at X high, not at max. This is points at X high, not at max. This is not what Codeex does. Sorry, this is not not what Codeex does. Sorry, this is not not what Codeex does. Sorry, this is not what Chat GBT does. I also can't help what Chat GBT does. I also can't help what Chat GBT does. I also can't help but notice they are hiding the max but notice they are hiding the max but notice they are hiding the max option now. They're listening to me.

  6. option now. They're listening to me. option now. They're listening to me. It's fun seeing in real time OpenAI It's fun seeing in real time OpenAI It's fun seeing in real time OpenAI making changes based on the things that making changes based on the things that making changes based on the things that I'm shouting about on the internet. But I'm shouting about on the internet. But I'm shouting about on the internet. But Ultra is still hidden here under effort Ultra is still hidden here under effort Ultra is still hidden here under effort where it never should have been in the where it never should have been in the where it never should have been in the first place. When we click it, you'll first place. When we click it, you'll first place. When we click it, you'll see all of these other changes under the see all of these other changes under the see all of these other changes under the hood. Ultra is using max reasoning hood. Ultra is using max reasoning hood. Ultra is using max reasoning levels, which is absurd because max levels, which is absurd because max levels, which is absurd because max reasoning burns tokens. I talked about reasoning burns tokens. I talked about reasoning burns tokens. I talked about this a bit in my previous cost savings this a bit in my previous cost savings this a bit in my previous cost savings video, but the TLDDR is that max does up video, but the TLDDR is that max does up video, but the TLDDR is that max does up to two times more token burn than X to two times more token burn than X to two times more token burn than X higher high. Depending on the bench, higher high. Depending on the bench, higher high. Depending on the bench, it's between 4% and 10% better at more it's between 4% and 10% better at more it's between 4% and 10% better at more than 2x the cost. So, I don't think max than 2x the cost. So, I don't think max than 2x the cost. So, I don't think max makes sense almost ever. But ultra isn't makes sense almost ever. But ultra isn't makes sense almost ever. But ultra isn't max. Ultra is nearly infinitely max. Ultra is nearly infinitely max. Ultra is nearly infinitely recursive maxes, which means it burns recursive maxes, which means it burns recursive maxes, which means it burns tokens. This is the primary reason Ultra tokens. This is the primary reason Ultra tokens. This is the primary reason Ultra is bad. The problem is that the parent is bad. The problem is that the parent is bad. The problem is that the parent agent and all of the sub aents are set agent and all of the sub aents are set agent and all of the sub aents are set to max reasoning. Well, Theo, this is an to max reasoning. Well, Theo, this is an to max reasoning. Well, Theo, this is an easy fix. Just tell the parent agent, easy fix. Just tell the parent agent, easy fix. Just tell the parent agent, the top level, to spin up the sub aents the top level, to spin up the sub aents the top level, to spin up the sub aents on a lower reasoning level. I'll have on a lower reasoning level. I'll have on a lower reasoning level. I'll have more to say on this in a bit, but the more to say on this in a bit, but the more to say on this in a bit, but the TLDDR for now is that the version of sub TLDDR for now is that the version of sub TLDDR for now is that the version of sub aents that soul uses does not currently aents that soul uses does not currently aents that soul uses does not currently allow for selection of reasoning levels.

  7. allow for selection of reasoning levels. allow for selection of reasoning levels. So, if you have ultra set, all of the So, if you have ultra set, all of the So, if you have ultra set, all of the children sub agents spawned are also children sub agents spawned are also children sub agents spawned are also going to be ultra. I crashed out pretty going to be ultra. I crashed out pretty going to be ultra. I crashed out pretty hard about this on Twitter going as far hard about this on Twitter going as far hard about this on Twitter going as far as to say that claude code is far ahead. as to say that claude code is far ahead. as to say that claude code is far ahead. At the very least, they should give me At the very least, they should give me At the very least, they should give me an interim solution where I can hard set an interim solution where I can hard set an interim solution where I can hard set sub aent reasoning levels. The current sub aent reasoning levels. The current sub aent reasoning levels. The current implementation is so bad that I think it implementation is so bad that I think it implementation is so bad that I think it has caused this perception shift where has caused this perception shift where has caused this perception shift where people think that the new models are way people think that the new models are way people think that the new models are way less efficient and will destroy your less efficient and will destroy your less efficient and will destroy your rate limits. The new models won't do rate limits. The new models won't do rate limits. The new models won't do that. At least with a little bit of that. At least with a little bit of that. At least with a little bit of careful use, they won't do that. Ultra careful use, they won't do that. Ultra careful use, they won't do that. Ultra absolutely will. I blew my first 5-hour absolutely will. I blew my first 5-hour absolutely will. I blew my first 5-hour limit in 20 minutes using ultra on fast limit in 20 minutes using ultra on fast limit in 20 minutes using ultra on fast mode cuz I didn't know how much crazier mode cuz I didn't know how much crazier mode cuz I didn't know how much crazier it would be. Instantaneously evaporated it would be. Instantaneously evaporated it would be. Instantaneously evaporated my limit. I burned a manual reset and I my limit. I burned a manual reset and I my limit. I burned a manual reset and I evaporated again another 40 minutes. I evaporated again another 40 minutes. I evaporated again another 40 minutes. I hit the 5-hour limit twice in under an hit the 5-hour limit twice in under an hit the 5-hour limit twice in under an hour. All because I had one ultra run hour. All because I had one ultra run hour. All because I had one ultra run going. Hitting your 5-hour limit in 20 going. Hitting your 5-hour limit in 20 going. Hitting your 5-hour limit in 20 minutes is bad, but what would be much minutes is bad, but what would be much minutes is bad, but what would be much worse is hitting your weekly limit in an worse is hitting your weekly limit in an worse is hitting your weekly limit in an hour and a half, which is dangerously hour and a half, which is dangerously hour and a half, which is dangerously possible right now because OpenAI is possible right now because OpenAI is possible right now because OpenAI is temporarily removing the 5-hour limit temporarily removing the 5-hour limit temporarily removing the 5-hour limit due to people hitting it too fast. This due to people hitting it too fast. This due to people hitting it too fast. This is a good change overall, but if you're is a good change overall, but if you're is a good change overall, but if you're using Ultra, it's a little scary because using Ultra, it's a little scary because using Ultra, it's a little scary because previously that 5-hour limit being hit previously that 5-hour limit being hit previously that 5-hour limit being hit is almost a warning, like a hey, be is almost a warning, like a hey, be is almost a warning, like a hey, be careful. you're going to overuse this careful. you're going to overuse this careful. you're going to overuse this and lose your weekly limit. And since and lose your weekly limit. And since and lose your weekly limit. And since the five hour is like 20 to 25% of your the five hour is like 20 to 25% of your the five hour is like 20 to 25% of your weekly, you would only lose 20 to 25% of weekly, you would only lose 20 to 25% of weekly, you would only lose 20 to 25% of your weekly usage. With this change, it your weekly usage. With this change, it your weekly usage. With this change, it is very much possible that Ultra could is very much possible that Ultra could is very much possible that Ultra could nuke your whole week. So, be careful.

  8. nuke your whole week. So, be careful. nuke your whole week. So, be careful. I'm realizing now that the order of I'm realizing now that the order of I'm realizing now that the order of events I planned here isn't actually events I planned here isn't actually events I planned here isn't actually going to make a lot of sense. So, I'm going to make a lot of sense. So, I'm going to make a lot of sense. So, I'm going to move how it can be fixed to be going to move how it can be fixed to be going to move how it can be fixed to be after my rant about the current state of after my rant about the current state of after my rant about the current state of Codeex. Oh boy. Codeex. Oh boy. Codeex. Oh boy. So as I said the sub aents version being So as I said the sub aents version being So as I said the sub aents version being used doesn't expose the necessary used doesn't expose the necessary used doesn't expose the necessary functionality. What do I mean by that? functionality. What do I mean by that? functionality. What do I mean by that? Well what I mean is that there are two Well what I mean is that there are two Well what I mean is that there are two versions of sub aents in codeex right versions of sub aents in codeex right versions of sub aents in codeex right now labeled v1 and v2 internally. The now labeled v1 and v2 internally. The now labeled v1 and v2 internally. The codec cli which is what powers the codec cli which is what powers the codec cli which is what powers the codeex desktop app is open source. So we codeex desktop app is open source. So we codeex desktop app is open source. So we can actually look through the code here. can actually look through the code here. can actually look through the code here. I have 56 soul in codec and claude code I have 56 soul in codec and claude code I have 56 soul in codec and claude code breaking down how the sub aent breaking down how the sub aent breaking down how the sub aent implementations work. so that we can implementations work. so that we can implementations work. so that we can compare based on their findings. Codex compare based on their findings. Codex compare based on their findings. Codex finished first. I'll look at cloud codes finished first. I'll look at cloud codes finished first. I'll look at cloud codes for the difference in a minute. I had for the difference in a minute. I had for the difference in a minute. I had them both going through the actual code them both going through the actual code them both going through the actual code base to try and break up the difference base to try and break up the difference base to try and break up the difference between these implementations. Obviously between these implementations. Obviously between these implementations. Obviously LLM speak so we'll get through and I LLM speak so we'll get through and I LLM speak so we'll get through and I will do my best to clean it up. V1 is will do my best to clean it up. V1 is will do my best to clean it up. V1 is like a dispatcher hiring temporary like a dispatcher hiring temporary like a dispatcher hiring temporary helpers by ticket number. V2 is like a helpers by ticket number. V2 is like a helpers by ticket number. V2 is like a named project team with an org chart and named project team with an org chart and named project team with an org chart and mailboxes.

  9. mailboxes. mailboxes. Good enough. V1 is very much a simple Good enough. V1 is very much a simple Good enough. V1 is very much a simple tool call that is a way the model can tool call that is a way the model can tool call that is a way the model can trigger another submodel to go do trigger another submodel to go do trigger another submodel to go do something. So let's say you tell the something. So let's say you tell the something. So let's say you tell the model to go review three PRs. It can model to go review three PRs. It can model to go review three PRs. It can spin up three sub aents, one for each PR spin up three sub aents, one for each PR spin up three sub aents, one for each PR with the instructions to go do that one with the instructions to go do that one with the instructions to go do that one thing and then when it's they're done, thing and then when it's they're done, thing and then when it's they're done, they'll all send up their findings and they'll all send up their findings and they'll all send up their findings and that top level agent will decide what to that top level agent will decide what to that top level agent will decide what to do from there. Side note, this page is do from there. Side note, this page is do from there. Side note, this page is ugly as sin. I have things to say about ugly as sin. I have things to say about ugly as sin. I have things to say about that soon, don't worry. But it starts that soon, don't worry. But it starts that soon, don't worry. But it starts here with five definitions worth here with five definitions worth here with five definitions worth knowing. They refer to the parent agent knowing. They refer to the parent agent knowing. They refer to the parent agent as the root agent. I think that's fine. as the root agent. I think that's fine. as the root agent. I think that's fine. Then there is the sub agents which are Then there is the sub agents which are Then there is the sub agents which are the things spawned by that root agent. the things spawned by that root agent. the things spawned by that root agent. Sub aents can also spawn themselves. Sub aents can also spawn themselves. Sub aents can also spawn themselves. We'll have to deal with that in a bit. We'll have to deal with that in a bit. We'll have to deal with that in a bit. Then there is the context which is all Then there is the context which is all Then there is the context which is all of the info that the top level agent and of the info that the top level agent and of the info that the top level agent and the sub aents have. Usually sub agents the sub aents have. Usually sub agents the sub aents have. Usually sub agents are spawned by writing a new prompt or are spawned by writing a new prompt or are spawned by writing a new prompt or summarizing the stuff going on in the summarizing the stuff going on in the summarizing the stuff going on in the main thread so that the sub aent has a main thread so that the sub aent has a main thread so that the sub aent has a limited set of context. Remember that limited set of context. Remember that limited set of context. Remember that because there are things going on here. because there are things going on here. because there are things going on here. The mailbox is a new concept in the V2 The mailbox is a new concept in the V2 The mailbox is a new concept in the V2 implementation where there are messages implementation where there are messages implementation where there are messages for the team as well as the ability to for the team as well as the ability to for the team as well as the ability to send messages between different agents.

  10. send messages between different agents. send messages between different agents. And then there is slots which is the And then there is slots which is the And then there is slots which is the number of agents that can be running and number of agents that can be running and number of agents that can be running and the shared workspace which is the files the shared workspace which is the files the shared workspace which is the files the agents are working on. This does not the agents are working on. This does not the agents are working on. This does not need to be in here. This is not a great need to be in here. This is not a great need to be in here. This is not a great description. What a surprise. The very description. What a surprise. The very description. What a surprise. The very least it made decent diagrams here with least it made decent diagrams here with least it made decent diagrams here with V1. you have that root level agent and V1. you have that root level agent and V1. you have that root level agent and it can spawn sub aents to do specific it can spawn sub aents to do specific it can spawn sub aents to do specific things that return a result when they're things that return a result when they're things that return a result when they're done. V2 is a complete overhaul where done. V2 is a complete overhaul where done. V2 is a complete overhaul where the different sub aents can talk to each the different sub aents can talk to each the different sub aents can talk to each other and spawn sub aents of their own. other and spawn sub aents of their own. other and spawn sub aents of their own. This sound familiar? Remember that This sound familiar? Remember that This sound familiar? Remember that infinite recursion thing I was talking infinite recursion thing I was talking infinite recursion thing I was talking about with Ultra? Yeah. There's a about with Ultra? Yeah. There's a about with Ultra? Yeah. There's a problem though. V1 is the finished problem though. V1 is the finished problem though. V1 is the finished implementation. This is what Codeex uses implementation. This is what Codeex uses implementation. This is what Codeex uses by default. And V2 is an overhaul that by default. And V2 is an overhaul that by default. And V2 is an overhaul that they are still working on. It's still a they are still working on. It's still a they are still working on. It's still a work in progress. V2 is very much work in progress. V2 is very much work in progress. V2 is very much unfinished and if you turn it on, you unfinished and if you turn it on, you unfinished and if you turn it on, you get errors for having V1 at all. get errors for having V1 at all. get errors for having V1 at all. Especially if you have anything changed Especially if you have anything changed Especially if you have anything changed about V1 like custom limits, custom about V1 like custom limits, custom about V1 like custom limits, custom instructions, whatnot. You can't have V1 instructions, whatnot. You can't have V1 instructions, whatnot. You can't have V1 and V2 on at the same time until now and V2 on at the same time until now and V2 on at the same time until now because it's a really really annoying because it's a really really annoying because it's a really really annoying exception. In the models cache JSON exception. In the models cache JSON exception. In the models cache JSON file, which is the file that Codeex uses file, which is the file that Codeex uses file, which is the file that Codeex uses to determine which models are available to determine which models are available to determine which models are available to you, there's a new field they added to you, there's a new field they added to you, there's a new field they added multi- aent version, which is V2 for multi- aent version, which is V2 for multi- aent version, which is V2 for Soul and Terra and V1 for everything Soul and Terra and V1 for everything Soul and Terra and V1 for everything else. So now it doesn't matter what you else. So now it doesn't matter what you else. So now it doesn't matter what you have configured or if you've manually have configured or if you've manually have configured or if you've manually opted into V2, the new models are always opted into V2, the new models are always opted into V2, the new models are always routing to V2. But V2 has some other routing to V2. But V2 has some other routing to V2. But V2 has some other quirks. With V1, Codex had access to quirks. With V1, Codex had access to quirks. With V1, Codex had access to this handful of tools for spawning,

  11. this handful of tools for spawning, this handful of tools for spawning, sending, waiting, closing, and resuming sending, waiting, closing, and resuming sending, waiting, closing, and resuming sub agents. They would create a separate sub agents. They would create a separate sub agents. They would create a separate thread effectively with its own context thread effectively with its own context thread effectively with its own context for the agent to go do its thing and for the agent to go do its thing and for the agent to go do its thing and then respond when it's done. The parent then respond when it's done. The parent then respond when it's done. The parent would wait for the agents to be complete would wait for the agents to be complete would wait for the agents to be complete and then summarize whatever it got back. and then summarize whatever it got back. and then summarize whatever it got back. V2 is quite different because now it is V2 is quite different because now it is V2 is quite different because now it is breaking down tasks instead of just sub breaking down tasks instead of just sub breaking down tasks instead of just sub agents. An agent requires a task name. A agents. An agent requires a task name. A agents. An agent requires a task name. A root child named research becomes root child named research becomes root child named research becomes /root/ressearch. /root/ressearch. /root/ressearch. its tests become slrootressearchests. its tests become slrootressearchests. its tests become slrootressearchests. So this is a pathbased naming scheme for So this is a pathbased naming scheme for So this is a pathbased naming scheme for agents to spawn sub aent layers. But the agents to spawn sub aent layers. But the agents to spawn sub aent layers. But the biggest issue for me is the way it biggest issue for me is the way it biggest issue for me is the way it handles context sharing. By default, all handles context sharing. By default, all handles context sharing. By default, all of the history in your main thread is of the history in your main thread is of the history in your main thread is going to be shared to all of the sub going to be shared to all of the sub going to be shared to all of the sub aents. Spoiler, this is really aents. Spoiler, this is really aents. Spoiler, this is really stupid. While 56 is way better at not stupid. While 56 is way better at not stupid. While 56 is way better at not falling for context pollution where some falling for context pollution where some falling for context pollution where some bad thing in the history affects how it bad thing in the history affects how it bad thing in the history affects how it behaves, it still does. This also is a behaves, it still does. This also is a behaves, it still does. This also is a massive increase in cost because the sub massive increase in cost because the sub massive increase in cost because the sub aents now have way more context. I would aents now have way more context. I would aents now have way more context. I would assume that part of why they did this is assume that part of why they did this is assume that part of why they did this is to keep the cache consistent between the to keep the cache consistent between the to keep the cache consistent between the parent agent and the sub aents, but I'm parent agent and the sub aents, but I'm parent agent and the sub aents, but I'm also pretty sure there are changes in also pretty sure there are changes in also pretty sure there are changes in the system prompt which is slightly the system prompt which is slightly the system prompt which is slightly higher, therefore busting the cache. Not higher, therefore busting the cache. Not higher, therefore busting the cache. Not positive about that, but I would be positive about that, but I would be positive about that, but I would be surprised. Either way, getting one cache surprised. Either way, getting one cache surprised. Either way, getting one cache right at the start of a new sub agent right at the start of a new sub agent right at the start of a new sub agent thread is far from a high cost. I think thread is far from a high cost. I think thread is far from a high cost. I think it's stupid to try and preserve cash it's stupid to try and preserve cash it's stupid to try and preserve cash between the parent or sorry, root agent between the parent or sorry, root agent between the parent or sorry, root agent and the sub agents. So, yeah, dumb. I and the sub agents. So, yeah, dumb. I and the sub agents. So, yeah, dumb. I don't like this at all. As I said, by don't like this at all. As I said, by don't like this at all. As I said, by default, it now takes all of your turns

  12. default, it now takes all of your turns default, it now takes all of your turns in all of the history when it goes to in all of the history when it goes to in all of the history when it goes to that sub agent. But you can set that sub agent. But you can set that sub agent. But you can set different amounts to share, like none or different amounts to share, like none or different amounts to share, like none or a number like three. And as I said, the a number like three. And as I said, the a number like three. And as I said, the default is the full history, which I default is the full history, which I default is the full history, which I think is really dumb. It also filters think is really dumb. It also filters think is really dumb. It also filters out the tool calls. So that breaks cash. out the tool calls. So that breaks cash. out the tool calls. So that breaks cash. Where things start to get really weird Where things start to get really weird Where things start to get really weird is this idea of mailboxes. Instead of is this idea of mailboxes. Instead of is this idea of mailboxes. Instead of just having a sub agent go do some work just having a sub agent go do some work just having a sub agent go do some work and then respond when it's done, they and then respond when it's done, they and then respond when it's done, they now can send typed messages between each now can send typed messages between each now can send typed messages between each other. Send message cues a note. other. Send message cues a note. other. Send message cues a note. Follow-up task gives an idle helper more Follow-up task gives an idle helper more Follow-up task gives an idle helper more work and can start its next turn. This work and can start its next turn. This work and can start its next turn. This allows for the parent agent to send allows for the parent agent to send allows for the parent agent to send additional work down to an existing sub additional work down to an existing sub additional work down to an existing sub agent when it decides it wants more agent when it decides it wants more agent when it decides it wants more done. This is one of those things where done. This is one of those things where done. This is one of those things where it like does genuinely sound really it like does genuinely sound really it like does genuinely sound really really cool. In reality, mixed results. really cool. In reality, mixed results. really cool. In reality, mixed results. The finished work is routed to the The finished work is routed to the The finished work is routed to the direct parent. So again, if they're direct parent. So again, if they're direct parent. So again, if they're nested, tests sends its results to nested, tests sends its results to nested, tests sends its results to research, which sends its results to research, which sends its results to research, which sends its results to root. Yeah. And waiting no longer means root. Yeah. And waiting no longer means root. Yeah. And waiting no longer means wait till it's done. It means wait till wait till it's done. It means wait till wait till it's done. It means wait till we get a message back. The root agent we get a message back. The root agent we get a message back. The root agent can also list all of the sub aents and can also list all of the sub aents and can also list all of the sub aents and even interrupt them if for some reason even interrupt them if for some reason even interrupt them if for some reason it wants to. If it got info from another it wants to. If it got info from another it wants to. If it got info from another sub agent, it's like, "Oh, that thing sub agent, it's like, "Oh, that thing sub agent, it's like, "Oh, that thing probably doesn't matter anymore. I can probably doesn't matter anymore. I can probably doesn't matter anymore. I can kill that." There is no depth limit in kill that." There is no depth limit in kill that." There is no depth limit in V2, so it could spawn infinitely nested V2, so it could spawn infinitely nested V2, so it could spawn infinitely nested sub aents, but it can only run four at a sub aents, but it can only run four at a sub aents, but it can only run four at a time by default, which I had changed, time by default, which I had changed, time by default, which I had changed, which is a big part of how I hit crazy which is a big part of how I hit crazy which is a big part of how I hit crazy usage so quickly. Random observation, I usage so quickly. Random observation, I usage so quickly. Random observation, I did run this prompt on both Codeex and did run this prompt on both Codeex and did run this prompt on both Codeex and Claude Code. Claude Code took a bit

  13. Claude Code. Claude Code took a bit Claude Code. Claude Code took a bit longer. Again, I was using 56 soul on longer. Again, I was using 56 soul on longer. Again, I was using 56 soul on both. If you feel like this design's a both. If you feel like this design's a both. If you feel like this design's a little bloated, you're not the only one. little bloated, you're not the only one. little bloated, you're not the only one. I love this message from JKF here. It I love this message from JKF here. It I love this message from JKF here. It sounds like too many people designed sub sounds like too many people designed sub sounds like too many people designed sub agents v2 and caused it to be bloated. I agents v2 and caused it to be bloated. I agents v2 and caused it to be bloated. I agree. I don't really like the design of agree. I don't really like the design of agree. I don't really like the design of V2, and I didn't use it much when I was V2, and I didn't use it much when I was V2, and I didn't use it much when I was testing 56 soul because again, they testing 56 soul because again, they testing 56 soul because again, they hadn't added this change where it would hadn't added this change where it would hadn't added this change where it would default route to V2. I did turn it on default route to V2. I did turn it on default route to V2. I did turn it on manually near the end of my testing out manually near the end of my testing out manually near the end of my testing out of curiosity and I noticed it would work of curiosity and I noticed it would work of curiosity and I noticed it would work a lot longer. Didn't get far enough on a lot longer. Didn't get far enough on a lot longer. Didn't get far enough on real work to really see a difference in real work to really see a difference in real work to really see a difference in the quality of the output, but it the quality of the output, but it the quality of the output, but it definitely goes a lot longer and spins definitely goes a lot longer and spins definitely goes a lot longer and spins up a lot more. Don't necessarily love up a lot more. Don't necessarily love up a lot more. Don't necessarily love that, but it can. All of that said, I've that, but it can. All of that said, I've that, but it can. All of that said, I've been using it more since and I'm not been using it more since and I'm not been using it more since and I'm not really impressed with the quality of the really impressed with the quality of the really impressed with the quality of the work and output I get from V2 with the work and output I get from V2 with the work and output I get from V2 with the current settings and setup in codecs. current settings and setup in codecs. current settings and setup in codecs. While I do love the ability to spawn sub While I do love the ability to spawn sub While I do love the ability to spawn sub agents to break up work and control the agents to break up work and control the agents to break up work and control the context a bit more, the changes to how context a bit more, the changes to how context a bit more, the changes to how context is managed and the changes to context is managed and the changes to context is managed and the changes to how messages are passed makes this how messages are passed makes this how messages are passed makes this noisier is the simplest I can put it.

  14. noisier is the simplest I can put it. noisier is the simplest I can put it. And it also burns way, way, way more And it also burns way, way, way more And it also burns way, way, way more tokens as a result. And this leads us to tokens as a result. And this leads us to tokens as a result. And this leads us to our final question of how can it be our final question of how can it be our final question of how can it be fixed. I know this stuff scares a lot of fixed. I know this stuff scares a lot of fixed. I know this stuff scares a lot of people like they're afraid if they don't people like they're afraid if they don't people like they're afraid if they don't get it right initially, they might miss get it right initially, they might miss get it right initially, they might miss out on this opportunity. they might fall out on this opportunity. they might fall out on this opportunity. they might fall behind all these things and a lot of you behind all these things and a lot of you behind all these things and a lot of you guys are trying so hard to stay on top guys are trying so hard to stay on top guys are trying so hard to stay on top of the best solutions. If you're here of the best solutions. If you're here of the best solutions. If you're here out of fear and not out of excitement out of fear and not out of excitement out of fear and not out of excitement and curiosity, the only thing I want you and curiosity, the only thing I want you and curiosity, the only thing I want you to take from this is that you should to take from this is that you should to take from this is that you should just wait a bit. The current just wait a bit. The current just wait a bit. The current implementation is bad and has rough implementation is bad and has rough implementation is bad and has rough edges, but the OpenAI team is incredible edges, but the OpenAI team is incredible edges, but the OpenAI team is incredible at taking feedback. I have been shocked at taking feedback. I have been shocked at taking feedback. I have been shocked as I film at just how many of the things as I film at just how many of the things as I film at just how many of the things I'm complaining to them about have been I'm complaining to them about have been I'm complaining to them about have been addressed between the sixish hours ago addressed between the sixish hours ago addressed between the sixish hours ago where I sent the complaint on Slack and where I sent the complaint on Slack and where I sent the complaint on Slack and now when it's not there anymore. If you now when it's not there anymore. If you now when it's not there anymore. If you just wait a little bit, these things just wait a little bit, these things just wait a little bit, these things will improve. It continuing to use the will improve. It continuing to use the will improve. It continuing to use the defaults and trusting that these defaults and trusting that these defaults and trusting that these companies will get it figured out companies will get it figured out companies will get it figured out eventually, you'll be fine. It's not eventually, you'll be fine. It's not eventually, you'll be fine. It's not that big a deal. I know I am crashing that big a deal. I know I am crashing that big a deal. I know I am crashing out hard about this, but I want you to out hard about this, but I want you to out hard about this, but I want you to realize this is because I am a hardcore realize this is because I am a hardcore realize this is because I am a hardcore nerd who cares too much and is just nerd who cares too much and is just nerd who cares too much and is just digging into the details because they digging into the details because they digging into the details because they want to. You don't have to do this to want to. You don't have to do this to want to. You don't have to do this to stay ahead. You can just use the tools stay ahead. You can just use the tools stay ahead. You can just use the tools as they work by default and they are as they work by default and they are as they work by default and they are still really good. With all of that still really good. With all of that still really good. With all of that said, I think OpenAI needs to rethink said, I think OpenAI needs to rethink said, I think OpenAI needs to rethink how they do sub agents. As I mentioned how they do sub agents. As I mentioned how they do sub agents. As I mentioned before, ultra in codeex is a pretty before, ultra in codeex is a pretty before, ultra in codeex is a pretty blatant copy of ultra code in claude blatant copy of ultra code in claude blatant copy of ultra code in claude code, at least in how it is meant to code, at least in how it is meant to code, at least in how it is meant to work where it's added to the reasoning work where it's added to the reasoning work where it's added to the reasoning effort slider. It is meant to trigger effort slider. It is meant to trigger effort slider. It is meant to trigger sub aents and be a way to force sub

  15. sub aents and be a way to force sub sub aents and be a way to force sub aents on. It basically just appends to aents on. It basically just appends to aents on. It basically just appends to the system prompt. Please, please, the system prompt. Please, please, the system prompt. Please, please, please use sub aents. But in Ultra on please use sub aents. But in Ultra on please use sub aents. But in Ultra on codeex, it is doing that through these codeex, it is doing that through these codeex, it is doing that through these weird sub aent tools like we've seen weird sub aent tools like we've seen weird sub aent tools like we've seen many times before. In cloud code, it's many times before. In cloud code, it's many times before. In cloud code, it's using workflows. I've talked about using workflows. I've talked about using workflows. I've talked about workflows a bunch before in my Claude workflows a bunch before in my Claude workflows a bunch before in my Claude code is good video and I plan to talk code is good video and I plan to talk code is good video and I plan to talk about them even more in my Claudeex about them even more in my Claudeex about them even more in my Claudeex video where I show how I'm using 56 soul video where I show how I'm using 56 soul video where I show how I'm using 56 soul in cloud code directly, not as a sub in cloud code directly, not as a sub in cloud code directly, not as a sub aent or a thing that cloud code calls aent or a thing that cloud code calls aent or a thing that cloud code calls literally as the model inside of cloud literally as the model inside of cloud literally as the model inside of cloud code directly. It's way better than I code directly. It's way better than I code directly. It's way better than I thought it would be. The reason I'm thought it would be. The reason I'm thought it would be. The reason I'm doing this is workflows. To put it doing this is workflows. To put it doing this is workflows. To put it simply, workflows are a way to simply, workflows are a way to simply, workflows are a way to programmatically define a bunch of sub programmatically define a bunch of sub programmatically define a bunch of sub aent stuff. In cloud code, you can aent stuff. In cloud code, you can aent stuff. In cloud code, you can actually save the file it writes because actually save the file it writes because actually save the file it writes because it's making a JavaScript file to do this it's making a JavaScript file to do this it's making a JavaScript file to do this work. And that's by defining a meta, work. And that's by defining a meta, work. And that's by defining a meta, which is a name, a description, and which is a name, a description, and which is a name, a description, and different phases because it has stages different phases because it has stages different phases because it has stages throughout its work. It then has throughout its work. It then has throughout its work. It then has schemas, which are the typed outputs schemas, which are the typed outputs schemas, which are the typed outputs that each of the phases should have. It that each of the phases should have. It that each of the phases should have. It then writes a prompt, and this is then writes a prompt, and this is then writes a prompt, and this is programmatic. So if it wants to insert programmatic. So if it wants to insert programmatic. So if it wants to insert things programmatically, use a map or things programmatically, use a map or things programmatically, use a map or some type of loop to add context in and some type of loop to add context in and some type of loop to add context in and do different steps, it can because it is do different steps, it can because it is do different steps, it can because it is just code. Once this is all defined, it just code. Once this is all defined, it just code. Once this is all defined, it can call the different phases. Sorry to can call the different phases. Sorry to can call the different phases. Sorry to my editor phase as I'm talking about my editor phase as I'm talking about my editor phase as I'm talking about phases. I'm sure that won't be confusing phases. I'm sure that won't be confusing phases. I'm sure that won't be confusing at all. So it starts phase reviewing. It at all. So it starts phase reviewing. It at all. So it starts phase reviewing. It defines reviews as a bunch of parallel defines reviews as a bunch of parallel defines reviews as a bunch of parallel agents. We have the first one here, agents. We have the first one here, agents. We have the first one here, which is a hard-coded prompt. Your which is a hard-coded prompt. Your which is a hard-coded prompt. Your assigned perspective is 56 soul at high assigned perspective is 56 soul at high assigned perspective is 56 soul at high reasoning effort label review 56 soul

  16. reasoning effort label review 56 soul reasoning effort label review 56 soul phase reviewing model 56 soul effort phase reviewing model 56 soul effort phase reviewing model 56 soul effort high schema review schema agent type high schema review schema agent type high schema review schema agent type general purpose and we do the same here general purpose and we do the same here general purpose and we do the same here with another one that's using terra as with another one that's using terra as with another one that's using terra as well and then one here using fable cuz I well and then one here using fable cuz I well and then one here using fable cuz I asked it to use all three of these once asked it to use all three of these once asked it to use all three of these once this phase is completed because remember this phase is completed because remember this phase is completed because remember we awaited it up here I know crazy we're we awaited it up here I know crazy we're we awaited it up here I know crazy we're reading code in a theo video again it reading code in a theo video again it reading code in a theo video again it starts the synthesizing phase where it starts the synthesizing phase where it starts the synthesizing phase where it awaits another agent call where it is awaits another agent call where it is awaits another agent call where it is passed these review results s as a JSON passed these review results s as a JSON passed these review results s as a JSON blob that's been stringified. And now blob that's been stringified. And now blob that's been stringified. And now another agent using Fable 5 on high is another agent using Fable 5 on high is another agent using Fable 5 on high is going to process all of the data it was going to process all of the data it was going to process all of the data it was given and respond following the given and respond following the given and respond following the synthesis schema. And this is how you synthesis schema. And this is how you synthesis schema. And this is how you can programmatically use sub aents. All can programmatically use sub aents. All can programmatically use sub aents. All of the stages are defined ahead of time. of the stages are defined ahead of time. of the stages are defined ahead of time. And depending on what the outputs of a And depending on what the outputs of a And depending on what the outputs of a given stage are or a given agent run given stage are or a given agent run given stage are or a given agent run are, it can decide if it should put the are, it can decide if it should put the are, it can decide if it should put the results in another run in another sub results in another run in another sub results in another run in another sub aent in a different stage. This one is aent in a different stage. This one is aent in a different stage. This one is really, really simple because there was really, really simple because there was really, really simple because there was only two phases. There's the review only two phases. There's the review only two phases. There's the review phase and the synthesizing phase. I've phase and the synthesizing phase. I've phase and the synthesizing phase. I've had many more complex ones, though. I've had many more complex ones, though. I've had many more complex ones, though. I've had some that had 12 plus phases. Here's had some that had 12 plus phases. Here's had some that had 12 plus phases. Here's one that had five. It was research, one that had five. It was research, one that had five. It was research, verify, synthesize, critique, and verify, synthesize, critique, and verify, synthesize, critique, and finalize.

  17. finalize. finalize. Each of those got a schema for how the Each of those got a schema for how the Each of those got a schema for how the output should be formatted. It was given output should be formatted. It was given output should be formatted. It was given a common prompt to append to the top of a common prompt to append to the top of a common prompt to append to the top of all of them because all of them need to all of them because all of them need to all of them because all of them need to know what the directory is and a bit of know what the directory is and a bit of know what the directory is and a bit of other information. It created these other information. It created these other information. It created these different topics that are keys and different topics that are keys and different topics that are keys and prompts for different things that we'll prompts for different things that we'll prompts for different things that we'll be doing throughout. It then starts the be doing throughout. It then starts the be doing throughout. It then starts the research phase. It goes through and research phase. It goes through and research phase. It goes through and defines multiple sub aents to do all of defines multiple sub aents to do all of defines multiple sub aents to do all of that. It clears out any that didn't have that. It clears out any that didn't have that. It clears out any that didn't have an actual response in case that happens. an actual response in case that happens. an actual response in case that happens. But you can also filter based on a given But you can also filter based on a given But you can also filter based on a given field. Like if you had a field in the field. Like if you had a field in the field. Like if you had a field in the response that was should keep response that was should keep response that was should keep researching true or false or needs researching true or false or needs researching true or false or needs another review or needs followup or another review or needs followup or another review or needs followup or solves problem whatever it decides to solves problem whatever it decides to solves problem whatever it decides to put in here. It can now use that content put in here. It can now use that content put in here. It can now use that content to determine where to pass off the work to determine where to pass off the work to determine where to pass off the work next because it is code. The problem next because it is code. The problem next because it is code. The problem with most sub aent implementations is with most sub aent implementations is with most sub aent implementations is that they're leaning too hard into tool that they're leaning too hard into tool that they're leaning too hard into tool calls. They work because the model will calls. They work because the model will calls. They work because the model will spit out a call spin-up sub agent with spit out a call spin-up sub agent with spit out a call spin-up sub agent with this information and then it does it and this information and then it does it and this information and then it does it and then it does another tool call for then it does another tool call for then it does another tool call for another and then it does it and then it another and then it does it and then it another and then it does it and then it does it again where the agent has to does it again where the agent has to does it again where the agent has to keep all of this context itself and spin keep all of this context itself and spin keep all of this context itself and spin everything up as part of the LLM's work.

  18. everything up as part of the LLM's work. everything up as part of the LLM's work. With workflows it's hardcoded but it's With workflows it's hardcoded but it's With workflows it's hardcoded but it's hardcoded on the fly. Previously, sub hardcoded on the fly. Previously, sub hardcoded on the fly. Previously, sub aents were just full-on hard-coded where aents were just full-on hard-coded where aents were just full-on hard-coded where certain harnesses cough cough, oh Mike, certain harnesses cough cough, oh Mike, certain harnesses cough cough, oh Mike, pi, cough, cough, open code had pi, cough, cough, open code had pi, cough, cough, open code had hardcoded sub aent descriptions like hardcoded sub aent descriptions like hardcoded sub aent descriptions like this is a researcher sub aent. It does this is a researcher sub aent. It does this is a researcher sub aent. It does these things. This is an implement sub these things. This is an implement sub these things. This is an implement sub aent. It does these things. Workflows is aent. It does these things. Workflows is aent. It does these things. Workflows is a really cool in between where on the a really cool in between where on the a really cool in between where on the fly it will create these different sub fly it will create these different sub fly it will create these different sub aent types and classes and what they're aent types and classes and what they're aent types and classes and what they're expected to return with and then create expected to return with and then create expected to return with and then create a flow, a workflow if you will, where a flow, a workflow if you will, where a flow, a workflow if you will, where each stage is determined based on what each stage is determined based on what each stage is determined based on what happened on the fly with code that was happened on the fly with code that was happened on the fly with code that was written ahead of time. This does a few written ahead of time. This does a few written ahead of time. This does a few things really well, but the biggest it things really well, but the biggest it things really well, but the biggest it does is it kind of puts a hard cap on does is it kind of puts a hard cap on does is it kind of puts a hard cap on when this all ends. Instead of just when this all ends. Instead of just when this all ends. Instead of just going forever, which is really possible going forever, which is really possible going forever, which is really possible with Codeex's Ultra implementation, with Codeex's Ultra implementation, with Codeex's Ultra implementation, there's a fixed number of phases. So there's a fixed number of phases. So there's a fixed number of phases. So eventually it will get to the end. It eventually it will get to the end. It eventually it will get to the end. It might spin up a shitload of sub agents might spin up a shitload of sub agents might spin up a shitload of sub agents in one field. For example, if you ask it in one field. For example, if you ask it in one field. For example, if you ask it to read three files and find things that to read three files and find things that to read three files and find things that are worth fixing, it might find 72 are worth fixing, it might find 72 are worth fixing, it might find 72 things to fix in the first file and it things to fix in the first file and it things to fix in the first file and it might spin up 72 sub aents in the next might spin up 72 sub aents in the next might spin up 72 sub aents in the next section, the next phase where it is section, the next phase where it is section, the next phase where it is fixing the things that it finds. That's fixing the things that it finds. That's fixing the things that it finds. That's fine though because it will eventually fine though because it will eventually fine though because it will eventually get to the end always. Ultra is much get to the end always. Ultra is much get to the end always. Ultra is much more likely to just go forever because more likely to just go forever because more likely to just go forever because it ends when the agents decide to end, it ends when the agents decide to end, it ends when the agents decide to end, not when the code is completed. So, what not when the code is completed. So, what not when the code is completed. So, what I'm trying to say is I think OpenAI I'm trying to say is I think OpenAI I'm trying to say is I think OpenAI copied the wrong parts. They copied the copied the wrong parts. They copied the copied the wrong parts. They copied the UX, which was bad, where they disguised

  19. UX, which was bad, where they disguised UX, which was bad, where they disguised it as a reasoning level where it isn't, it as a reasoning level where it isn't, it as a reasoning level where it isn't, and they didn't copy the implementation, and they didn't copy the implementation, and they didn't copy the implementation, which is good, because workflows are an which is good, because workflows are an which is good, because workflows are an awesome way to break up work awesome way to break up work awesome way to break up work programmatically. They are so awesome programmatically. They are so awesome programmatically. They are so awesome that I'm causing problems for a lot of that I'm causing problems for a lot of that I'm causing problems for a lot of people because of it. And I have a whole people because of it. And I have a whole people because of it. And I have a whole dedicated video coming up about my usage dedicated video coming up about my usage dedicated video coming up about my usage of 56 soul in Claude Code. It's probably of 56 soul in Claude Code. It's probably of 56 soul in Claude Code. It's probably going to be the next video I post. So, going to be the next video I post. So, going to be the next video I post. So, make sure you subscribe and hit that make sure you subscribe and hit that make sure you subscribe and hit that bell if you want to see it as soon as it bell if you want to see it as soon as it bell if you want to see it as soon as it drops. I think you'll be surprised. drops. I think you'll be surprised. drops. I think you'll be surprised. Spoiler, I'm going to spend the majority Spoiler, I'm going to spend the majority Spoiler, I'm going to spend the majority of this video crashing out about the of this video crashing out about the of this video crashing out about the system prompt in Codeex because it is so system prompt in Codeex because it is so system prompt in Codeex because it is so much worse than I thought it would be. much worse than I thought it would be. much worse than I thought it would be. One last thing, shout out to Maria for One last thing, shout out to Maria for One last thing, shout out to Maria for making this awesome demo of what the making this awesome demo of what the making this awesome demo of what the model selector should look like in model selector should look like in model selector should look like in Codeex. Sorry, chat GPT. You have the Codeex. Sorry, chat GPT. You have the Codeex. Sorry, chat GPT. You have the different models in different effort different models in different effort different models in different effort levels, but you also have ultra as a levels, but you also have ultra as a levels, but you also have ultra as a switch on the side. Because again, ultra switch on the side. Because again, ultra switch on the side. Because again, ultra isn't an effort level. Ultra is a skill. isn't an effort level. Ultra is a skill. isn't an effort level. Ultra is a skill. And it seems like everybody from TBO to And it seems like everybody from TBO to And it seems like everybody from TBO to Dominic to even GDB himself agrees that Dominic to even GDB himself agrees that Dominic to even GDB himself agrees that as silly as this is, something like this as silly as this is, something like this as silly as this is, something like this does make sense. So for now, don't use does make sense. So for now, don't use does make sense. So for now, don't use ultra. If you really want to get this ultra. If you really want to get this ultra. If you really want to get this type of thing, keep an eye out for my type of thing, keep an eye out for my type of thing, keep an eye out for my video on how to use workflows in cloud video on how to use workflows in cloud video on how to use workflows in cloud code with GPD56 soul because even Tibo code with GPD56 soul because even Tibo code with GPD56 soul because even Tibo supports my chaos here. Hope this was supports my chaos here. Hope this was supports my chaos here. Hope this was helpful and until next time, peace helpful and until next time, peace helpful and until next time, peace nerds.

Summary

The discussion centers on OpenAI's GBD 5.6 release, particularly its "Ultra" reasoning level, and criticizes its rollout and perceived misleading marketing. The speaker argues that "Ultra" is not a true reasoning level but a poorly implemented feature, blaming Anthropic for setting a precedent of misleading AI capabilities. The practical takeaway is to be skeptical of new AI features and understand their true nature before relying on them for critical tasks.

View original episode ↗