This is absolute chaos...
Read full transcript 23 segments
-
GBT 56 is an incredible model, but it's GBT 56 is an incredible model, but it's also incredible at burning through also incredible at burning through also incredible at burning through people's rate limits. It used to be people's rate limits. It used to be people's rate limits. It used to be really hard to hit 5-hour limits on GBT really hard to hit 5-hour limits on GBT really hard to hit 5-hour limits on GBT 55, even X high. Now at 56 soul, it's 55, even X high. Now at 56 soul, it's 55, even X high. Now at 56 soul, it's draining super fast, even on medium draining super fast, even on medium draining super fast, even on medium reasoning. I've already burned through reasoning. I've already burned through reasoning. I've already burned through three 5 hour limits. Seriously, this has three 5 hour limits. Seriously, this has three 5 hour limits. Seriously, this has to stop. I've now set 56 from high to to stop. I've now set 56 from high to to stop. I've now set 56 from high to medium, not fast mode. And I'm still medium, not fast mode. And I'm still medium, not fast mode. And I'm still burning through my rates at an insane burning through my rates at an insane burning through my rates at an insane rate. My 5 hours are almost gone again. rate. My 5 hours are almost gone again. rate. My 5 hours are almost gone again. I've hit the limits every single time I've hit the limits every single time I've hit the limits every single time since Soul was released, and I don't since Soul was released, and I don't since Soul was released, and I don't think I ever hit it with 55, bro. I've think I ever hit it with 55, bro. I've think I ever hit it with 55, bro. I've waxed a weekly and a half on one message waxed a weekly and a half on one message waxed a weekly and a half on one message in X high. With GBD55 on the $100 plan, in X high. With GBD55 on the $100 plan, in X high. With GBD55 on the $100 plan, I could barely use up my quota. Now, one I could barely use up my quota. Now, one I could barely use up my quota. Now, one small PR with 56 soul consumes about small PR with 56 soul consumes about small PR with 56 soul consumes about half of my 5h hour limit. Not going to half of my 5h hour limit. Not going to half of my 5h hour limit. Not going to lie, GBD56 soul is great and all, but lie, GBD56 soul is great and all, but lie, GBD56 soul is great and all, but the usage drain is whack. Used to be the usage drain is whack. Used to be the usage drain is whack. Used to be able to work consistently all day with able to work consistently all day with able to work consistently all day with 55 on XI. Now, I can't even get through 55 on XI. Now, I can't even get through 55 on XI. Now, I can't even get through a single 3-hour session without hitting a single 3-hour session without hitting a single 3-hour session without hitting limits. Whack. Not good for long-term limits. Whack. Not good for long-term limits. Whack. Not good for long-term work. As you can tell, people aren't work. As you can tell, people aren't work. As you can tell, people aren't happy. And I'll be frank, a lot of this happy. And I'll be frank, a lot of this happy. And I'll be frank, a lot of this is OpenAI's fault. There are some things is OpenAI's fault. There are some things is OpenAI's fault. There are some things in codeex right now that just don't make in codeex right now that just don't make in codeex right now that just don't make sense. I want to do my best to help you sense. I want to do my best to help you sense. I want to do my best to help you guys out to make sure you can get the guys out to make sure you can get the guys out to make sure you can get the most out of 56 without hitting these most out of 56 without hitting these most out of 56 without hitting these limits. I made a bunch of these changes limits. I made a bunch of these changes limits. I made a bunch of these changes myself and I've actually noticed the myself and I've actually noticed the myself and I've actually noticed the quality of outputs going up while also quality of outputs going up while also quality of outputs going up while also using as little as a fourth or a fifth using as little as a fourth or a fifth using as little as a fourth or a fifth as many tokens as I was before. I took as many tokens as I was before. I took as many tokens as I was before. I took the time to write an article sharing my the time to write an article sharing my the time to write an article sharing my advice here, but since then quite a bit advice here, but since then quite a bit advice here, but since then quite a bit has changed even though it's been 24 has changed even though it's been 24 has changed even though it's been 24 hours. I've also seen a bunch of really hours. I've also seen a bunch of really hours. I've also seen a bunch of really bad advice going around that has been bad advice going around that has been bad advice going around that has been confirmed by many, including the actual confirmed by many, including the actual confirmed by many, including the actual team at OpenAI to not be good. So, if team at OpenAI to not be good. So, if team at OpenAI to not be good. So, if you want to get the most out of your you want to get the most out of your you want to get the most out of your subscription plans with Codeex, I hope
-
subscription plans with Codeex, I hope subscription plans with Codeex, I hope this is helpful. And even if you don't, this is helpful. And even if you don't, this is helpful. And even if you don't, maybe you're using something else like maybe you're using something else like maybe you're using something else like Cursor or Claude Code, this video will Cursor or Claude Code, this video will Cursor or Claude Code, this video will still have a lot of tips that could help still have a lot of tips that could help still have a lot of tips that could help you be better at using those in an you be better at using those in an you be better at using those in an efficient, effective way. I have never efficient, effective way. I have never efficient, effective way. I have never burned as many tokens as I have with 56 burned as many tokens as I have with 56 burned as many tokens as I have with 56 soul and I've really come to know how soul and I've really come to know how soul and I've really come to know how this model works, what its strengths and this model works, what its strengths and this model works, what its strengths and weaknesses are, and what gets it to weaknesses are, and what gets it to weaknesses are, and what gets it to burn. I wanted to take the time to share burn. I wanted to take the time to share burn. I wanted to take the time to share this all with you. But this this all with you. But this this all with you. But this experimentation was not cheap. So I hope experimentation was not cheap. So I hope experimentation was not cheap. So I hope you can pardon a quick break for today's you can pardon a quick break for today's you can pardon a quick break for today's sponsor. As AI writes more and more of sponsor. As AI writes more and more of sponsor. As AI writes more and more of our code, the chances of it getting our code, the chances of it getting our code, the chances of it getting things wrong goes up, too. And as smart things wrong goes up, too. And as smart things wrong goes up, too. And as smart as the models are, there are certain as the models are, there are certain as the models are, there are certain places you just shouldn't take the risk. places you just shouldn't take the risk. places you just shouldn't take the risk. And if you're trying to land enterprise And if you're trying to land enterprise And if you're trying to land enterprise customers, the O layer is one of the customers, the O layer is one of the customers, the O layer is one of the worst places to take that risk. That's worst places to take that risk. That's worst places to take that risk. That's why work OS has always been a great why work OS has always been a great why work OS has always been a great option. They really understand option. They really understand option. They really understand enterprise and they don't compromise on enterprise and they don't compromise on enterprise and they don't compromise on the developer experience to get there. the developer experience to get there. the developer experience to get there. They have really good SDKs, really good They have really good SDKs, really good They have really good SDKs, really good integrations, really good docs, and integrations, really good docs, and integrations, really good docs, and everything else you need to set up your everything else you need to set up your everything else you need to set up your apps in a great way. They also give you apps in a great way. They also give you apps in a great way. They also give you up to a million users for free, which is up to a million users for free, which is up to a million users for free, which is just insane. And I can't imagine many of just insane. And I can't imagine many of just insane. And I can't imagine many of us having that problem. I I wish I had us having that problem. I I wish I had us having that problem. I I wish I had enough users to hit the limits and need enough users to hit the limits and need enough users to hit the limits and need to pay money, but I don't. So, it's a to pay money, but I don't. So, it's a to pay money, but I don't. So, it's a really generous offer. But nowadays, really generous offer. But nowadays, really generous offer. But nowadays, users aren't your only users. Your users users aren't your only users. Your users users aren't your only users. Your users have agents. How are those agents going have agents. How are those agents going have agents. How are those agents going to O? And if a company like Microsoft to O? And if a company like Microsoft to O? And if a company like Microsoft wants their agents to use your service, wants their agents to use your service, wants their agents to use your service, how are they going to get in? Work OS is how are they going to get in? Work OS is how are they going to get in? Work OS is one of the few companies thinking about one of the few companies thinking about one of the few companies thinking about this, which is why they partnered with this, which is why they partnered with this, which is why they partnered with many others to build OMD. They're many others to build OMD. They're many others to build OMD. They're working with companies like Cloudflare working with companies like Cloudflare working with companies like Cloudflare and Firecrawl to get all the pieces and Firecrawl to get all the pieces and Firecrawl to get all the pieces necessary so your agents can off for you necessary so your agents can off for you necessary so your agents can off for you can create accounts for you and then let can create accounts for you and then let can create accounts for you and then let you attach them to your service account you attach them to your service account you attach them to your service account whatever in all of the logical ways that
-
whatever in all of the logical ways that whatever in all of the logical ways that are needed. To be very very clear, OMD are needed. To be very very clear, OMD are needed. To be very very clear, OMD is not a works feature. It's an open is not a works feature. It's an open is not a works feature. It's an open standard that they helped build. They're standard that they helped build. They're standard that they helped build. They're the authors of the protocol and the authors of the protocol and the authors of the protocol and obviously that means you can oneclick obviously that means you can oneclick obviously that means you can oneclick turn it on in your work OS apps, but turn it on in your work OS apps, but turn it on in your work OS apps, but this is a standard that anyone could this is a standard that anyone could this is a standard that anyone could implement and I'm excited for everyone implement and I'm excited for everyone implement and I'm excited for everyone else too as well. Work OS is great else too as well. Work OS is great else too as well. Work OS is great because they let you sell to enterprises because they let you sell to enterprises because they let you sell to enterprises now, but they also set you up for a now, but they also set you up for a now, but they also set you up for a future where the agents are signing up future where the agents are signing up future where the agents are signing up instead. Get enterprise proof and agent instead. Get enterprise proof and agent instead. Get enterprise proof and agent ready at soyv.link/workos. ready at soyv.link/workos. ready at soyv.link/workos. Time to talk about how to use soul Time to talk about how to use soul Time to talk about how to use soul without hitting limits. The model is without hitting limits. The model is without hitting limits. The model is great and the $200 plan is still great. great and the $200 plan is still great. great and the $200 plan is still great. Even the $100 plan is pretty generous. Even the $100 plan is pretty generous. Even the $100 plan is pretty generous. I've not sure what the state of the $20 I've not sure what the state of the $20 I've not sure what the state of the $20 plan is right now, but the $100 one is plan is right now, but the $100 one is plan is right now, but the $100 one is fine. The $200 one is really good as fine. The $200 one is really good as fine. The $200 one is really good as long as you don't make certain mistakes. long as you don't make certain mistakes. long as you don't make certain mistakes. And again, to your credit, OpenAI is not And again, to your credit, OpenAI is not And again, to your credit, OpenAI is not doing this right. There are a lot of doing this right. There are a lot of doing this right. There are a lot of little things that exist in codecs right little things that exist in codecs right little things that exist in codecs right now that are kind of forcing you to now that are kind of forcing you to now that are kind of forcing you to overuse your usage and this is their overuse your usage and this is their overuse your usage and this is their fault. I feel bad if people in chat are fault. I feel bad if people in chat are fault. I feel bad if people in chat are saying the $20 plan is actually pretty saying the $20 plan is actually pretty saying the $20 plan is actually pretty good for coding. That's good to hear. I good for coding. That's good to hear. I good for coding. That's good to hear. I One person said it's horrible, but One person said it's horrible, but One person said it's horrible, but others are saying it's fine. So yeah, others are saying it's fine. So yeah, others are saying it's fine. So yeah, take it with a grain of salt. Apparently take it with a grain of salt. Apparently take it with a grain of salt. Apparently it's okay. I am on the $200 plan and I it's okay. I am on the $200 plan and I it's okay. I am on the $200 plan and I was hitting limits aggressively when was hitting limits aggressively when was hitting limits aggressively when they made this move. Also, full they made this move. Also, full they made this move. Also, full disclosure, when I was testing 56, it disclosure, when I was testing 56, it disclosure, when I was testing 56, it didn't count against my usage, which is didn't count against my usage, which is didn't count against my usage, which is why I was able to push the model so hard why I was able to push the model so hard why I was able to push the model so hard and do a genuinely absurd amount of and do a genuinely absurd amount of and do a genuinely absurd amount of inference on it. But as soon as it came inference on it. But as soon as it came inference on it. But as soon as it came back, just a day or two before it went
-
back, just a day or two before it went back, just a day or two before it went public, I started hitting limits public, I started hitting limits public, I started hitting limits immediately. It was 20 minutes after immediately. It was 20 minutes after immediately. It was 20 minutes after they said, "Okay, you have it again." they said, "Okay, you have it again." they said, "Okay, you have it again." That I hit a limit for the first time. That I hit a limit for the first time. That I hit a limit for the first time. And the reason I hit that limit was a And the reason I hit that limit was a And the reason I hit that limit was a certain new feature called Ultra. I have certain new feature called Ultra. I have certain new feature called Ultra. I have a dedicated video about Ultra coming a dedicated video about Ultra coming a dedicated video about Ultra coming very soon. It might even be the next very soon. It might even be the next very soon. It might even be the next video after this one. So, as always, video after this one. So, as always, video after this one. So, as always, make sure you subscribe and hit that make sure you subscribe and hit that make sure you subscribe and hit that bell if you want to better understand bell if you want to better understand bell if you want to better understand all of these pieces. If I was to talk all of these pieces. If I was to talk all of these pieces. If I was to talk about Ultra in depth right now, this about Ultra in depth right now, this about Ultra in depth right now, this video would be two plus hours long. I video would be two plus hours long. I video would be two plus hours long. I want this to be to the point and easy to want this to be to the point and easy to want this to be to the point and easy to apply to your work and send to apply to your work and send to apply to your work and send to co-workers and whatnot. So, I won't do co-workers and whatnot. So, I won't do co-workers and whatnot. So, I won't do that here. But know that Ultra is not that here. But know that Ultra is not that here. But know that Ultra is not worth using at the very least until you worth using at the very least until you worth using at the very least until you wait for that follow-up video. So, for wait for that follow-up video. So, for wait for that follow-up video. So, for now, Ultra, no. Ignore this. Don't touch now, Ultra, no. Ignore this. Don't touch now, Ultra, no. Ignore this. Don't touch it for now. This might change in the it for now. This might change in the it for now. This might change in the future. I am hopeful that it will, but future. I am hopeful that it will, but future. I am hopeful that it will, but for now, Ultra is best avoided, not for now, Ultra is best avoided, not for now, Ultra is best avoided, not used. With Ultra out of the way, I do used. With Ultra out of the way, I do used. With Ultra out of the way, I do want to talk about a thing that just want to talk about a thing that just want to talk about a thing that just changed cuz it's quite important. Open changed cuz it's quite important. Open changed cuz it's quite important. Open eye is aware of the issues. I've been eye is aware of the issues. I've been eye is aware of the issues. I've been working with them a bunch on getting working with them a bunch on getting working with them a bunch on getting these things fixed. Hopefully, more of these things fixed. Hopefully, more of these things fixed. Hopefully, more of them are fixed before this video is them are fixed before this video is them are fixed before this video is live. Probably not too many though.
-
live. Probably not too many though. live. Probably not too many though. We'll still have useful tips long term, We'll still have useful tips long term, We'll still have useful tips long term, but I want to give you guys context on but I want to give you guys context on but I want to give you guys context on where things are at the moment of where things are at the moment of where things are at the moment of filming. TBO posted this morning the filming. TBO posted this morning the filming. TBO posted this morning the following update, which is very relevant following update, which is very relevant following update, which is very relevant to what we're talking about here. The to what we're talking about here. The to what we're talking about here. The last 48 hours of codeex and chatgbt work last 48 hours of codeex and chatgbt work last 48 hours of codeex and chatgbt work have been intense. There's three have been intense. There's three have been intense. There's three important updates. First is that they've important updates. First is that they've important updates. First is that they've temporarily removed the 5hour usage temporarily removed the 5hour usage temporarily removed the 5hour usage limits for all plus business and pro limits for all plus business and pro limits for all plus business and pro plans. If you're not familiar, your sub plans. If you're not familiar, your sub plans. If you're not familiar, your sub is broken up into two limits. You have a is broken up into two limits. You have a is broken up into two limits. You have a 5h hour limit that resets every 5 hours 5h hour limit that resets every 5 hours 5h hour limit that resets every 5 hours starting from when you send a message. starting from when you send a message. starting from when you send a message. So let's say you do a bunch of work at So let's say you do a bunch of work at So let's say you do a bunch of work at the start of the day. You start at 9:00 the start of the day. You start at 9:00 the start of the day. You start at 9:00 a.m. You have a limit to how much you a.m. You have a limit to how much you a.m. You have a limit to how much you can use between 9:00 a.m. and 2:00 p.m. can use between 9:00 a.m. and 2:00 p.m. can use between 9:00 a.m. and 2:00 p.m. And then let's say you hit that limit at And then let's say you hit that limit at And then let's say you hit that limit at 1:00 p.m. You have an hour until you can 1:00 p.m. You have an hour until you can 1:00 p.m. You have an hour until you can use it again. But if you don't start use it again. But if you don't start use it again. But if you don't start until 3, the next 5 hour limit doesn't until 3, the next 5 hour limit doesn't until 3, the next 5 hour limit doesn't start until 3:00. One of the tricks I start until 3:00. One of the tricks I start until 3:00. One of the tricks I have is I have a cron job running that have is I have a cron job running that have is I have a cron job running that does something every 5 hours so that I'm does something every 5 hours so that I'm does something every 5 hours so that I'm always burning one of those 5 hour always burning one of those 5 hour always burning one of those 5 hour limits. It uses a very small model with limits. It uses a very small model with limits. It uses a very small model with a really simple like hello, hi prompt. a really simple like hello, hi prompt. a really simple like hello, hi prompt. But now you don't have to, at least in But now you don't have to, at least in But now you don't have to, at least in the interim, because the 5-hour limit is the interim, because the 5-hour limit is the interim, because the 5-hour limit is gone for now. The weekly limit is the gone for now. The weekly limit is the gone for now. The weekly limit is the one to be scared about. The weekly limit one to be scared about. The weekly limit one to be scared about. The weekly limit is equal to roughly four or five of the is equal to roughly four or five of the is equal to roughly four or five of the 5-hour limits. So if you go hard enough 5-hour limits. So if you go hard enough 5-hour limits. So if you go hard enough to max out the 5-hour limit four plus to max out the 5-hour limit four plus to max out the 5-hour limit four plus times, you will hit the weekly limit.
-
times, you will hit the weekly limit. times, you will hit the weekly limit. Weekly is the only one applying right Weekly is the only one applying right Weekly is the only one applying right now, which is good because the 5-hour now, which is good because the 5-hour now, which is good because the 5-hour limits were a little too easy to hit. limits were a little too easy to hit. limits were a little too easy to hit. But it's bad because if you accidentally But it's bad because if you accidentally But it's bad because if you accidentally ran an ultra fast run that is way ran an ultra fast run that is way ran an ultra fast run that is way heavier than expected, the five hour heavier than expected, the five hour heavier than expected, the five hour limit would have stopped it before it limit would have stopped it before it limit would have stopped it before it used more than 25% of your weekly. Now used more than 25% of your weekly. Now used more than 25% of your weekly. Now it won't. So it's possible to run one it won't. So it's possible to run one it won't. So it's possible to run one prompt and blow through your whole prompt and blow through your whole prompt and blow through your whole weekly limit. I'm tempted to do it as an weekly limit. I'm tempted to do it as an weekly limit. I'm tempted to do it as an example and then burn a reset, but I'm example and then burn a reset, but I'm example and then burn a reset, but I'm I'm hoarding my resets. I'll be real, I'm hoarding my resets. I'll be real, I'm hoarding my resets. I'll be real, guys. So with all of this done, let's go guys. So with all of this done, let's go guys. So with all of this done, let's go through the rest here. The second point through the rest here. The second point through the rest here. The second point he made is rolling out changes that will he made is rolling out changes that will he made is rolling out changes that will make 56 soul more efficient across the make 56 soul more efficient across the make 56 soul more efficient across the board and it will be reflected in less board and it will be reflected in less board and it will be reflected in less usage being used so it can take you usage being used so it can take you usage being used so it can take you further exact impact to be quantified further exact impact to be quantified further exact impact to be quantified and shared. Still working on that there. and shared. Still working on that there. and shared. Still working on that there. And his most excited point here is that And his most excited point here is that And his most excited point here is that they have 6 million active users and they have 6 million active users and they have 6 million active users and they're about to land a usage reset. they're about to land a usage reset. they're about to land a usage reset. That one already landed. So as was That one already landed. So as was That one already landed. So as was established prior, they got rid of the 5 established prior, they got rid of the 5 established prior, they got rid of the 5 hour for now and we just have the weekly hour for now and we just have the weekly hour for now and we just have the weekly limit. I fired one prompt off before limit. I fired one prompt off before limit. I fired one prompt off before streaming, so I don't have much used streaming, so I don't have much used streaming, so I don't have much used here yet, but I'll probably burn through here yet, but I'll probably burn through here yet, but I'll probably burn through this in the next day or two, and I'll be this in the next day or two, and I'll be this in the next day or two, and I'll be sure to share how, why, and what I'm sure to share how, why, and what I'm sure to share how, why, and what I'm doing to minimize and maximize usage.
-
doing to minimize and maximize usage. doing to minimize and maximize usage. With all this established, it's time to With all this established, it's time to With all this established, it's time to give practical advice. Part of how I give practical advice. Part of how I give practical advice. Part of how I want to think about this is with 55 want to think about this is with 55 want to think about this is with 55 versus 56. When 55 came out, I was far versus 56. When 55 came out, I was far versus 56. When 55 came out, I was far from the biggest fan. I saw how capable from the biggest fan. I saw how capable from the biggest fan. I saw how capable it could be, but I didn't love using it it could be, but I didn't love using it it could be, but I didn't love using it because it felt like it would just lose because it felt like it would just lose because it felt like it would just lose track of what I was doing. It was really track of what I was doing. It was really track of what I was doing. It was really easy to screw up its context. If it read easy to screw up its context. If it read easy to screw up its context. If it read the wrong file and had something in its the wrong file and had something in its the wrong file and had something in its history, it would fixate on that instead history, it would fixate on that instead history, it would fixate on that instead of doing what I asked it to do. And this of doing what I asked it to do. And this of doing what I asked it to do. And this happened a lot. It also was really, happened a lot. It also was really, happened a lot. It also was really, really bad at compaction, which was just really bad at compaction, which was just really bad at compaction, which was just obnoxious when you had longunning obnoxious when you had longunning obnoxious when you had longunning threads. I've never had more threads threads. I've never had more threads threads. I've never had more threads than I did with 55 because I felt like I than I did with 55 because I felt like I than I did with 55 because I felt like I had to in order to keep it on task. None had to in order to keep it on task. None had to in order to keep it on task. None of that is why 55 was cheaper, though. of that is why 55 was cheaper, though. of that is why 55 was cheaper, though. The main reason 55 was cheaper is that The main reason 55 was cheaper is that The main reason 55 was cheaper is that it stopped and asked for permission or it stopped and asked for permission or it stopped and asked for permission or feedback all of the time. If you asked feedback all of the time. If you asked feedback all of the time. If you asked 55 to write a plan, it would it would 55 to write a plan, it would it would 55 to write a plan, it would it would say, "Here's step one. Here's step two. say, "Here's step one. Here's step two. say, "Here's step one. Here's step two. Here's step three through 10." And then Here's step three through 10." And then Here's step three through 10." And then you would say, "Okay, go do all of the you would say, "Okay, go do all of the you would say, "Okay, go do all of the steps." It would get through step one. steps." It would get through step one. steps." It would get through step one. It would get halfway through step two.
-
It would get halfway through step two. It would get halfway through step two. And then it would stop and say, "Okay, I And then it would stop and say, "Okay, I And then it would stop and say, "Okay, I finished step one and I'm halfway finished step one and I'm halfway finished step one and I'm halfway through step two. Let me know when I can through step two. Let me know when I can through step two. Let me know when I can keep going." It's like, "I never told keep going." It's like, "I never told keep going." It's like, "I never told you to stop, bro." The result of this you to stop, bro." The result of this you to stop, bro." The result of this behavior is that 55 didn't really use behavior is that 55 didn't really use behavior is that 55 didn't really use your limits heavily on a per message your limits heavily on a per message your limits heavily on a per message basis. From my experience, a single basis. From my experience, a single basis. From my experience, a single message on 55, even on the highest message on 55, even on the highest message on 55, even on the highest reasoning levels, would use between reasoning levels, would use between reasoning levels, would use between like.1% and 2% of your 5h hour limit per like.1% and 2% of your 5h hour limit per like.1% and 2% of your 5h hour limit per message. It wasn't that bad. 56 fixed message. It wasn't that bad. 56 fixed message. It wasn't that bad. 56 fixed these problems, but in doing such these problems, but in doing such these problems, but in doing such massively increased how much usage massively increased how much usage massively increased how much usage you're getting. 55 was a price bump from you're getting. 55 was a price bump from you're getting. 55 was a price bump from 54. They went from $15 per mill out to 54. They went from $15 per mill out to 54. They went from $15 per mill out to 30 per mill out, which would have hurt a 30 per mill out, which would have hurt a 30 per mill out, which would have hurt a lot more if it wasn't for this sudden lot more if it wasn't for this sudden lot more if it wasn't for this sudden stopping behavior from 55. 56 fixed that stopping behavior from 55. 56 fixed that stopping behavior from 55. 56 fixed that behavior, but as a result, it can go behavior, but as a result, it can go behavior, but as a result, it can go much longer per message, which can much longer per message, which can much longer per message, which can result in much bigger usage per message. result in much bigger usage per message. result in much bigger usage per message. For my experience, on higher reasoning For my experience, on higher reasoning For my experience, on higher reasoning levels, especially X high and max, 56 levels, especially X high and max, 56 levels, especially X high and max, 56 can use up to 15% of my 5-hour limit in can use up to 15% of my 5-hour limit in can use up to 15% of my 5-hour limit in one message. This is a big jump one message. This is a big jump one message. This is a big jump obviously, but honestly, it's not too obviously, but honestly, it's not too obviously, but honestly, it's not too too bad because I would end up using too bad because I would end up using too bad because I would end up using similar amounts of 55, but it would be similar amounts of 55, but it would be similar amounts of 55, but it would be prompt after prompt after prompt because prompt after prompt after prompt because prompt after prompt after prompt because each message only would use so much of each message only would use so much of each message only would use so much of my limits. I didn't find it too brutal my limits. I didn't find it too brutal my limits. I didn't find it too brutal and I would often use it with fast. Fast and I would often use it with fast. Fast and I would often use it with fast. Fast mode is pretty cool because you get 1.5x mode is pretty cool because you get 1.5x mode is pretty cool because you get 1.5x faster speeds, but has a catch. You go faster speeds, but has a catch. You go faster speeds, but has a catch. You go through your limit 2.5x faster. And this through your limit 2.5x faster. And this through your limit 2.5x faster. And this wasn't too bad when I was using 55
-
wasn't too bad when I was using 55 wasn't too bad when I was using 55 because when I sent one message with 55, because when I sent one message with 55, because when I sent one message with 55, the 2.5x more burn would make that 2% the 2.5x more burn would make that 2% the 2.5x more burn would make that 2% into 5%. into 5%. into 5%. That wasn't too bad. 2.5x burn on a That wasn't too bad. 2.5x burn on a That wasn't too bad. 2.5x burn on a message that was 15% though, that's a message that was 15% though, that's a message that was 15% though, that's a little more brutal. That is closer to little more brutal. That is closer to little more brutal. That is closer to half of your 5hour limit from one half of your 5hour limit from one half of your 5hour limit from one message. And that's not even with Ultra. message. And that's not even with Ultra. message. And that's not even with Ultra. That is scary. And that is a problem I That is scary. And that is a problem I That is scary. And that is a problem I think a lot of people are having right think a lot of people are having right think a lot of people are having right now is that they are continuing to use now is that they are continuing to use now is that they are continuing to use the model the way they did with 55. And the model the way they did with 55. And the model the way they did with 55. And with 55 since each message was less with 55 since each message was less with 55 since each message was less usage, fast mode didn't feel too bad. So usage, fast mode didn't feel too bad. So usage, fast mode didn't feel too bad. So if you're hitting limits and you're if you're hitting limits and you're if you're hitting limits and you're using fast mode, please stop. It's not using fast mode, please stop. It's not using fast mode, please stop. It's not that much faster. The reason I think that much faster. The reason I think that much faster. The reason I think fast mode felt so good with 55 is fast mode felt so good with 55 is fast mode felt so good with 55 is because it would stop itself constantly because it would stop itself constantly because it would stop itself constantly and you had to talk back. I never and you had to talk back. I never and you had to talk back. I never watched my agents run quite as much as I watched my agents run quite as much as I watched my agents run quite as much as I did with 55. both because its context did with 55. both because its context did with 55. both because its context was constantly getting screwed up and I was constantly getting screwed up and I was constantly getting screwed up and I had to prune it, clean it, make a new had to prune it, clean it, make a new had to prune it, clean it, make a new thread, but also because it stopped all thread, but also because it stopped all thread, but also because it stopped all the time, so I would have to respond and the time, so I would have to respond and the time, so I would have to respond and tell it what to do next. As such, the tell it what to do next. As such, the tell it what to do next. As such, the speed mattered a lot because I was speed mattered a lot because I was speed mattered a lot because I was sitting there waiting for it to respond.
-
sitting there waiting for it to respond. sitting there waiting for it to respond. With 56, I spin it up and then I go do With 56, I spin it up and then I go do With 56, I spin it up and then I go do something else. Maybe I spin up another something else. Maybe I spin up another something else. Maybe I spin up another thread. Maybe I go do code reviews. thread. Maybe I go do code reviews. thread. Maybe I go do code reviews. Maybe I check email. Maybe I play Power Maybe I check email. Maybe I play Power Maybe I check email. Maybe I play Power World. I don't care. With 56, I do other World. I don't care. With 56, I do other World. I don't care. With 56, I do other things. With 55, I have to sit there and things. With 55, I have to sit there and things. With 55, I have to sit there and watch. So fast mode felt important. With watch. So fast mode felt important. With watch. So fast mode felt important. With 56, it's running for so long anyways 56, it's running for so long anyways 56, it's running for so long anyways that the inference is barely the slow that the inference is barely the slow that the inference is barely the slow part. It's usually the tool calls, the part. It's usually the tool calls, the part. It's usually the tool calls, the test runs, it's doing, all the other test runs, it's doing, all the other test runs, it's doing, all the other things it has to do. I've barely noticed things it has to do. I've barely noticed things it has to do. I've barely noticed a difference in speed since turning off a difference in speed since turning off a difference in speed since turning off fast mode, but I have noticed a massive fast mode, but I have noticed a massive fast mode, but I have noticed a massive shift in the amount of usage that I'm shift in the amount of usage that I'm shift in the amount of usage that I'm hitting with it. For context, there's hitting with it. For context, there's hitting with it. For context, there's lots of people in chat, including those lots of people in chat, including those lots of people in chat, including those on my team like Maria, that says almost on my team like Maria, that says almost on my team like Maria, that says almost every prompt and thread they've done of every prompt and thread they've done of every prompt and thread they've done of 56 has lasted over 8 hours. Yeah, you 56 has lasted over 8 hours. Yeah, you 56 has lasted over 8 hours. Yeah, you can get these models going for a while can get these models going for a while can get these models going for a while even without using slash goal. To be even without using slash goal. To be even without using slash goal. To be clear, I almost felt like slashgoal was clear, I almost felt like slashgoal was clear, I almost felt like slashgoal was necessary with 55 because otherwise it necessary with 55 because otherwise it necessary with 55 because otherwise it would stop and the goal would basically would stop and the goal would basically would stop and the goal would basically tell it to keep going over and over tell it to keep going over and over tell it to keep going over and over instead of having a human do it with 56. instead of having a human do it with 56. instead of having a human do it with 56. Not necessary. I never use slash goal Not necessary. I never use slash goal Not necessary. I never use slash goal anymore unless I'm just trying to burn anymore unless I'm just trying to burn anymore unless I'm just trying to burn tokens. Generally with 56, it will tokens. Generally with 56, it will tokens. Generally with 56, it will complete the thing without needing extra complete the thing without needing extra complete the thing without needing extra encouragement. So now establish two encouragement. So now establish two encouragement. So now establish two things. Don't touch ultra. You should things. Don't touch ultra. You should things. Don't touch ultra. You should probably turn off fast mode at this probably turn off fast mode at this probably turn off fast mode at this point because it doesn't matter as much point because it doesn't matter as much point because it doesn't matter as much as it used to, but there's more that we as it used to, but there's more that we as it used to, but there's more that we can learn from. Reasoning level should can learn from. Reasoning level should can learn from. Reasoning level should be an easy enough one to cover. I'll be an easy enough one to cover. I'll be an easy enough one to cover. I'll dive through this super quick. You have dive through this super quick. You have dive through this super quick. You have five options. They've been rebranded, five options. They've been rebranded, five options. They've been rebranded, but they used to be low, medium, high, X but they used to be low, medium, high, X but they used to be low, medium, high, X high, and max. You might think Altra is high, and max. You might think Altra is high, and max. You might think Altra is a reasoning level. It's not. Again,
-
a reasoning level. It's not. Again, a reasoning level. It's not. Again, video crashing out about ultra coming video crashing out about ultra coming video crashing out about ultra coming very, very soon. Hit that button if you very, very soon. Hit that button if you very, very soon. Hit that button if you haven't at the bottom. Subscribing makes haven't at the bottom. Subscribing makes haven't at the bottom. Subscribing makes it more likely you see that when it it more likely you see that when it it more likely you see that when it drops. So, with Altra removed, we have drops. So, with Altra removed, we have drops. So, with Altra removed, we have these five options. I did a really good these five options. I did a really good these five options. I did a really good analogy about what these are and how analogy about what these are and how analogy about what these are and how they work on my most recent podcast they work on my most recent podcast they work on my most recent podcast episode that should hopefully be out in episode that should hopefully be out in episode that should hopefully be out in the next day or two. If you're not the next day or two. If you're not the next day or two. If you're not already listening to the podcast, check already listening to the podcast, check already listening to the podcast, check that out as well. Ben and I just nerd that out as well. Ben and I just nerd that out as well. Ben and I just nerd out about the details here and it lets out about the details here and it lets out about the details here and it lets me go a little longer than I try to on me go a little longer than I try to on me go a little longer than I try to on the videos. Trying to keep these shorter the videos. Trying to keep these shorter the videos. Trying to keep these shorter lately. Hope you appreciate the effort. lately. Hope you appreciate the effort. lately. Hope you appreciate the effort. So, instead of breaking down all of the So, instead of breaking down all of the So, instead of breaking down all of the details here, I'll give it to you details here, I'll give it to you details here, I'll give it to you simple. Ignore everything past this line simple. Ignore everything past this line simple. Ignore everything past this line for now. X high can be really cool. Max for now. X high can be really cool. Max for now. X high can be really cool. Max is less likely to be low, medium, and is less likely to be low, medium, and is less likely to be low, medium, and high are all very good options. I mostly high are all very good options. I mostly high are all very good options. I mostly just default to high right now because just default to high right now because just default to high right now because the model's so efficient. Even on high the model's so efficient. Even on high the model's so efficient. Even on high reasoning, if the task is simple, it reasoning, if the task is simple, it reasoning, if the task is simple, it will stop really quickly and not reason will stop really quickly and not reason will stop really quickly and not reason too much. When I was benchmarking the too much. When I was benchmarking the too much. When I was benchmarking the model, I thought there were bugs in how model, I thought there were bugs in how model, I thought there were bugs in how I was passing the reasoning level I was passing the reasoning level I was passing the reasoning level because the difference between low and because the difference between low and because the difference between low and high was like 50 to 100 tokens at most, high was like 50 to 100 tokens at most, high was like 50 to 100 tokens at most, which just didn't seem right at all.
-
which just didn't seem right at all. which just didn't seem right at all. Turns out it is actually just good at Turns out it is actually just good at Turns out it is actually just good at that. it will use less tokens if the that. it will use less tokens if the that. it will use less tokens if the task is simple. So leaving it on high task is simple. So leaving it on high task is simple. So leaving it on high has been fine for me. I do want to look has been fine for me. I do want to look has been fine for me. I do want to look at the numbers quick though. So let's do at the numbers quick though. So let's do at the numbers quick though. So let's do that with deepsw SWE. This is my that with deepsw SWE. This is my that with deepsw SWE. This is my favorite code benchmark right now. Very favorite code benchmark right now. Very favorite code benchmark right now. Very much subject to change eventually. I'll much subject to change eventually. I'll much subject to change eventually. I'll throw in Fable for a comparison here. So throw in Fable for a comparison here. So throw in Fable for a comparison here. So with GPT 56 soul on low you score 45% with GPT 56 soul on low you score 45% with GPT 56 soul on low you score 45% and in this bench it cost a dollar per and in this bench it cost a dollar per and in this bench it cost a dollar per task. Medium bumps all the way from a task. Medium bumps all the way from a task. Medium bumps all the way from a 45% to a 61% for a $186 per task. And 45% to a 61% for a $186 per task. And 45% to a 61% for a $186 per task. And then high bumps you up to a 69. Nice. then high bumps you up to a 69. Nice. then high bumps you up to a 69. Nice. Which puts you neck andneck with Fable Which puts you neck andneck with Fable Which puts you neck andneck with Fable on this benchmark which gets a 70. But on this benchmark which gets a 70. But on this benchmark which gets a 70. But the cost is $347 per task versus Fable the cost is $347 per task versus Fable the cost is $347 per task versus Fable at the same score being $13 per task. So at the same score being $13 per task. So at the same score being $13 per task. So that is super efficient. All three of that is super efficient. All three of that is super efficient. All three of these are meaningful jumps in score. You these are meaningful jumps in score. You these are meaningful jumps in score. You get from a 45% on low to a 61 on medium get from a 45% on low to a 61 on medium get from a 45% on low to a 61 on medium to a 69 on high. But then we see the to a 69 on high. But then we see the to a 69 on high. But then we see the next two levels like X high which goes next two levels like X high which goes next two levels like X high which goes from a 69 nice to a 71 less nice but from a 69 nice to a 71 less nice but from a 69 nice to a 71 less nice but also bumps the price meaningfully going also bumps the price meaningfully going also bumps the price meaningfully going from 347 to 470 per task. And while from 347 to 470 per task. And while from 347 to 470 per task. And while we're not paying API rates when we use we're not paying API rates when we use we're not paying API rates when we use it through the subscription, the API it through the subscription, the API it through the subscription, the API rates are a good measure of what you're rates are a good measure of what you're rates are a good measure of what you're burning. So yeah, that's not great. And burning. So yeah, that's not great. And burning. So yeah, that's not great. And then when we bump up to max, the price then when we bump up to max, the price then when we bump up to max, the price doubles to $8.39 doubles to $8.39 doubles to $8.39 for an additional two percentage points.
-
for an additional two percentage points. for an additional two percentage points. I don't think that's good. Medium is I don't think that's good. Medium is I don't think that's good. Medium is $1.86 and gets a 61. Max is $839 and $1.86 and gets a 61. Max is $839 and $1.86 and gets a 61. Max is $839 and gets a 73. And again, high is 347 for gets a 73. And again, high is 347 for gets a 73. And again, high is 347 for 69%. That's a more than double cost for 69%. That's a more than double cost for 69%. That's a more than double cost for a 4 percentage bump. Not worth it. Just a 4 percentage bump. Not worth it. Just a 4 percentage bump. Not worth it. Just stick with high. If you disagree, that's stick with high. If you disagree, that's stick with high. If you disagree, that's fine. I just hope you have the budget to fine. I just hope you have the budget to fine. I just hope you have the budget to handle your disagreement there. And for handle your disagreement there. And for handle your disagreement there. And for full transparency, there are other full transparency, there are other full transparency, there are other benches that have shown different benches that have shown different benches that have shown different numbers. For example, in cursor bench, numbers. For example, in cursor bench, numbers. For example, in cursor bench, high only scored a 63.5 high only scored a 63.5 high only scored a 63.5 and max scored a 67.2 and the price gap and max scored a 67.2 and the price gap and max scored a 67.2 and the price gap there was 279 to 569. So, it wasn't there was 279 to 569. So, it wasn't there was 279 to 569. So, it wasn't quite a doubling of price and it was a quite a doubling of price and it was a quite a doubling of price and it was a much more meaningful percentage bump. much more meaningful percentage bump. much more meaningful percentage bump. So, this will depend on the work you're So, this will depend on the work you're So, this will depend on the work you're doing. It's also worth noting that in doing. It's also worth noting that in doing. It's also worth noting that in Cursor Bench 32, Fable ended up scoring Cursor Bench 32, Fable ended up scoring Cursor Bench 32, Fable ended up scoring quite a bit higher than Soul did, which quite a bit higher than Soul did, which quite a bit higher than Soul did, which we'll talk more about Soul versus Fable we'll talk more about Soul versus Fable we'll talk more about Soul versus Fable in the near future. That will be a big in the near future. That will be a big in the near future. That will be a big video, too. I don't want to be video, too. I don't want to be video, too. I don't want to be sidetracked by that. The point I'm sidetracked by that. The point I'm sidetracked by that. The point I'm trying to make here is high is really trying to make here is high is really trying to make here is high is really good. And past high is when you lose good. And past high is when you lose good. And past high is when you lose this like vertical line where the cost this like vertical line where the cost this like vertical line where the cost jump isn't that big and the success jump jump isn't that big and the success jump jump isn't that big and the success jump is. I just decided to quickly check the is. I just decided to quickly check the is. I just decided to quickly check the artificial analysis intelligence versus artificial analysis intelligence versus artificial analysis intelligence versus cost index. And when I turn off cost index. And when I turn off cost index. And when I turn off everything other than 56 soul and fable everything other than 56 soul and fable everything other than 56 soul and fable 5, 56 soul high is the only thing in the 5, 56 soul high is the only thing in the 5, 56 soul high is the only thing in the good quadrant of cost to intelligence.
-
good quadrant of cost to intelligence. good quadrant of cost to intelligence. It does meaningfully go up for each It does meaningfully go up for each It does meaningfully go up for each bump. But again, 56 soul seems to be bump. But again, 56 soul seems to be bump. But again, 56 soul seems to be this really solid end of the vertical this really solid end of the vertical this really solid end of the vertical curve. So reasoning levels are now curve. So reasoning levels are now curve. So reasoning levels are now covered. Now we need to talk about what covered. Now we need to talk about what covered. Now we need to talk about what is probably the biggest behavioral is probably the biggest behavioral is probably the biggest behavioral difference with the 56 model, difference with the 56 model, difference with the 56 model, specifically both Soul and Tyra. We need specifically both Soul and Tyra. We need specifically both Soul and Tyra. We need to talk about sub agents. There's a lot to talk about sub agents. There's a lot to talk about sub agents. There's a lot of different ways to implement sub of different ways to implement sub of different ways to implement sub agents, but you generally need a tool agents, but you generally need a tool agents, but you generally need a tool that allows for your main top level that allows for your main top level that allows for your main top level agent to spin up other work for other agent to spin up other work for other agent to spin up other work for other agents to complete. I could say that agents to complete. I could say that agents to complete. I could say that Codex's implementation of sub agents Codex's implementation of sub agents Codex's implementation of sub agents isn't good, but that wouldn't be fair isn't good, but that wouldn't be fair isn't good, but that wouldn't be fair because they have two implementations of because they have two implementations of because they have two implementations of sub aents called V1 and V2, and they're sub aents called V1 and V2, and they're sub aents called V1 and V2, and they're both not good. I have a lot to say about both not good. I have a lot to say about both not good. I have a lot to say about that. Again, the Ultra video will go that. Again, the Ultra video will go that. Again, the Ultra video will go more in depth there. I will resist the more in depth there. I will resist the more in depth there. I will resist the urge to crash out hard about it. But urge to crash out hard about it. But urge to crash out hard about it. But what I will say with relatively high what I will say with relatively high what I will say with relatively high confidence is use with caution. There's confidence is use with caution. There's confidence is use with caution. There's a problem though. You might want to use a problem though. You might want to use a problem though. You might want to use sub agents with caution, but 56 was sub agents with caution, but 56 was sub agents with caution, but 56 was trained to use them with less caution. trained to use them with less caution. trained to use them with less caution. As such, it's not uncommon for 56 to As such, it's not uncommon for 56 to As such, it's not uncommon for 56 to spin up sub aents for things where it spin up sub aents for things where it spin up sub aents for things where it doesn't really need it. If you notice doesn't really need it. If you notice doesn't really need it. If you notice this happening, I don't think you should this happening, I don't think you should this happening, I don't think you should rush to go change anything about this rush to go change anything about this rush to go change anything about this just yet because the sub agents can be just yet because the sub agents can be just yet because the sub agents can be pretty good. And I'm also hoping that pretty good. And I'm also hoping that pretty good. And I'm also hoping that the Codex team fixes the issues I have the Codex team fixes the issues I have the Codex team fixes the issues I have with it as they currently stand. But if with it as they currently stand. But if with it as they currently stand. But if you are noticing limit burn after you are noticing limit burn after you are noticing limit burn after applying all of these other applying all of these other applying all of these other recommendations and you're seeing sub recommendations and you're seeing sub recommendations and you're seeing sub agents spinning up a bunch, I would agents spinning up a bunch, I would agents spinning up a bunch, I would recommend doing a slight adjustment to recommend doing a slight adjustment to recommend doing a slight adjustment to your global agents MD file. The easiest your global agents MD file. The easiest your global agents MD file. The easiest fix for now is to add the following to fix for now is to add the following to fix for now is to add the following to your agents MD. Only use sub agents if
-
your agents MD. Only use sub agents if your agents MD. Only use sub agents if the user explicitly requests them. This the user explicitly requests them. This the user explicitly requests them. This change will pretty much stop them change will pretty much stop them change will pretty much stop them entirely unless you request, and I entirely unless you request, and I entirely unless you request, and I usually request when I want them. You'll usually request when I want them. You'll usually request when I want them. You'll eventually figure out when to or not to eventually figure out when to or not to eventually figure out when to or not to use sub agents for work, like what work use sub agents for work, like what work use sub agents for work, like what work justifies splitting things up that way. justifies splitting things up that way. justifies splitting things up that way. So, don't be hesitant to experiment with So, don't be hesitant to experiment with So, don't be hesitant to experiment with them. I think they're really cool, them. I think they're really cool, them. I think they're really cool, especially when applied carefully. But especially when applied carefully. But especially when applied carefully. But right now, 56 is a little too eager. If right now, 56 is a little too eager. If right now, 56 is a little too eager. If you need to tone it down, here's the you need to tone it down, here's the you need to tone it down, here's the strategy to do such. It is worth noting strategy to do such. It is worth noting strategy to do such. It is worth noting that the sub aent implementation in that the sub aent implementation in that the sub aent implementation in other tools like claude code and cursor other tools like claude code and cursor other tools like claude code and cursor is meaningfully better. So much so that is meaningfully better. So much so that is meaningfully better. So much so that I've actually done some hacks in order I've actually done some hacks in order I've actually done some hacks in order to get cloud code working with my codec to get cloud code working with my codec to get cloud code working with my codec sub. And I've been really really happy sub. And I've been really really happy sub. And I've been really really happy with the results there. And crazy with the results there. And crazy with the results there. And crazy enough, TBO even blessed it saying that enough, TBO even blessed it saying that enough, TBO even blessed it saying that if anyone is banned for using my weird if anyone is banned for using my weird if anyone is banned for using my weird hacks to get my codec sub into cloud hacks to get my codec sub into cloud hacks to get my codec sub into cloud code that he will give you a reset. So code that he will give you a reset. So code that he will give you a reset. So yeah, this is a blessed solution. Very yeah, this is a blessed solution. Very yeah, this is a blessed solution. Very fun video on that coming in the near fun video on that coming in the near fun video on that coming in the near future. Just wanted you to know that future. Just wanted you to know that future. Just wanted you to know that it's worth exploring other options. I it's worth exploring other options. I it's worth exploring other options. I know Pi is a decent one as well. If know Pi is a decent one as well. If know Pi is a decent one as well. If you're curious enough to explore, but I you're curious enough to explore, but I you're curious enough to explore, but I do expect Codex to fix this in the not do expect Codex to fix this in the not do expect Codex to fix this in the not tooistant future. So, if you're doing it tooistant future. So, if you're doing it tooistant future. So, if you're doing it out of fear, just wait. If you're doing out of fear, just wait. If you're doing out of fear, just wait. If you're doing it out of curiosity, go nuts. One other it out of curiosity, go nuts. One other it out of curiosity, go nuts. One other thing I forgot with the reasoning level thing I forgot with the reasoning level thing I forgot with the reasoning level bit, and I'll just be super quick about bit, and I'll just be super quick about bit, and I'll just be super quick about this, is model selection. TLDDR, don't this, is model selection. TLDDR, don't this, is model selection. TLDDR, don't use Luna. It's just not for us to use use Luna. It's just not for us to use use Luna. It's just not for us to use for code. It's really useful as a thing for code. It's really useful as a thing for code. It's really useful as a thing that you access programmatically, like that you access programmatically, like that you access programmatically, like you're using it over API to filter data you're using it over API to filter data you're using it over API to filter data and whatnot. Not really what you should and whatnot. Not really what you should and whatnot. Not really what you should be using for coding. Terra seems to be a be using for coding. Terra seems to be a be using for coding. Terra seems to be a good middle ground, but honestly, I
-
good middle ground, but honestly, I good middle ground, but honestly, I would rather just use soul on low or would rather just use soul on low or would rather just use soul on low or medium. They're really efficient. I just medium. They're really efficient. I just medium. They're really efficient. I just recommend using soul high as your recommend using soul high as your recommend using soul high as your default. And if you find things that it default. And if you find things that it default. And if you find things that it goes longer than it should for, try out goes longer than it should for, try out goes longer than it should for, try out so low, maybe play with Terra. But I so low, maybe play with Terra. But I so low, maybe play with Terra. But I have been totally fine just using soul have been totally fine just using soul have been totally fine just using soul for everything. And once I made these for everything. And once I made these for everything. And once I made these other changes, I was not coming close to other changes, I was not coming close to other changes, I was not coming close to my limits. But there is one last change my limits. But there is one last change my limits. But there is one last change we have to talk about. The most we have to talk about. The most we have to talk about. The most important tip in this whole video. This important tip in this whole video. This important tip in this whole video. This is definitely the change that affected is definitely the change that affected is definitely the change that affected how Soul worked for me the most. And I'm how Soul worked for me the most. And I'm how Soul worked for me the most. And I'm really excited to share more about it really excited to share more about it really excited to share more about it right after a quick break for a sponsor. right after a quick break for a sponsor. right after a quick break for a sponsor. Today's ad's going to be a little Today's ad's going to be a little Today's ad's going to be a little different. You've already heard me talk different. You've already heard me talk different. You've already heard me talk about G2I. They help you hire worldclass about G2I. They help you hire worldclass about G2I. They help you hire worldclass engineers at whatever scale and time engineers at whatever scale and time engineers at whatever scale and time frame you need. Normally, I tell you all frame you need. Normally, I tell you all frame you need. Normally, I tell you all about how crazy it is that they can get about how crazy it is that they can get about how crazy it is that they can get you engineers in under a week that are you engineers in under a week that are you engineers in under a week that are actually experienced and ready to go. actually experienced and ready to go. actually experienced and ready to go. But instead, I'm going to tell you about But instead, I'm going to tell you about But instead, I'm going to tell you about one of their customers, Bataround, one of their customers, Bataround, one of their customers, Bataround, because they've hired nine engineers because they've hired nine engineers because they've hired nine engineers through G2I, and they learned about G2I through G2I, and they learned about G2I through G2I, and they learned about G2I through me. So, I think that's pretty through me. So, I think that's pretty through me. So, I think that's pretty cool. These guys needed to hire. And if cool. These guys needed to hire. And if cool. These guys needed to hire. And if you've ever hired, you know how grueling you've ever hired, you know how grueling you've ever hired, you know how grueling it is. You just lose all of your spare it is. You just lose all of your spare it is. You just lose all of your spare time, and no work gets done. If you're time, and no work gets done. If you're time, and no work gets done. If you're spending time hiring, you're not spending time hiring, you're not spending time hiring, you're not spending time working. And it's this spending time working. And it's this spending time working. And it's this weird unintuitive thing where when you weird unintuitive thing where when you weird unintuitive thing where when you realize you need more help, you end up realize you need more help, you end up realize you need more help, you end up needing more help sooner because you're needing more help sooner because you're needing more help sooner because you're stuck doing other stuff. and they wanted stuck doing other stuff. and they wanted stuck doing other stuff. and they wanted to hire, but they couldn't afford to to hire, but they couldn't afford to to hire, but they couldn't afford to lose bandwidth. Batteron's leadership lose bandwidth. Batteron's leadership lose bandwidth. Batteron's leadership team had decades of experience hiring team had decades of experience hiring team had decades of experience hiring engineers in other sectors, but their engineers in other sectors, but their engineers in other sectors, but their network in the media space was small. It network in the media space was small. It network in the media space was small. It was getting hard to find qualified was getting hard to find qualified was getting hard to find qualified candidates, which is why they reached candidates, which is why they reached candidates, which is why they reached out to G2I. They wanted fewer rums from out to G2I. They wanted fewer rums from out to G2I. They wanted fewer rums from better candidates that they could more better candidates that they could more better candidates that they could more easily evaluate. G2I already had a large easily evaluate. G2I already had a large easily evaluate. G2I already had a large network of engineers ready to go, and
-
network of engineers ready to go, and network of engineers ready to go, and they filtered to the best ones fit for they filtered to the best ones fit for they filtered to the best ones fit for the specific needs of Batteround in this the specific needs of Batteround in this the specific needs of Batteround in this mixed media space. With G2I, Bataround mixed media space. With G2I, Bataround mixed media space. With G2I, Bataround has contracted nine engineers, had a 90% has contracted nine engineers, had a 90% has contracted nine engineers, had a 90% success rate, reduced leadership time success rate, reduced leadership time success rate, reduced leadership time spent screening candidates, integrated spent screening candidates, integrated spent screening candidates, integrated engineers directly through their engineers directly through their engineers directly through their existing Slack workflows, and they've existing Slack workflows, and they've existing Slack workflows, and they've minimized disruption through the clean minimized disruption through the clean minimized disruption through the clean transitions and fast back fills when transitions and fast back fills when transitions and fast back fills when they need to shift around their staff. they need to shift around their staff. they need to shift around their staff. Battle round themselves said that a 50% Battle round themselves said that a 50% Battle round themselves said that a 50% success rate would feel like a success rate would feel like a success rate would feel like a superpower to them, and now they're up superpower to them, and now they're up superpower to them, and now they're up to 90. Waste less time and hire better to 90. Waste less time and hire better to 90. Waste less time and hire better engineers at soy.link/g2i. engineers at soy.link/g2i. engineers at soy.link/g2i. It is worth noting that a lot of It is worth noting that a lot of It is worth noting that a lot of analysis is showing that Terra isn't analysis is showing that Terra isn't analysis is showing that Terra isn't really a great option, which is really a great option, which is really a great option, which is disappointing. I was quite excited about disappointing. I was quite excited about disappointing. I was quite excited about it. Artificial analysis said the it. Artificial analysis said the it. Artificial analysis said the following. 56 Sol Luna are ahead of following. 56 Sol Luna are ahead of following. 56 Sol Luna are ahead of Terra at every point on the intelligence Terra at every point on the intelligence Terra at every point on the intelligence versus cost per task chart. This is the versus cost per task chart. This is the versus cost per task chart. This is the chart that I had just shown before. And chart that I had just shown before. And chart that I had just shown before. And here you can see that Luna on max scores here you can see that Luna on max scores here you can see that Luna on max scores slightly better than Soul on low while slightly better than Soul on low while slightly better than Soul on low while being just barely more expensive. It being just barely more expensive. It being just barely more expensive. It almost looks like a pretty consistent almost looks like a pretty consistent almost looks like a pretty consistent line from Luna, low, mid, high, XH high, line from Luna, low, mid, high, XH high, line from Luna, low, mid, high, XH high, and then solo, low, mid, high, XH high.
-
and then solo, low, mid, high, XH high. and then solo, low, mid, high, XH high. But Terra underneath here not looking But Terra underneath here not looking But Terra underneath here not looking great. The model needs to know when to great. The model needs to know when to great. The model needs to know when to stop. As I was hinting at earlier, 56 is stop. As I was hinting at earlier, 56 is stop. As I was hinting at earlier, 56 is very eager to get work done. It will very eager to get work done. It will very eager to get work done. It will keep going as far as it possibly can keep going as far as it possibly can keep going as far as it possibly can unless you give it a reason to stop unless you give it a reason to stop unless you give it a reason to stop going. I gave some good examples in my going. I gave some good examples in my going. I gave some good examples in my article, so I'm going to reuse those article, so I'm going to reuse those article, so I'm going to reuse those here. This is the style of prompts I've here. This is the style of prompts I've here. This is the style of prompts I've learned to write when using Soul because learned to write when using Soul because learned to write when using Soul because without it, it can go further than I without it, it can go further than I without it, it can go further than I want it to pretty often. Example one is want it to pretty often. Example one is want it to pretty often. Example one is the following. I want you to build this the following. I want you to build this the following. I want you to build this new feature. Start by writing a plan. new feature. Start by writing a plan. new feature. Start by writing a plan. When you finish the plan, stop and ask When you finish the plan, stop and ask When you finish the plan, stop and ask for feedback before proceeding. I have for feedback before proceeding. I have for feedback before proceeding. I have told the model explicitly, do this thing told the model explicitly, do this thing told the model explicitly, do this thing and be done at this point. That doesn't and be done at this point. That doesn't and be done at this point. That doesn't mean you have to give the model prompts mean you have to give the model prompts mean you have to give the model prompts that are less work though. You can have that are less work though. You can have that are less work though. You can have the stop point be way, way, way further the stop point be way, way, way further the stop point be way, way, way further on. For example, the plan looks great. on. For example, the plan looks great. on. For example, the plan looks great. Let's build it out. Use computer use to Let's build it out. Use computer use to Let's build it out. Use computer use to test your implementation. Keep going test your implementation. Keep going test your implementation. Keep going until the code works and you're happy until the code works and you're happy until the code works and you're happy with the implementation. Put up a PR, with the implementation. Put up a PR, with the implementation. Put up a PR, babysit it for the first set of review babysit it for the first set of review babysit it for the first set of review comments, and address them. Stop after comments, and address them. Stop after comments, and address them. Stop after the first set of review comments. I'll the first set of review comments. I'll the first set of review comments. I'll handle it from there. Previously with handle it from there. Previously with handle it from there. Previously with 55, the model stopped way too often.
-
55, the model stopped way too often. 55, the model stopped way too often. With 56, it's so unlikely to stop that I With 56, it's so unlikely to stop that I With 56, it's so unlikely to stop that I feel like I have to put up the stop feel like I have to put up the stop feel like I have to put up the stop signs myself. But now that I've been signs myself. But now that I've been signs myself. But now that I've been doing that, it feels much better. It's doing that, it feels much better. It's doing that, it feels much better. It's nice that I can start building my own nice that I can start building my own nice that I can start building my own intuition for how far the model can go intuition for how far the model can go intuition for how far the model can go without my intervention and also that I without my intervention and also that I without my intervention and also that I can kind of choose when to stop it can kind of choose when to stop it can kind of choose when to stop it rather than trying to do this through rather than trying to do this through rather than trying to do this through reasoning levels or changes to your reasoning levels or changes to your reasoning levels or changes to your harness. So much of this can just be harness. So much of this can just be harness. So much of this can just be done via prompting. And I will say I am done via prompting. And I will say I am done via prompting. And I will say I am proud of myself for getting this far in proud of myself for getting this far in proud of myself for getting this far in the video without my instructions being the video without my instructions being the video without my instructions being prompt better because your prompt can prompt better because your prompt can prompt better because your prompt can absolutely cost you more money, but absolutely cost you more money, but absolutely cost you more money, but there's other things that I thought were there's other things that I thought were there's other things that I thought were more important. This piece though, this more important. This piece though, this more important. This piece though, this can change how you use the model can change how you use the model can change how you use the model fundamentally. When you realize that the fundamentally. When you realize that the fundamentally. When you realize that the end is a thing you can define in the end is a thing you can define in the end is a thing you can define in the prompt rather than a thing you have to prompt rather than a thing you have to prompt rather than a thing you have to set up with tools, you can let the model set up with tools, you can let the model set up with tools, you can let the model go way further. And the challenge I go way further. And the challenge I go way further. And the challenge I would present to you is see how far off would present to you is see how far off would present to you is see how far off you can make that stop sign without the you can make that stop sign without the you can make that stop sign without the results in the quality going down. I results in the quality going down. I results in the quality going down. I didn't talk about this one too much didn't talk about this one too much didn't talk about this one too much here, but it is really fun. You can let here, but it is really fun. You can let here, but it is really fun. You can let other agents steer. For example, I set other agents steer. For example, I set other agents steer. For example, I set up my Claude code to allow for it to up my Claude code to allow for it to up my Claude code to allow for it to call Codeex. I've went much further with call Codeex. I've went much further with call Codeex. I've went much further with this since, and that's the Codeex in this since, and that's the Codeex in this since, and that's the Codeex in Claude Code video coming soon. We're Claude Code video coming soon. We're Claude Code video coming soon. We're actually using Soul as the model in actually using Soul as the model in actually using Soul as the model in Claude Code. Seriously, Claudeex is way Claude Code. Seriously, Claudeex is way Claude Code. Seriously, Claudeex is way better than I expected it to be. And as better than I expected it to be. And as better than I expected it to be. And as silly as it is seeing this 56 soul silly as it is seeing this 56 soul silly as it is seeing this 56 soul there, it's been working great for me.
-
there, it's been working great for me. there, it's been working great for me. But most importantly here, the thing I But most importantly here, the thing I But most importantly here, the thing I want to encourage is more want to encourage is more want to encourage is more experimentation. You shouldn't be trying experimentation. You shouldn't be trying experimentation. You shouldn't be trying to copy my skills, my agent MD, and my to copy my skills, my agent MD, and my to copy my skills, my agent MD, and my prompts directly because this should be prompts directly because this should be prompts directly because this should be different for everyone. You should different for everyone. You should different for everyone. You should experiment more with these different experiment more with these different experiment more with these different things. Try prompting different ways. things. Try prompting different ways. things. Try prompting different ways. Try different reasoning levels. Try Try different reasoning levels. Try Try different reasoning levels. Try adjusting your codeex agents MD and your adjusting your codeex agents MD and your adjusting your codeex agents MD and your Claude MD files a bit to see how it Claude MD files a bit to see how it Claude MD files a bit to see how it changes behavior. You should be spending changes behavior. You should be spending changes behavior. You should be spending a good bit of time in the docex andcloud a good bit of time in the docex andcloud a good bit of time in the docex andcloud directories on your computer if you're directories on your computer if you're directories on your computer if you're using these agents and models. It's so using these agents and models. It's so using these agents and models. It's so fun to customize because it's not like fun to customize because it's not like fun to customize because it's not like you have to be deep in the weeds of the you have to be deep in the weeds of the you have to be deep in the weeds of the details of how it works and all these details of how it works and all these details of how it works and all these config files. It's a markdown file and config files. It's a markdown file and config files. It's a markdown file and changes to that can fundamentally change changes to that can fundamentally change changes to that can fundamentally change the experience that you're having. And the experience that you're having. And the experience that you're having. And the more you move things around, the the more you move things around, the the more you move things around, the more you channel our inner Dan Abramov more you channel our inner Dan Abramov more you channel our inner Dan Abramov to experiment and shift until things to experiment and shift until things to experiment and shift until things feel the way you want them to and have feel the way you want them to and have feel the way you want them to and have your own feelings change in the process, your own feelings change in the process, your own feelings change in the process, that's what's going to give you the best that's what's going to give you the best that's what's going to give you the best experience using these things. If you experience using these things. If you experience using these things. If you blindly follow my recommendations or you blindly follow my recommendations or you blindly follow my recommendations or you go set up somebody else's go set up somebody else's go set up somebody else's recommendations, like you install any of recommendations, like you install any of recommendations, like you install any of the oh my whatevers, you're never going the oh my whatevers, you're never going the oh my whatevers, you're never going to learn what does and doesn't work and to learn what does and doesn't work and to learn what does and doesn't work and you're going to be limited by the you're going to be limited by the you're going to be limited by the creativity of someone else. Go play with creativity of someone else. Go play with creativity of someone else. Go play with it yourself. I don't have a bunch of it yourself. I don't have a bunch of it yourself. I don't have a bunch of other people's skills and things other people's skills and things other people's skills and things installed. I just dick around and change installed. I just dick around and change installed. I just dick around and change things and read the traces from my things and read the traces from my things and read the traces from my agents when they don't do what I expect agents when they don't do what I expect agents when they don't do what I expect and make adjustments until they do. I am and make adjustments until they do. I am and make adjustments until they do. I am currently deep in the process of currently deep in the process of currently deep in the process of rewriting my agents MD. And this is all rewriting my agents MD. And this is all rewriting my agents MD. And this is all by hand. Not a line of this was LM
-
by hand. Not a line of this was LM by hand. Not a line of this was LM generated. And I think that is one of generated. And I think that is one of generated. And I think that is one of the best places to take the time because the best places to take the time because the best places to take the time because you'll be amazed how much these things you'll be amazed how much these things you'll be amazed how much these things can change how the model feels to use. can change how the model feels to use. can change how the model feels to use. What I'm trying to say is that you need What I'm trying to say is that you need What I'm trying to say is that you need to stop looking for someone else's to stop looking for someone else's to stop looking for someone else's solutions on a shelf for you to buy, solutions on a shelf for you to buy, solutions on a shelf for you to buy, copy, paste, and hope for the best. You copy, paste, and hope for the best. You copy, paste, and hope for the best. You should use resources like this video and should use resources like this video and should use resources like this video and like these other examples people have as like these other examples people have as like these other examples people have as inspiration to get it the way you want inspiration to get it the way you want inspiration to get it the way you want yourself. I mentioned at the start that yourself. I mentioned at the start that yourself. I mentioned at the start that I've been seeing some bad advice. This I've been seeing some bad advice. This I've been seeing some bad advice. This is probably the worst one. While I don't is probably the worst one. While I don't is probably the worst one. While I don't trust OpenAI to configure everything trust OpenAI to configure everything trust OpenAI to configure everything perfectly in codeex for us, they perfectly in codeex for us, they perfectly in codeex for us, they absolutely covered the context windows absolutely covered the context windows absolutely covered the context windows themselves just fine. I've seen a lot of themselves just fine. I've seen a lot of themselves just fine. I've seen a lot of people making recommendations like this, people making recommendations like this, people making recommendations like this, specifying that you should limit the specifying that you should limit the specifying that you should limit the context window and when it autocompacts context window and when it autocompacts context window and when it autocompacts manually inside of your config. Don't do manually inside of your config. Don't do manually inside of your config. Don't do this. The model was trained on specific this. The model was trained on specific this. The model was trained on specific levels of compaction in tools like levels of compaction in tools like levels of compaction in tools like codecs. This is going to make the model codecs. This is going to make the model codecs. This is going to make the model dumber and it's not going to save you dumber and it's not going to save you dumber and it's not going to save you any money. If anything, it's going to any money. If anything, it's going to any money. If anything, it's going to cause compaction to happen more than it cause compaction to happen more than it cause compaction to happen more than it should and compaction is expensive. Tibo should and compaction is expensive. Tibo should and compaction is expensive. Tibo jumped in and said as much. This is not jumped in and said as much. This is not jumped in and said as much. This is not correct. Do not do this if you do not correct. Do not do this if you do not correct. Do not do this if you do not understand exactly what you're doing. We understand exactly what you're doing. We understand exactly what you're doing. We do not charge extra above 270K context do not charge extra above 270K context do not charge extra above 270K context and the context threshold has been tuned and the context threshold has been tuned and the context threshold has been tuned for 56 soul to be perfect with a default for 56 soul to be perfect with a default for 56 soul to be perfect with a default limit. It is trust him. Don't just limit. It is trust him. Don't just limit. It is trust him. Don't just blindly follow random advice on Twitter.
-
blindly follow random advice on Twitter. blindly follow random advice on Twitter. At the very least, make sure the people At the very least, make sure the people At the very least, make sure the people giving the advice aren't getting their giving the advice aren't getting their giving the advice aren't getting their advice countered by people at OpenAI. advice countered by people at OpenAI. advice countered by people at OpenAI. I hate filming live sometimes because I hate filming live sometimes because I hate filming live sometimes because literally right after I finished that literally right after I finished that literally right after I finished that part and was ready to go offline, Tibo part and was ready to go offline, Tibo part and was ready to go offline, Tibo tweeted the following. We've landed tweeted the following. We've landed tweeted the following. We've landed inference optimizations to make it inference optimizations to make it inference optimizations to make it cheaper to run. It's going to be 10% cheaper to run. It's going to be 10% cheaper to run. It's going to be 10% savings. Cool. Awesome. They also savings. Cool. Awesome. They also savings. Cool. Awesome. They also noticed that changing the context size noticed that changing the context size noticed that changing the context size limit in the product to 372K up from limit in the product to 372K up from limit in the product to 372K up from 272K resulted in more usage being 272K resulted in more usage being 272K resulted in more usage being charged than intended. This means that charged than intended. This means that charged than intended. This means that their usage tracking was broken even their usage tracking was broken even their usage tracking was broken even though they planned to let this stay. though they planned to let this stay. though they planned to let this stay. They've temporarily reverted to 272K and They've temporarily reverted to 272K and They've temporarily reverted to 272K and will roll back in the days to come. This will roll back in the days to come. This will roll back in the days to come. This should be a big change in how fast your should be a big change in how fast your should be a big change in how fast your usage drains. Cool. Awesome. Stupid that usage drains. Cool. Awesome. Stupid that usage drains. Cool. Awesome. Stupid that we're here, but here we are. They we're here, but here we are. They we're here, but here we are. They confirmed that the leaks around juice confirmed that the leaks around juice confirmed that the leaks around juice values being changed were true, but values being changed were true, but values being changed were true, but they've been reverted since. And they they've been reverted since. And they they've been reverted since. And they also call out that there's been more use also call out that there's been more use also call out that there's been more use of multi- aent than intended in high and of multi- aent than intended in high and of multi- aent than intended in high and x high reasoning efforts and they're x high reasoning efforts and they're x high reasoning efforts and they're fixing it going forward. Also fixing fixing it going forward. Also fixing fixing it going forward. Also fixing small other things they notice with auto small other things they notice with auto small other things they notice with auto review where they can be more efficient. review where they can be more efficient. review where they can be more efficient. Reporting on the fly is annoying, but Reporting on the fly is annoying, but Reporting on the fly is annoying, but you now have the additional context.
-
you now have the additional context. you now have the additional context. That all said, every piece of advice I That all said, every piece of advice I That all said, every piece of advice I gave in this is still true. So don't gave in this is still true. So don't gave in this is still true. So don't think this means my advice is invalid. think this means my advice is invalid. think this means my advice is invalid. Just know that this is a moving target. Just know that this is a moving target. Just know that this is a moving target. Best practices are still best practices. Best practices are still best practices. Best practices are still best practices. One more point in favor of lower One more point in favor of lower One more point in favor of lower reasoning levels. The Open Code folks reasoning levels. The Open Code folks reasoning levels. The Open Code folks are obsessed with 56. They barely even are obsessed with 56. They barely even are obsessed with 56. They barely even use Fable. They like 56 so much. They use Fable. They like 56 so much. They use Fable. They like 56 so much. They like it so much that most of OpenAI has like it so much that most of OpenAI has like it so much that most of OpenAI has been promoting their posts throughout. I been promoting their posts throughout. I been promoting their posts throughout. I assume they were probably using it on assume they were probably using it on assume they were probably using it on higher reasoning levels because who higher reasoning levels because who higher reasoning levels because who wouldn't initially hit an issue though, wouldn't initially hit an issue though, wouldn't initially hit an issue though, one that I've hit when I was testing one that I've hit when I was testing one that I've hit when I was testing these things over API. They changed the these things over API. They changed the these things over API. They changed the key in the JSON for reasoning levels at key in the JSON for reasoning levels at key in the JSON for reasoning levels at some point and a lot of tools don't pass some point and a lot of tools don't pass some point and a lot of tools don't pass through reasoning level properly. I ran through reasoning level properly. I ran through reasoning level properly. I ran into this a lot during benchmarking. into this a lot during benchmarking. into this a lot during benchmarking. Turns out Open Code did too. And during Turns out Open Code did too. And during Turns out Open Code did too. And during all of their usage, a month of usage, all of their usage, a month of usage, all of their usage, a month of usage, they realized that they had it they realized that they had it they realized that they had it misconfigured to always use medium, even misconfigured to always use medium, even misconfigured to always use medium, even though they thought they were using XH though they thought they were using XH though they thought they were using XH high or other levels. And despite that, high or other levels. And despite that, high or other levels. And despite that, the whole team agreed it was their the whole team agreed it was their the whole team agreed it was their favorite model. Definitely an favorite model. Definitely an favorite model. Definitely an endorsement for medium as the default. endorsement for medium as the default. endorsement for medium as the default. So if you do like Open Code and or you So if you do like Open Code and or you So if you do like Open Code and or you trust the Open Code team, they have trust the Open Code team, they have trust the Open Code team, they have confirmed medium is incredible. Medium confirmed medium is incredible. Medium confirmed medium is incredible. Medium and high are both really, really good and high are both really, really good and high are both really, really good options. So, go play, go configure, go options. So, go play, go configure, go options. So, go play, go configure, go customize, and see what you can do. I customize, and see what you can do. I customize, and see what you can do. I have a feeling you'll be surprised. Keep have a feeling you'll be surprised. Keep have a feeling you'll be surprised. Keep playing and keep prompting.
Summary
The main topic is the high token consumption of GBT 56, contrasted with older models like GBT 55. The transcript highlights the struggle to stay within rate limits despite adjusting settings. The practical takeaway is that users need to optimize their prompts and usage to conserve tokens and maximize their subscription benefits, as OpenAI's current implementation is inefficient.