← Back
Theo September 18, 2026 27m

Please stop using stupid models

Not yet indexed — Search & Ask will be available once this episode finishes processing.

Read full transcript 20 segments
  1. If you enjoy the videos where everyone If you enjoy the videos where everyone calls me a paid shill, you're going to calls me a paid shill, you're going to calls me a paid shill, you're going to love this one because I kind of have to love this one because I kind of have to love this one because I kind of have to glaze for a second because I see some glaze for a second because I see some glaze for a second because I see some comments that are just so dumb that it comments that are just so dumb that it comments that are just so dumb that it makes me realize the majority of makes me realize the majority of makes me realize the majority of engineers just aren't using agents right engineers just aren't using agents right engineers just aren't using agents right at all. And I mean that when I say it. at all. And I mean that when I say it. at all. And I mean that when I say it. I'm inspired to make this video because I'm inspired to make this video because I'm inspired to make this video because of a post from another engineer that I of a post from another engineer that I of a post from another engineer that I generally quite respect, but is so dumb generally quite respect, but is so dumb generally quite respect, but is so dumb that I need to talk to you guys about it that I need to talk to you guys about it that I need to talk to you guys about it because I know a lot of y'all feel the because I know a lot of y'all feel the because I know a lot of y'all feel the same way and you feel it so deeply that same way and you feel it so deeply that same way and you feel it so deeply that you accuse me of being a shill because you accuse me of being a shill because you accuse me of being a shill because of how deep my feelings here go. The of how deep my feelings here go. The of how deep my feelings here go. The take is as follows. from David Kramer take is as follows. from David Kramer take is as follows. from David Kramer aka Zeg, the founder of Sentry. Everyone aka Zeg, the founder of Sentry. Everyone aka Zeg, the founder of Sentry. Everyone using Fable or Astra, which is most of using Fable or Astra, which is most of using Fable or Astra, which is most of us on subs because no one can afford it, us on subs because no one can afford it, us on subs because no one can afford it, should try switching back to the other should try switching back to the other should try switching back to the other high reasoning models like Opus and high reasoning models like Opus and high reasoning models like Opus and Soul. You'll likely realize your tasks Soul. You'll likely realize your tasks Soul. You'll likely realize your tasks don't perform any differently. To which don't perform any differently. To which don't perform any differently. To which I responded, this is your worst take of I responded, this is your worst take of I responded, this is your worst take of all time, which is kind of crazy. all time, which is kind of crazy. all time, which is kind of crazy. Kramer's had some wild ones in the past, Kramer's had some wild ones in the past, Kramer's had some wild ones in the past, and I genuinely believe you suck at and I genuinely believe you suck at and I genuinely believe you suck at prompting if you believe this. So, if prompting if you believe this. So, if prompting if you believe this. So, if what Kramer just said resonates with what Kramer just said resonates with what Kramer just said resonates with you, if you actually think or have you, if you actually think or have you, if you actually think or have experienced going back a model gen or experienced going back a model gen or experienced going back a model gen or down a model tier and not notice a down a model tier and not notice a down a model tier and not notice a difference in the way you do work, you difference in the way you do work, you difference in the way you do work, you also suck at prompting. And this isn't also suck at prompting. And this isn't also suck at prompting. And this isn't like a disprovable thing the other way like a disprovable thing the other way like a disprovable thing the other way where you can't prove I'm wrong, cuz where you can't prove I'm wrong, cuz where you can't prove I'm wrong, cuz there is literally no way to do that.

  2. there is literally no way to do that. there is literally no way to do that. But I can very easily prove you're But I can very easily prove you're But I can very easily prove you're wrong, which is what I'm really excited wrong, which is what I'm really excited wrong, which is what I'm really excited to do right after a quick break for to do right after a quick break for to do right after a quick break for today's sponsor. I have a challenge for today's sponsor. I have a challenge for today's sponsor. I have a challenge for you. Next time you're reviewing a big PR you. Next time you're reviewing a big PR you. Next time you're reviewing a big PR and you're not sure if it's ready to go and you're not sure if it's ready to go and you're not sure if it's ready to go or not, ask your agent on your machine or not, ask your agent on your machine or not, ask your agent on your machine to pull it down, play with it, test it, to pull it down, play with it, test it, to pull it down, play with it, test it, and make sure everything works as and make sure everything works as and make sure everything works as expected, there's a good chance it's expected, there's a good chance it's expected, there's a good chance it's going to find things that both you and going to find things that both you and going to find things that both you and your agents wouldn't have otherwise. your agents wouldn't have otherwise. your agents wouldn't have otherwise. Because if you're just looking at the Because if you're just looking at the Because if you're just looking at the code, you're not going to be able to code, you're not going to be able to code, you're not going to be able to find too much. This is a mistake that a find too much. This is a mistake that a find too much. This is a mistake that a lot of the AI code review bots make. lot of the AI code review bots make. lot of the AI code review bots make. They think they can know everything and They think they can know everything and They think they can know everything and how it works just by reading the code. how it works just by reading the code. how it works just by reading the code. And reality is not that simple. Code And reality is not that simple. Code And reality is not that simple. Code breaks in ways that are not clearly breaks in ways that are not clearly breaks in ways that are not clearly visible just from reading code. This is visible just from reading code. This is visible just from reading code. This is why I think T-Rex by Grappile is so damn why I think T-Rex by Grappile is so damn why I think T-Rex by Grappile is so damn cool. These guys realize that agents cool. These guys realize that agents cool. These guys realize that agents know a lot more about code when they can know a lot more about code when they can know a lot more about code when they can actually run it than they can possibly actually run it than they can possibly actually run it than they can possibly guess by just staring at it. And that's guess by just staring at it. And that's guess by just staring at it. And that's why T-Rex uses sandboxes to actually why T-Rex uses sandboxes to actually why T-Rex uses sandboxes to actually test your changes before leaving review test your changes before leaving review test your changes before leaving review comments. Grapile already has a comments. Grapile already has a comments. Grapile already has a best-in-class system for understanding best-in-class system for understanding best-in-class system for understanding the context of changes because they know the context of changes because they know the context of changes because they know your whole codebase. They monitor it your whole codebase. They monitor it your whole codebase. They monitor it closely and they've indexed the hell out closely and they've indexed the hell out closely and they've indexed the hell out of it in order to make good insights of it in order to make good insights of it in order to make good insights happen. But now they can also test the happen. But now they can also test the happen. But now they can also test the code which makes it so much more code which makes it so much more code which makes it so much more powerful. And this goes so much further powerful. And this goes so much further powerful. And this goes so much further than just letting the reviewer run in a than just letting the reviewer run in a than just letting the reviewer run in a sandbox. It's actually kind of the sandbox. It's actually kind of the sandbox. It's actually kind of the opposite. It's more that the reviewer opposite. It's more that the reviewer opposite. It's more that the reviewer orchestrator can spin up and run orchestrator can spin up and run orchestrator can spin up and run sandboxes with sub agents for whatever sandboxes with sub agents for whatever sandboxes with sub agents for whatever theories it has about what might be theories it has about what might be theories it has about what might be wrong. Some reviews might need no wrong. Some reviews might need no wrong. Some reviews might need no sandboxes, some reviews might need 10.

  3. sandboxes, some reviews might need 10. sandboxes, some reviews might need 10. The orchestrator will figure out what is The orchestrator will figure out what is The orchestrator will figure out what is needed to verify your changes. This needed to verify your changes. This needed to verify your changes. This helps you ship with way more confidence, helps you ship with way more confidence, helps you ship with way more confidence, and not just because it gives you a and not just because it gives you a and not just because it gives you a thumbs up or thumbs down, but because it thumbs up or thumbs down, but because it thumbs up or thumbs down, but because it will respond with images and videos of will respond with images and videos of will respond with images and videos of the things it tests. So, if you want to the things it tests. So, if you want to the things it tests. So, if you want to make sure that your dropdown disappears make sure that your dropdown disappears make sure that your dropdown disappears correctly or that the signup flow works correctly or that the signup flow works correctly or that the signup flow works end to end, a simple approval message end to end, a simple approval message end to end, a simple approval message should not be enough. A video proving should not be enough. A video proving should not be enough. A video proving the changes worked is so much more the changes worked is so much more the changes worked is so much more valuable if you're trying to ship with valuable if you're trying to ship with valuable if you're trying to ship with confidence. We all know coding agents confidence. We all know coding agents confidence. We all know coding agents work way better if they have a real work way better if they have a real work way better if they have a real computer. Turns out review agents do as computer. Turns out review agents do as computer. Turns out review agents do as well. Get your review agents the support well. Get your review agents the support well. Get your review agents the support that they need at soy.link/grapile. that they need at soy.link/grapile. that they need at soy.link/grapile. Let's just break this down piece by Let's just break this down piece by Let's just break this down piece by piece because if I do the whole thing at piece because if I do the whole thing at piece because if I do the whole thing at once, people are going to read into the once, people are going to read into the once, people are going to read into the parts I don't focus on and say that's parts I don't focus on and say that's parts I don't focus on and say that's why I'm wrong. So, let's start with a why I'm wrong. So, let's start with a why I'm wrong. So, let's start with a thing that I just want to get out there thing that I just want to get out there thing that I just want to get out there right now. The Astra and Soul portion is right now. The Astra and Soul portion is right now. The Astra and Soul portion is very different from the Fable versus very different from the Fable versus very different from the Fable versus Opus and Soul portion here. Astra can do Opus and Soul portion here. Astra can do Opus and Soul portion here. Astra can do things no other model in history could. things no other model in history could. things no other model in history could. It is incredible, but it can also screw It is incredible, but it can also screw It is incredible, but it can also screw up in ways I haven't seen since I last up in ways I haven't seen since I last up in ways I haven't seen since I last was using a Gemini model seriously, was using a Gemini model seriously, was using a Gemini model seriously, which was like 2025. So, if your feeling which was like 2025. So, if your feeling which was like 2025. So, if your feeling here is purely around Astra, because here is purely around Astra, because here is purely around Astra, because yes, sometimes it does incredible yes, sometimes it does incredible yes, sometimes it does incredible things, but sometimes it does weird or things, but sometimes it does weird or things, but sometimes it does weird or annoying things, then I see why going annoying things, then I see why going annoying things, then I see why going back to soul would feel not even not back to soul would feel not even not back to soul would feel not even not bad, but in some real cases somewhat bad, but in some real cases somewhat bad, but in some real cases somewhat good. And I understand why people will good. And I understand why people will good. And I understand why people will be moving from Astra to soul. So, I'm be moving from Astra to soul. So, I'm be moving from Astra to soul. So, I'm jumping in that one immediately because jumping in that one immediately because jumping in that one immediately because I know people will be feeling that and I I know people will be feeling that and I I know people will be feeling that and I get it. I even have a diagram. Let me get it. I even have a diagram. Let me get it. I even have a diagram. Let me find it. Here it is. This diagram was find it. Here it is. This diagram was find it. Here it is. This diagram was meant to show roughly how I feel about

  4. meant to show roughly how I feel about meant to show roughly how I feel about the quality of responses over a large the quality of responses over a large the quality of responses over a large set of responses with Fable and Astra. set of responses with Fable and Astra. set of responses with Fable and Astra. Astra at its best can do things Fable Astra at its best can do things Fable Astra at its best can do things Fable never could. Astra at its worst makes me never could. Astra at its worst makes me never could. Astra at its worst makes me question why I'm using AI to code at all question why I'm using AI to code at all question why I'm using AI to code at all because it can make some real [ __ ] because it can make some real [ __ ] because it can make some real [ __ ] dumb decisions and assumptions. dumb decisions and assumptions. dumb decisions and assumptions. Generally speaking though, one of the Generally speaking though, one of the Generally speaking though, one of the benefits you get from these frontier benefits you get from these frontier benefits you get from these frontier models is that the gap between the worst models is that the gap between the worst models is that the gap between the worst and the best gets closed and the floor and the best gets closed and the floor and the best gets closed and the floor goes up. Actually think this diagram is goes up. Actually think this diagram is goes up. Actually think this diagram is a really good starting point for the a really good starting point for the a really good starting point for the things I wanted to try and communicate things I wanted to try and communicate things I wanted to try and communicate here. Generally speaking, y'all focus here. Generally speaking, y'all focus here. Generally speaking, y'all focus too much on the high points with models too much on the high points with models too much on the high points with models and not enough on the lows. And I'll be and not enough on the lows. And I'll be and not enough on the lows. And I'll be real, both of these assumptions have real, both of these assumptions have real, both of these assumptions have gotten me in trouble. There have been gotten me in trouble. There have been gotten me in trouble. There have been times where I focused too much on what times where I focused too much on what times where I focused too much on what the best looked like for a model, things the best looked like for a model, things the best looked like for a model, things like Opus 5, and I thought it was really like Opus 5, and I thought it was really like Opus 5, and I thought it was really good when it wasn't, and it misbehaved good when it wasn't, and it misbehaved good when it wasn't, and it misbehaved far more often than I had known at the far more often than I had known at the far more often than I had known at the time. And then there's times where I time. And then there's times where I time. And then there's times where I focus on how pathetic the model behaves focus on how pathetic the model behaves focus on how pathetic the model behaves sometimes, and just ignore it outright sometimes, and just ignore it outright sometimes, and just ignore it outright after that. Even though there are real after that. Even though there are real after that. Even though there are real strengths, I would argue to an extent strengths, I would argue to an extent strengths, I would argue to an extent that my disdain towards Gemini comes that my disdain towards Gemini comes that my disdain towards Gemini comes from the floor being so low and so from the floor being so low and so from the floor being so low and so consistent where it just it'll read the consistent where it just it'll read the consistent where it just it'll read the same file 26 times before making a same file 26 times before making a same file 26 times before making a change because it's a [ __ ] model. But change because it's a [ __ ] model. But change because it's a [ __ ] model. But sometimes it can name skateboard tricks sometimes it can name skateboard tricks sometimes it can name skateboard tricks really well. So it ceiling is in really well. So it ceiling is in really well. So it ceiling is in interesting places. But you got to think interesting places. But you got to think interesting places. But you got to think about the ceiling and the floor. And we about the ceiling and the floor. And we about the ceiling and the floor. And we all have a bad habit of thinking too all have a bad habit of thinking too all have a bad habit of thinking too much about the ceiling. I bring this up much about the ceiling. I bring this up much about the ceiling. I bring this up because the first huge benefit from this because the first huge benefit from this because the first huge benefit from this new era of models and from frontier new era of models and from frontier new era of models and from frontier stuff. The thing that you don't get as stuff. The thing that you don't get as stuff. The thing that you don't get as much from open weight models and from

  5. much from open weight models and from much from open weight models and from other labs other than anthropomic and other labs other than anthropomic and other labs other than anthropomic and open AI generally is a higher floor. And open AI generally is a higher floor. And open AI generally is a higher floor. And I personally find that raising the floor I personally find that raising the floor I personally find that raising the floor is way more beneficial than raising the is way more beneficial than raising the is way more beneficial than raising the ceiling. I don't care if the model ceiling. I don't care if the model ceiling. I don't care if the model solves novel math problems if it doesn't solves novel math problems if it doesn't solves novel math problems if it doesn't know what I mean when I say revert. Yes, know what I mean when I say revert. Yes, know what I mean when I say revert. Yes, Astra has actually been confused about Astra has actually been confused about Astra has actually been confused about what the word revert meant for me what the word revert meant for me what the word revert meant for me before. I don't care how good an agent before. I don't care how good an agent before. I don't care how good an agent is at building a 3D environment in is at building a 3D environment in is at building a 3D environment in Blender if it can't center an icon in a Blender if it can't center an icon in a Blender if it can't center an icon in a div, which yes, I have also seen Astra div, which yes, I have also seen Astra div, which yes, I have also seen Astra fail to do. Astra has kind of eroded the fail to do. Astra has kind of eroded the fail to do. Astra has kind of eroded the conversation that I want to have here, conversation that I want to have here, conversation that I want to have here, which is why I'm choosing to do it now which is why I'm choosing to do it now which is why I'm choosing to do it now because Astra has the traits of the best because Astra has the traits of the best because Astra has the traits of the best models and some of the traits of the models and some of the traits of the models and some of the traits of the worst, which makes it hard to recommend worst, which makes it hard to recommend worst, which makes it hard to recommend in the way I want to. So, I'm going to in the way I want to. So, I'm going to in the way I want to. So, I'm going to do a thing I don't want to do. do a thing I don't want to do. do a thing I don't want to do. I'm going to clone this diagram because I'm going to clone this diagram because I'm going to clone this diagram because I'm going to delete Astra from it and I'm going to delete Astra from it and I'm going to delete Astra from it and I'm going to reabel it to what it is, I'm going to reabel it to what it is, I'm going to reabel it to what it is, Gemini 3.8 Hey, Flash. This is how I Gemini 3.8 Hey, Flash. This is how I Gemini 3.8 Hey, Flash. This is how I actually feel. And the fact that Astra actually feel. And the fact that Astra actually feel. And the fact that Astra can perform as poorly as Flash ever is can perform as poorly as Flash ever is can perform as poorly as Flash ever is pathetic. And the engineers involved pathetic. And the engineers involved pathetic. And the engineers involved should feel bad and fix it. And should feel bad and fix it. And should feel bad and fix it. And thankfully, they do. I know for a fact thankfully, they do. I know for a fact thankfully, they do. I know for a fact they're going to fix it. And this little they're going to fix it. And this little they're going to fix it. And this little spike here is why Flash had those couple spike here is why Flash had those couple spike here is why Flash had those couple of good benchmarks that made it look of good benchmarks that made it look of good benchmarks that made it look really good when it wasn't. I'll do a really good when it wasn't. I'll do a really good when it wasn't. I'll do a more realistic comparison here to get my more realistic comparison here to get my more realistic comparison here to get my point across because no one should be point across because no one should be point across because no one should be using a Gemini thing. So, we're going to using a Gemini thing. So, we're going to using a Gemini thing. So, we're going to talk about Fable versus Opus because it talk about Fable versus Opus because it talk about Fable versus Opus because it seems like people believe Opus can do

  6. seems like people believe Opus can do seems like people believe Opus can do the work that they are doing. And that the work that they are doing. And that the work that they are doing. And that is true a lot of the time if your line is true a lot of the time if your line is true a lot of the time if your line for quality is here. If this is the for quality is here. If this is the for quality is here. If this is the prompts you're sending, if your prompts prompts you're sending, if your prompts prompts you're sending, if your prompts are things like here's the ticket. I are things like here's the ticket. I are things like here's the ticket. I want you to find the file and make this want you to find the file and make this want you to find the file and make this change and tell me when it's done so change and tell me when it's done so change and tell me when it's done so that I can test it. Then it is unlikely that I can test it. Then it is unlikely that I can test it. Then it is unlikely Opus or Fable or even a Gemini model is Opus or Fable or even a Gemini model is Opus or Fable or even a Gemini model is going to struggle too too much to do it. going to struggle too too much to do it. going to struggle too too much to do it. And if you have low tolerance for And if you have low tolerance for And if you have low tolerance for failure, which I'll admit even I do, failure, which I'll admit even I do, failure, which I'll admit even I do, when I ask the model to do a thing and when I ask the model to do a thing and when I ask the model to do a thing and it fails to do the thing, it pisses me it fails to do the thing, it pisses me it fails to do the thing, it pisses me off. Like it actually angers me. Which off. Like it actually angers me. Which off. Like it actually angers me. Which means if one of these dips is worse than means if one of these dips is worse than means if one of these dips is worse than the others, a bad thing happens. My bar the others, a bad thing happens. My bar the others, a bad thing happens. My bar gets lowered. This is where I'm gets lowered. This is where I'm gets lowered. This is where I'm comfortable prompting because if I tried comfortable prompting because if I tried comfortable prompting because if I tried a harder prompt or a prompt that a harder prompt or a prompt that a harder prompt or a prompt that involved more work and the model failed involved more work and the model failed involved more work and the model failed to do it, I now think models can't do to do it, I now think models can't do to do it, I now think models can't do that. So, I lower my expectations, the that. So, I lower my expectations, the that. So, I lower my expectations, the things I prompt for, the ways I prompt, things I prompt for, the ways I prompt, things I prompt for, the ways I prompt, and the most important detail, we'll be and the most important detail, we'll be and the most important detail, we'll be talking about this a lot, the width of talking about this a lot, the width of talking about this a lot, the width of my prompt, not the depth, not the my prompt, not the depth, not the my prompt, not the depth, not the difficulty of the thing, but the amount difficulty of the thing, but the amount difficulty of the thing, but the amount of things and the distance from the of things and the distance from the of things and the distance from the start to the end. Wider prompting start to the end. Wider prompting start to the end. Wider prompting requires higher floors. And to be clear, requires higher floors. And to be clear, requires higher floors. And to be clear, I'm not trying to say Zigg's bar is set I'm not trying to say Zigg's bar is set I'm not trying to say Zigg's bar is set this low. I'm guessing Z's bar is set this low. I'm guessing Z's bar is set this low. I'm guessing Z's bar is set hereish where sometimes it slightly hereish where sometimes it slightly hereish where sometimes it slightly disappoints, sometimes much more rarely, disappoints, sometimes much more rarely, disappoints, sometimes much more rarely, but sometimes it really disappoints, but but sometimes it really disappoints, but but sometimes it really disappoints, but generally things are solid. This is also generally things are solid. This is also generally things are solid. This is also what makes Astra so annoying is the

  7. what makes Astra so annoying is the what makes Astra so annoying is the random spikes into dumb are so random random spikes into dumb are so random random spikes into dumb are so random and so spiky that you don't really know and so spiky that you don't really know and so spiky that you don't really know how to gauge what level to operate at. how to gauge what level to operate at. how to gauge what level to operate at. This is why when a new Frontier model This is why when a new Frontier model This is why when a new Frontier model comes out, I immediately try to reset comes out, I immediately try to reset comes out, I immediately try to reset instead of just sending the same prompts instead of just sending the same prompts instead of just sending the same prompts I sent before, I always start with a set I sent before, I always start with a set I sent before, I always start with a set that I've tried on other things that that I've tried on other things that that I've tried on other things that failed to see if it can succeed. failed to see if it can succeed. failed to see if it can succeed. Usually, when I do those tests, I'll see Usually, when I do those tests, I'll see Usually, when I do those tests, I'll see a few things it did better at than a few things it did better at than a few things it did better at than expected, and from there, I can start to expected, and from there, I can start to expected, and from there, I can start to figure out what new things I can do that figure out what new things I can do that figure out what new things I can do that I couldn't. And again, I want to be I couldn't. And again, I want to be I couldn't. And again, I want to be clear here. If your goal is to go from clear here. If your goal is to go from clear here. If your goal is to go from Jira ticket to code, Zeke is right. If I Jira ticket to code, Zeke is right. If I Jira ticket to code, Zeke is right. If I have a well- formatted Jira ticket that have a well- formatted Jira ticket that have a well- formatted Jira ticket that lists what files the code is in and what lists what files the code is in and what lists what files the code is in and what exact behavior exists that shouldn't, I exact behavior exists that shouldn't, I exact behavior exists that shouldn't, I can throw that at Claude and get an can throw that at Claude and get an can throw that at Claude and get an answer relatively reliably. That's not answer relatively reliably. That's not answer relatively reliably. That's not what I'm talking about here. What I'm what I'm talking about here. What I'm what I'm talking about here. What I'm talking about is I get a DM on my phone talking about is I get a DM on my phone talking about is I get a DM on my phone from a bug a user had in T3 Code. So, I from a bug a user had in T3 Code. So, I from a bug a user had in T3 Code. So, I screenshot it. I paste it to my agent in screenshot it. I paste it to my agent in screenshot it. I paste it to my agent in T3 Code on my phone and say, "Fix this. T3 Code on my phone and say, "Fix this. T3 Code on my phone and say, "Fix this. Test it. Record a video showing it works Test it. Record a video showing it works Test it. Record a video showing it works now and link me the PR when you're done.

  8. now and link me the PR when you're done. now and link me the PR when you're done. Babysit it until all the issues that Babysit it until all the issues that Babysit it until all the issues that come up in review are addressed. That come up in review are addressed. That come up in review are addressed. That isn't harder than what I said before. isn't harder than what I said before. isn't harder than what I said before. That isn't harder than going and editing That isn't harder than going and editing That isn't harder than going and editing the code from the Jira ticket, but the code from the Jira ticket, but the code from the Jira ticket, but having the whole end to end where the having the whole end to end where the having the whole end to end where the model can go from a vague screenshot of model can go from a vague screenshot of model can go from a vague screenshot of what's wrong to a real functioning what's wrong to a real functioning what's wrong to a real functioning solution with a poll request that has a solution with a poll request that has a solution with a poll request that has a video proving it worked without my video proving it worked without my video proving it worked without my intervention at all. That is the intervention at all. That is the intervention at all. That is the capability that I'm excited about. Is it capability that I'm excited about. Is it capability that I'm excited about. Is it cool I can demo it making a crazy 3D cool I can demo it making a crazy 3D cool I can demo it making a crazy 3D game? Yeah. But it's way cooler that I game? Yeah. But it's way cooler that I game? Yeah. But it's way cooler that I can send it a screenshot of a problem can send it a screenshot of a problem can send it a screenshot of a problem and then go do something else and in an and then go do something else and in an and then go do something else and in an hour when I check in, it hasn't lost hour when I check in, it hasn't lost hour when I check in, it hasn't lost track of what it's doing. Has it burned track of what it's doing. Has it burned track of what it's doing. Has it burned more tokens? Yeah. I don't care though more tokens? Yeah. I don't care though more tokens? Yeah. I don't care though because despite the tokens being because despite the tokens being because despite the tokens being expensive, so are my engineers. So is expensive, so are my engineers. So is expensive, so are my engineers. So is our time. If Julius can be three times our time. If Julius can be three times our time. If Julius can be three times more productive by spending his salary more productive by spending his salary more productive by spending his salary in tokens, that's a no-brainer. Hell, he in tokens, that's a no-brainer. Hell, he in tokens, that's a no-brainer. Hell, he could do 2x salary in tokens because could do 2x salary in tokens because could do 2x salary in tokens because people like Julius are rare. And I would people like Julius are rare. And I would people like Julius are rare. And I would rather Julius ship three times more than rather Julius ship three times more than rather Julius ship three times more than risk it hiring two more engineers that risk it hiring two more engineers that risk it hiring two more engineers that will be more likely to get in his way will be more likely to get in his way will be more likely to get in his way than help him ship faster. But this is than help him ship faster. But this is than help him ship faster. But this is the like core point I really want to the like core point I really want to the like core point I really want to drive home here. If your tasks aren't drive home here. If your tasks aren't drive home here. If your tasks aren't super super narrow, and this is again super super narrow, and this is again super super narrow, and this is again how I want to think about this. Let's how I want to think about this. Let's how I want to think about this. Let's draw vertical lines like this instead. I draw vertical lines like this instead. I draw vertical lines like this instead. I find for most devs, and this is like find for most devs, and this is like find for most devs, and this is like I'll ask them to show me their prompts I'll ask them to show me their prompts I'll ask them to show me their prompts and show me their histories. I find most and show me their histories. I find most and show me their histories. I find most of their prompts look something like of their prompts look something like of their prompts look something like this. They are small and concentrated

  9. this. They are small and concentrated this. They are small and concentrated and only require a little bit of how the and only require a little bit of how the and only require a little bit of how the model operates. You only are relying on model operates. You only are relying on model operates. You only are relying on the model for so long. You might even be the model for so long. You might even be the model for so long. You might even be watching the thread as it goes. You know watching the thread as it goes. You know watching the thread as it goes. You know what? I'm going to do a poll quick. I what? I'm going to do a poll quick. I what? I'm going to do a poll quick. I want to know, do you watch the agent want to know, do you watch the agent want to know, do you watch the agent while it works? I will admit I'm a while it works? I will admit I'm a while it works? I will admit I'm a little disappointed in these results. I little disappointed in these results. I little disappointed in these results. I was hoping for almost never to be a was hoping for almost never to be a was hoping for almost never to be a clear win. You guys need to stop paying clear win. You guys need to stop paying clear win. You guys need to stop paying so much attention to your agents. I've so much attention to your agents. I've so much attention to your agents. I've been saying this for a while now. You been saying this for a while now. You been saying this for a while now. You guys should push your limits and your guys should push your limits and your guys should push your limits and your trust and see where it goes. I like trust and see where it goes. I like trust and see where it goes. I like this. I only watch if the first output this. I only watch if the first output this. I only watch if the first output is horribly wrong. That is a great is horribly wrong. That is a great is horribly wrong. That is a great mindset because watching isn't just mindset because watching isn't just mindset because watching isn't just trying to keep it from going doing wrong trying to keep it from going doing wrong trying to keep it from going doing wrong again. It's also helping you learn why again. It's also helping you learn why again. It's also helping you learn why it's going wrong. And even better, you it's going wrong. And even better, you it's going wrong. And even better, you can ask. You can say, "I expected you to can ask. You can say, "I expected you to can ask. You can say, "I expected you to do this, but you did this other thing do this, but you did this other thing do this, but you did this other thing instead. What led you there?" It won't instead. What led you there?" It won't instead. What led you there?" It won't get it perfectly. The models rarely know get it perfectly. The models rarely know get it perfectly. The models rarely know exactly why they did a thing, but they exactly why they did a thing, but they exactly why they did a thing, but they will usually indicate what signals they will usually indicate what signals they will usually indicate what signals they found and what tools they called and found and what tools they called and found and what tools they called and what files they read that led them in what files they read that led them in what files they read that led them in the wrong direction. I actually like the the wrong direction. I actually like the the wrong direction. I actually like the way Jamon framed this here. He likes way Jamon framed this here. He likes way Jamon framed this here. He likes doing overnight autonomous work, not doing overnight autonomous work, not doing overnight autonomous work, not because it gets a ton of code written, because it gets a ton of code written, because it gets a ton of code written, but because it will hit a bunch of those but because it will hit a bunch of those but because it will hit a bunch of those fail cases so that he can architect his fail cases so that he can architect his fail cases so that he can architect his code base and his systems so that when code base and his systems so that when code base and his systems so that when he's working more in the loop doing like he's working more in the loop doing like he's working more in the loop doing like one or two hour threads instead of 5 to one or two hour threads instead of 5 to one or two hour threads instead of 5 to 10 hour threads. If you have a speed 10 hour threads. If you have a speed 10 hour threads. If you have a speed bump in your codebase that agents hit on bump in your codebase that agents hit on bump in your codebase that agents hit on average once every two hours, then average once every two hours, then average once every two hours, then they'll hit on average two to three they'll hit on average two to three they'll hit on average two to three times every 5 hours. And if you do these times every 5 hours. And if you do these times every 5 hours. And if you do these super long runs, you can hit those super long runs, you can hit those super long runs, you can hit those failures faster and fix them. Learn from

  10. failures faster and fix them. Learn from failures faster and fix them. Learn from the bad runs. Make changes based on the the bad runs. Make changes based on the the bad runs. Make changes based on the bad runs. You should adjust your prompts bad runs. You should adjust your prompts bad runs. You should adjust your prompts a little, but you should adjust your a little, but you should adjust your a little, but you should adjust your codebase a lot. If agents are screwing codebase a lot. If agents are screwing codebase a lot. If agents are screwing up in your codebase, then a new dev up in your codebase, then a new dev up in your codebase, then a new dev would, too. Like if you took a super would, too. Like if you took a super would, too. Like if you took a super experienced dev that's never worked in experienced dev that's never worked in experienced dev that's never worked in your codebase before and they couldn't your codebase before and they couldn't your codebase before and they couldn't contribute by the end of the day, that's contribute by the end of the day, that's contribute by the end of the day, that's on you, not the agent. And if you are on you, not the agent. And if you are on you, not the agent. And if you are giving these very specific, precise giving these very specific, precise giving these very specific, precise instructions, hell, if you're I have a instructions, hell, if you're I have a instructions, hell, if you're I have a new poll actually, this one's going to new poll actually, this one's going to new poll actually, this one's going to hurt me. This one's going to really hurt hurt me. This one's going to really hurt hurt me. This one's going to really hurt me. Do you still mention names of files me. Do you still mention names of files me. Do you still mention names of files in your prompts? Often, sometimes, in your prompts? Often, sometimes, in your prompts? Often, sometimes, basically, never. Actually, never. This basically, never. Actually, never. This basically, never. Actually, never. This one's going to be very eye opening for one's going to be very eye opening for one's going to be very eye opening for me. Actually, I'm overcorrecting due to me. Actually, I'm overcorrecting due to me. Actually, I'm overcorrecting due to green field T3 code. We got 12 KPRs and green field T3 code. We got 12 KPRs and green field T3 code. We got 12 KPRs and 300,000 users. If I do anything wrong, 300,000 users. If I do anything wrong, 300,000 users. If I do anything wrong, we get flamed immediately. And I've also we get flamed immediately. And I've also we get flamed immediately. And I've also had phenomenal luck using this to had phenomenal luck using this to had phenomenal luck using this to maintain other huge code bases, too. But maintain other huge code bases, too. But maintain other huge code bases, too. But I really don't believe this like it I really don't believe this like it I really don't believe this like it doesn't work this way in real companies. doesn't work this way in real companies. doesn't work this way in real companies. No, it absolutely [ __ ] does. I just I No, it absolutely [ __ ] does. I just I No, it absolutely [ __ ] does. I just I don't believe it. I've helped enough don't believe it. I've helped enough don't believe it. I've helped enough bigger companies get this right. Let's bigger companies get this right. Let's bigger companies get this right. Let's see the results here. Okay, you guys see the results here. Okay, you guys see the results here. Okay, you guys have redeemed yourself. I feel much have redeemed yourself. I feel much have redeemed yourself. I feel much better now. The people in the often better now. The people in the often better now. The people in the often section I refer to as Atlassian devs and section I refer to as Atlassian devs and section I refer to as Atlassian devs and the ones at the bottom section I refer the ones at the bottom section I refer the ones at the bottom section I refer to as realistic ones. And I would guess to as realistic ones. And I would guess to as realistic ones. And I would guess that the ones at the top have a much that the ones at the top have a much that the ones at the top have a much better time with Opus and are confused better time with Opus and are confused better time with Opus and are confused about why people like things like you about why people like things like you about why people like things like you know Ael and Aster so much. The reason know Ael and Aster so much. The reason know Ael and Aster so much. The reason we like these new models isn't because

  11. we like these new models isn't because we like these new models isn't because we are pushing them to their absolute we are pushing them to their absolute we are pushing them to their absolute limits to make new sciences up. We just limits to make new sciences up. We just limits to make new sciences up. We just like that they're stupid less often. We like that they're stupid less often. We like that they're stupid less often. We like that they can go longer and make like that they can go longer and make like that they can go longer and make fewer mistakes and need less guidance to fewer mistakes and need less guidance to fewer mistakes and need less guidance to do the right thing. And to go back to my do the right thing. And to go back to my do the right thing. And to go back to my chart here, the thing I'm trying to chart here, the thing I'm trying to chart here, the thing I'm trying to emphasize is you should be striving to emphasize is you should be striving to emphasize is you should be striving to have one prompt, use more of the window have one prompt, use more of the window have one prompt, use more of the window of what the model's capable of. If you of what the model's capable of. If you of what the model's capable of. If you know where the rough areas are and you know where the rough areas are and you know where the rough areas are and you can smooth those out with changes to can smooth those out with changes to can smooth those out with changes to your codebase or just not using the your codebase or just not using the your codebase or just not using the model for those things and you find ways model for those things and you find ways model for those things and you find ways to go more horizontal, maybe instead of to go more horizontal, maybe instead of to go more horizontal, maybe instead of investigating the codebase and telling investigating the codebase and telling investigating the codebase and telling the agent which files to touch, you tell the agent which files to touch, you tell the agent which files to touch, you tell it what the problem is and tell it to it what the problem is and tell it to it what the problem is and tell it to find the files and change them itself. find the files and change them itself. find the files and change them itself. Maybe instead of telling it to let you Maybe instead of telling it to let you Maybe instead of telling it to let you know when it's done so you can build it know when it's done so you can build it know when it's done so you can build it on your phone, tell it to push the build on your phone, tell it to push the build on your phone, tell it to push the build to your phone when it's made the to your phone when it's made the to your phone when it's made the changes. Maybe tell it to run it in the changes. Maybe tell it to run it in the changes. Maybe tell it to run it in the simulator first, verify it, and then simulator first, verify it, and then simulator first, verify it, and then push it to my phone after so I can do push it to my phone after so I can do push it to my phone after so I can do one last check if I'm still concerned. one last check if I'm still concerned. one last check if I'm still concerned. Or just tell it to throw a video in the Or just tell it to throw a video in the Or just tell it to throw a video in the poll request so you know it worked. poll request so you know it worked. poll request so you know it worked. That's what we do most of the time now That's what we do most of the time now That's what we do most of the time now with T3 code. Here's a PR from Maria.

  12. with T3 code. Here's a PR from Maria. with T3 code. Here's a PR from Maria. This was a bunch of fixes to provider This was a bunch of fixes to provider This was a bunch of fixes to provider history issues when people were using history issues when people were using history issues when people were using the rewind features in various the rewind features in various the rewind features in various harnesses. Maria wrote none of this PR harnesses. Maria wrote none of this PR harnesses. Maria wrote none of this PR and it merged pretty quickly after she and it merged pretty quickly after she and it merged pretty quickly after she filed it because it was a good PR and I filed it because it was a good PR and I filed it because it was a good PR and I could open it. I could look at what it could open it. I could look at what it could open it. I could look at what it changed because it visualized what it changed because it visualized what it changed because it visualized what it changed because the model did all that. changed because the model did all that. changed because the model did all that. And by the time a human is bothered, it And by the time a human is bothered, it And by the time a human is bothered, it is much more likely the thing works. So is much more likely the thing works. So is much more likely the thing works. So these are the the the two core things I these are the the the two core things I these are the the the two core things I really want to push you guys to think really want to push you guys to think really want to push you guys to think more about. The first is how long can more about. The first is how long can more about. The first is how long can the model run without your input? And the model run without your input? And the model run without your input? And the second similar but not exactly the the second similar but not exactly the the second similar but not exactly the same and I want to clarify the same and I want to clarify the same and I want to clarify the differences. How likely is it that the differences. How likely is it that the differences. How likely is it that the thing works by the time a human gets thing works by the time a human gets thing works by the time a human gets involved? Again, these two things are involved? Again, these two things are involved? Again, these two things are not defined by using the models to not defined by using the models to not defined by using the models to rebuild 3D worlds. They're not defined rebuild 3D worlds. They're not defined rebuild 3D worlds. They're not defined by how well they can port all of by how well they can port all of by how well they can port all of Electron to Rust or anything. They're Electron to Rust or anything. They're Electron to Rust or anything. They're defined by how likely the model screws defined by how likely the model screws defined by how likely the model screws up. And your goal is to make it less up. And your goal is to make it less up. And your goal is to make it less likely that the model screws up in any likely that the model screws up in any likely that the model screws up in any given window. So, if in your experience, given window. So, if in your experience, given window. So, if in your experience, if you let the model go for 30 minutes, if you let the model go for 30 minutes, if you let the model go for 30 minutes, it usually hits bugs. It usually runs it usually hits bugs. It usually runs it usually hits bugs. It usually runs into problems, usually does stupid [ __ ] into problems, usually does stupid [ __ ] into problems, usually does stupid [ __ ] fix that. make it not do that because I fix that. make it not do that because I fix that. make it not do that because I have had runs go for six hours with no have had runs go for six hours with no have had runs go for six hours with no intervention that merged 10 minutes intervention that merged 10 minutes intervention that merged 10 minutes after I filed the PR because they did after I filed the PR because they did after I filed the PR because they did exactly what they were supposed to and exactly what they were supposed to and exactly what they were supposed to and the model had verified it itself before the model had verified it itself before the model had verified it itself before I got pulled in. But to go back here I got pulled in. But to go back here I got pulled in. But to go back here with the response quality, if your with the response quality, if your with the response quality, if your prompts look like this and the necessary prompts look like this and the necessary prompts look like this and the necessary bar for you is here, then yes, bar for you is here, then yes, bar for you is here, then yes, absolutely the difference between Fable

  13. absolutely the difference between Fable absolutely the difference between Fable and Opus isn't very big because both and Opus isn't very big because both and Opus isn't very big because both massively clear your bar in this window. massively clear your bar in this window. massively clear your bar in this window. But watch what happens when I move the But watch what happens when I move the But watch what happens when I move the end point over more and more. Oh no, now end point over more and more. Oh no, now end point over more and more. Oh no, now my bar isn't being met. And all of a my bar isn't being met. And all of a my bar isn't being met. And all of a sudden, if I make the task wider, not sudden, if I make the task wider, not sudden, if I make the task wider, not harder, wider, which means longer. It harder, wider, which means longer. It harder, wider, which means longer. It does more. The likelihood we hit one of does more. The likelihood we hit one of does more. The likelihood we hit one of those edges in something like Opus where those edges in something like Opus where those edges in something like Opus where it does something stupid goes up as time it does something stupid goes up as time it does something stupid goes up as time goes up. And the thing that makes the goes up. And the thing that makes the goes up. And the thing that makes the Frontier model special is that they hit Frontier model special is that they hit Frontier model special is that they hit those floors less and the floor is those floors less and the floor is those floors less and the floor is raised meaningfully. I don't like Fable raised meaningfully. I don't like Fable raised meaningfully. I don't like Fable because it's way smarter. I like Fable because it's way smarter. I like Fable because it's way smarter. I like Fable because it's less dumb. And those are because it's less dumb. And those are because it's less dumb. And those are different things. Less dumb and smart different things. Less dumb and smart different things. Less dumb and smart are almost opposites. They are the are almost opposites. They are the are almost opposites. They are the opposite ends of the spectrum. And if opposite ends of the spectrum. And if opposite ends of the spectrum. And if you're thinking of models as their peak you're thinking of models as their peak you're thinking of models as their peak capability and not their worst capability and not their worst capability and not their worst capabilities, you're not talking about capabilities, you're not talking about capabilities, you're not talking about them the right way. And as such, it's them the right way. And as such, it's them the right way. And as such, it's really hard for me to take anyone really hard for me to take anyone really hard for me to take anyone seriously when they say Opus will seriously when they say Opus will seriously when they say Opus will perform just as well for your tasks perform just as well for your tasks perform just as well for your tasks because it means that their tasks are because it means that their tasks are because it means that their tasks are really, really short and simple. And to really, really short and simple. And to really, really short and simple. And to be very clear, I have nothing against be very clear, I have nothing against be very clear, I have nothing against simple. I love using agents for simple simple. I love using agents for simple simple. I love using agents for simple stuff. That's what I do most of the stuff. That's what I do most of the stuff. That's what I do most of the time. But it's long simple stuff. At time. But it's long simple stuff. At time. But it's long simple stuff. At every generation bump, the amount of every generation bump, the amount of every generation bump, the amount of time until the model is 50% likely to time until the model is 50% likely to time until the model is 50% likely to have done something stupid goes down have done something stupid goes down have done something stupid goes down exponentially. And if you don't feel exponentially. And if you don't feel exponentially. And if you don't feel this way, or maybe you've been prompting this way, or maybe you've been prompting this way, or maybe you've been prompting this way occasionally and you're not this way occasionally and you're not this way occasionally and you're not happy with the results, then you have a happy with the results, then you have a happy with the results, then you have a great opportunity to make real great opportunity to make real great opportunity to make real improvements here. Maria had some good

  14. improvements here. Maria had some good improvements here. Maria had some good comments on this that I want to bring comments on this that I want to bring comments on this that I want to bring up. She spent 3 to four days going over up. She spent 3 to four days going over up. She spent 3 to four days going over her traces and refining skills and her traces and refining skills and her traces and refining skills and whatnot after seeing what the model does whatnot after seeing what the model does whatnot after seeing what the model does and doesn't do right and made all of and doesn't do right and made all of and doesn't do right and made all of these adjustments, and now she can just these adjustments, and now she can just these adjustments, and now she can just fire a single prompt and get a PR landed fire a single prompt and get a PR landed fire a single prompt and get a PR landed instantly. Yeah, it's great. Okay, she instantly. Yeah, it's great. Okay, she instantly. Yeah, it's great. Okay, she credits potato, not me. Fair. I get it. credits potato, not me. Fair. I get it. credits potato, not me. Fair. I get it. I'm trying to push these same things I'm trying to push these same things I'm trying to push these same things more. What percent of devs have a spend more. What percent of devs have a spend more. What percent of devs have a spend of many hundreds of dollars a month just of many hundreds of dollars a month just of many hundreds of dollars a month just to play with it? Most full-time devs can to play with it? Most full-time devs can to play with it? Most full-time devs can afford a $200 sub to Claude and a $200 afford a $200 sub to Claude and a $200 afford a $200 sub to Claude and a $200 sub to Codeex. And most of them probably sub to Codeex. And most of them probably sub to Codeex. And most of them probably work somewhere that is willing to pay work somewhere that is willing to pay work somewhere that is willing to pay for those as well. So yeah, that gets for those as well. So yeah, that gets for those as well. So yeah, that gets you 8 grand of Claude tokens and 12 you 8 grand of Claude tokens and 12 you 8 grand of Claude tokens and 12 grand of tokens from OpenAI. You got a grand of tokens from OpenAI. You got a grand of tokens from OpenAI. You got a lot of wiggle room for not a lot of lot of wiggle room for not a lot of lot of wiggle room for not a lot of money. And I know that is a lot for money. And I know that is a lot for money. And I know that is a lot for people who aren't in western countries, people who aren't in western countries, people who aren't in western countries, who aren't full-time devs, who are who aren't full-time devs, who are who aren't full-time devs, who are younger, who are students, etc. But younger, who are students, etc. But younger, who are students, etc. But those aren't the people I'm talking those aren't the people I'm talking those aren't the people I'm talking about here. If you are using the dumber about here. If you are using the dumber about here. If you are using the dumber models because that's what you can models because that's what you can models because that's what you can afford, you should be very careful which afford, you should be very careful which afford, you should be very careful which ones you use because sometimes the ones you use because sometimes the ones you use because sometimes the cheaper model ends up more expensive.

  15. cheaper model ends up more expensive. cheaper model ends up more expensive. But that's not who I'm talking about But that's not who I'm talking about But that's not who I'm talking about here. I'm talking about people like here. I'm talking about people like here. I'm talking about people like Kramer who concluded, I think he's wrong Kramer who concluded, I think he's wrong Kramer who concluded, I think he's wrong here because I should stick to YouTube here because I should stick to YouTube here because I should stick to YouTube videos and don't know anything about videos and don't know anything about videos and don't know anything about engineering. or people like David K who engineering. or people like David K who engineering. or people like David K who I love. He built Xstate which is one of I love. He built Xstate which is one of I love. He built Xstate which is one of the best state management libraries in the best state management libraries in the best state management libraries in the whole webdev world saying you don't the whole webdev world saying you don't the whole webdev world saying you don't need Astra Soul or Fable for most things need Astra Soul or Fable for most things need Astra Soul or Fable for most things which is true but my time is more which is true but my time is more which is true but my time is more valuable than Fables. So if I downgrade valuable than Fables. So if I downgrade valuable than Fables. So if I downgrade to a cheaper model I have to put more to a cheaper model I have to put more to a cheaper model I have to put more time in before it fires and more time in time in before it fires and more time in time in before it fires and more time in when it stops. And the more I let the when it stops. And the more I let the when it stops. And the more I let the model chew out both sides there and get model chew out both sides there and get model chew out both sides there and get involved earlier and pull me back in involved earlier and pull me back in involved earlier and pull me back in later, the better things are. I also saw later, the better things are. I also saw later, the better things are. I also saw somebody in chat pushing back on me somebody in chat pushing back on me somebody in chat pushing back on me saying exponential here. Let's see. This saying exponential here. Let's see. This saying exponential here. Let's see. This will take a bit, so I'll have to record will take a bit, so I'll have to record will take a bit, so I'll have to record an extra later for it. I want you to go an extra later for it. I want you to go an extra later for it. I want you to go through a set of my prompts and the through a set of my prompts and the through a set of my prompts and the responses from, I don't know, January responses from, I don't know, January responses from, I don't know, January versus now, maybe February if I don't versus now, maybe February if I don't versus now, maybe February if I don't have enough history. And I want you to have enough history. And I want you to have enough history. And I want you to figure out from a reasonably randomized figure out from a reasonably randomized figure out from a reasonably randomized set how long my average prompt ran for.

  16. set how long my average prompt ran for. set how long my average prompt ran for. So, from when I sent the prompt to when So, from when I sent the prompt to when So, from when I sent the prompt to when the response stopped generating, how has the response stopped generating, how has the response stopped generating, how has the amount of time changed from January the amount of time changed from January the amount of time changed from January or February to now? Great. My Vibe proxy or February to now? Great. My Vibe proxy or February to now? Great. My Vibe proxy is quite broken. You know what? I got is quite broken. You know what? I got is quite broken. You know what? I got some Opus usage. Let's let Opus do some Opus usage. Let's let Opus do some Opus usage. Let's let Opus do something for once. I saw somebody in something for once. I saw somebody in something for once. I saw somebody in chat say they thought it would be 10 to chat say they thought it would be 10 to chat say they thought it would be 10 to 15% longer. And I feel like I am insane. 15% longer. And I feel like I am insane. 15% longer. And I feel like I am insane. I didn't trust models to run for more I didn't trust models to run for more I didn't trust models to run for more than 15 minutes just a few months ago. than 15 minutes just a few months ago. than 15 minutes just a few months ago. It even was like GBD55 which was It even was like GBD55 which was It even was like GBD55 which was generally better with agentic stuff. It generally better with agentic stuff. It generally better with agentic stuff. It stopped so often that I found its runs stopped so often that I found its runs stopped so often that I found its runs were actually kind of shorter overall were actually kind of shorter overall were actually kind of shorter overall that the amount of time these sessions that the amount of time these sessions that the amount of time these sessions ran for went down with 55 and then 56 ran for went down with 55 and then 56 ran for went down with 55 and then 56 suddenly I could let it run way longer. suddenly I could let it run way longer. suddenly I could let it run way longer. And now my agent runs averaged probably And now my agent runs averaged probably And now my agent runs averaged probably 2x longer, but the top 1% longest ones 2x longer, but the top 1% longest ones 2x longer, but the top 1% longest ones are at least 10 times longer. I've had are at least 10 times longer. I've had are at least 10 times longer. I've had things run for 2 days straight with no things run for 2 days straight with no things run for 2 days straight with no issues and I could barely get a thing to issues and I could barely get a thing to issues and I could barely get a thing to run for an hour before. Just got the run for an hour before. Just got the run for an hour before. Just got the numbers in and I didn't have as much numbers in and I didn't have as much numbers in and I didn't have as much data as I was hoping. I only have my data as I was hoping. I only have my data as I was hoping. I only have my logs since March because I did a logs since March because I did a logs since March because I did a computer move and didn't back up my computer move and didn't back up my computer move and didn't back up my agent history cuz I didn't care much agent history cuz I didn't care much agent history cuz I didn't care much yet. And here are the results. Median yet. And here are the results. Median yet. And here are the results. Median prompt went from 53 seconds to 2 minutes prompt went from 53 seconds to 2 minutes prompt went from 53 seconds to 2 minutes and 20 seconds. So median more than and 20 seconds. So median more than and 20 seconds. So median more than doubled in length. P95 went from a bit doubled in length. P95 went from a bit doubled in length. P95 went from a bit under 7 minutes to over 16 minutes and under 7 minutes to over 16 minutes and under 7 minutes to over 16 minutes and 20 seconds. You understand, right?

  17. 20 seconds. You understand, right? 20 seconds. You understand, right? That's from April to now. And here you That's from April to now. And here you That's from April to now. And here you can see over time this is actually can see over time this is actually can see over time this is actually really useful. In March my 5% longest really useful. In March my 5% longest really useful. In March my 5% longest requests were 9 minutes long. Then it requests were 9 minutes long. Then it requests were 9 minutes long. Then it went to 11 then 12. And then May to June went to 11 then 12. And then May to June went to 11 then 12. And then May to June this is when we started to get Fable and this is when we started to get Fable and this is when we started to get Fable and Soul. We went from 12 minutes to 22 Soul. We went from 12 minutes to 22 Soul. We went from 12 minutes to 22 minutes. Nearly doubled month overmonth minutes. Nearly doubled month overmonth minutes. Nearly doubled month overmonth just from the new models. So yes it is just from the new models. So yes it is just from the new models. So yes it is exponential. The rate at which the exponential. The rate at which the exponential. The rate at which the length your prompts can go for is length your prompts can go for is length your prompts can go for is massively skyrocketing. Chat's hopping massively skyrocketing. Chat's hopping massively skyrocketing. Chat's hopping in to agree here. I can confidently say in to agree here. I can confidently say in to agree here. I can confidently say that mine went from 5 to 15 minutes to 1 that mine went from 5 to 15 minutes to 1 that mine went from 5 to 15 minutes to 1 to four hours. Oh, sorry. It was the to four hours. Oh, sorry. It was the to four hours. Oh, sorry. It was the floor improved 10 to 15%. I didn't say floor improved 10 to 15%. I didn't say floor improved 10 to 15%. I didn't say the floor improved exponentially. I said the floor improved exponentially. I said the floor improved exponentially. I said the impact of it improved exponentially. the impact of it improved exponentially. the impact of it improved exponentially. Here, let let's do the math out here. Here, let let's do the math out here. Here, let let's do the math out here. Let's say you have something that fails Let's say you have something that fails Let's say you have something that fails 5% of the time in a 10-minute window. 5% of the time in a 10-minute window. 5% of the time in a 10-minute window. That means that you have a 95% chance of That means that you have a 95% chance of That means that you have a 95% chance of success in that same window. What success in that same window. What success in that same window. What happens if you want to run for 30 happens if you want to run for 30 happens if you want to run for 30 minutes? You all know how this math minutes? You all know how this math minutes? You all know how this math works, right? 0.95 to the power of works, right? 0.95 to the power of works, right? 0.95 to the power of three. Going for 10 minutes to 30 three. Going for 10 minutes to 30 three. Going for 10 minutes to 30 minutes changes your failure rate from minutes changes your failure rate from minutes changes your failure rate from 5% to 15%. Let's say you want to go for 5% to 15%. Let's say you want to go for 5% to 15%. Let's say you want to go for an hour. Oh god, now I'm at 73%. 2 an hour. Oh god, now I'm at 73%. 2 an hour. Oh god, now I'm at 73%. 2 hours, now you're at a 50% fail rate. 4 hours, now you're at a 50% fail rate. 4 hours, now you're at a 50% fail rate. 4 hours and now you're at a 30% success.

  18. hours and now you're at a 30% success. hours and now you're at a 30% success. Let's just slightly bump this. Let's say Let's just slightly bump this. Let's say Let's just slightly bump this. Let's say you improved the floor by 2%. It instead you improved the floor by 2%. It instead you improved the floor by 2%. It instead of failing 5% of the time in 10 minutes, of failing 5% of the time in 10 minutes, of failing 5% of the time in 10 minutes, it's now 3% of the time in 10 minutes. it's now 3% of the time in 10 minutes. it's now 3% of the time in 10 minutes. That bumps us here to 97. Oh wow, that's That bumps us here to 97. Oh wow, that's That bumps us here to 97. Oh wow, that's kind of crazy. That's only a 50% fail kind of crazy. That's only a 50% fail kind of crazy. That's only a 50% fail rate at 4 hours, but it's only a 2% rate at 4 hours, but it's only a 2% rate at 4 hours, but it's only a 2% difference, wasn't it? Oh yeah, it was a difference, wasn't it? Oh yeah, it was a difference, wasn't it? Oh yeah, it was a 70% fail rate with a 5% every 10 70% fail rate with a 5% every 10 70% fail rate with a 5% every 10 minutes. And now it's only 50%. So that minutes. And now it's only 50%. So that minutes. And now it's only 50%. So that 2% change ends up being 20% at the time 2% change ends up being 20% at the time 2% change ends up being 20% at the time scale of 4 hours. That's a 2% scale of 4 hours. That's a 2% scale of 4 hours. That's a 2% difference. Now imagine it's 10 to 15% difference. Now imagine it's 10 to 15% difference. Now imagine it's 10 to 15% like you said it was. Oh man, that's an like you said it was. Oh man, that's an like you said it was. Oh man, that's an exponential change in how long you can exponential change in how long you can exponential change in how long you can run. When you make these small cuts to run. When you make these small cuts to run. When you make these small cuts to fail rates in given time windows, you fail rates in given time windows, you fail rates in given time windows, you exponentially increase the distance that exponentially increase the distance that exponentially increase the distance that it can run for. AD guy said he could run it can run for. AD guy said he could run it can run for. AD guy said he could run for 48 hours straight back in February, for 48 hours straight back in February, for 48 hours straight back in February, but he'd have to spend several hours but he'd have to spend several hours but he'd have to spend several hours building up specs. They're not going to building up specs. They're not going to building up specs. They're not going to do things that long with more improvised do things that long with more improvised do things that long with more improvised and shorter prep periods. I don't think and shorter prep periods. I don't think and shorter prep periods. I don't think you really could do this before, even you really could do this before, even you really could do this before, even with really good specs, because the with really good specs, because the with really good specs, because the coherency the model has over time wasn't coherency the model has over time wasn't coherency the model has over time wasn't great. And as crazy cool as Ralph loops great. And as crazy cool as Ralph loops great. And as crazy cool as Ralph loops were, the models weren't good enough at were, the models weren't good enough at were, the models weren't good enough at compaction or keeping track of what compaction or keeping track of what compaction or keeping track of what they've done in the past or leaving they've done in the past or leaving they've done in the past or leaving reminders of what they've tried. And the reminders of what they've tried. And the reminders of what they've tried. And the results ended up being still very, very results ended up being still very, very results ended up being still very, very high failure rates. Those have dropped high failure rates. Those have dropped high failure rates. Those have dropped exponentially. You can do things like exponentially. You can do things like exponentially. You can do things like write these specs to help keep it write these specs to help keep it write these specs to help keep it somewhat more on track, but it only somewhat more on track, but it only somewhat more on track, but it only helped so much. And it didn't bump these

  19. helped so much. And it didn't bump these helped so much. And it didn't bump these failure rates often enough. No matter failure rates often enough. No matter failure rates often enough. No matter how much work you put in, the result of how much work you put in, the result of how much work you put in, the result of specking out a run and letting it go for specking out a run and letting it go for specking out a run and letting it go for 8 hours isn't T3 code. The result of 8 hours isn't T3 code. The result of 8 hours isn't T3 code. The result of that is cursor 2 and cursor 3 where that is cursor 2 and cursor 3 where that is cursor 2 and cursor 3 where everything broke as soon as you looked everything broke as soon as you looked everything broke as soon as you looked at it too closely. And now that models at it too closely. And now that models at it too closely. And now that models are good enough, cursor starting to get are good enough, cursor starting to get are good enough, cursor starting to get stable because they don't want to be in stable because they don't want to be in stable because they don't want to be in the loop. They want to run it for 4 to 8 the loop. They want to run it for 4 to 8 the loop. They want to run it for 4 to 8 hours even if it's bad. And now that 4 hours even if it's bad. And now that 4 hours even if it's bad. And now that 4 to 8 hour runs are way more likely to to 8 hour runs are way more likely to to 8 hour runs are way more likely to come out good. Suddenly cursor functions come out good. Suddenly cursor functions come out good. Suddenly cursor functions again. Obviously Lawrence to credit again. Obviously Lawrence to credit again. Obviously Lawrence to credit there to some extent too. But yeah, there to some extent too. But yeah, there to some extent too. But yeah, meaningful difference. So while I deeply meaningful difference. So while I deeply meaningful difference. So while I deeply respect Zeg and David K, I genuinely respect Zeg and David K, I genuinely respect Zeg and David K, I genuinely think both are still prompting like think both are still prompting like think both are still prompting like we're in February. And the reason why we're in February. And the reason why we're in February. And the reason why they're doing that is they're not they're doing that is they're not they're doing that is they're not valuing their time properly. They feel valuing their time properly. They feel valuing their time properly. They feel good putting that extra effort in at the good putting that extra effort in at the good putting that extra effort in at the start and the end because as a great dev start and the end because as a great dev start and the end because as a great dev before the thing that got you to level before the thing that got you to level before the thing that got you to level up, the thing that got you from a good up, the thing that got you from a good up, the thing that got you from a good contributor to a good leader was doing contributor to a good leader was doing contributor to a good leader was doing more of the prep before the code started more of the prep before the code started more of the prep before the code started and doing more of the vetting after. So and doing more of the vetting after. So and doing more of the vetting after. So it's even more uncomfortable to give it's even more uncomfortable to give it's even more uncomfortable to give that up to the agent. So the parties I that up to the agent. So the parties I that up to the agent. So the parties I see falling for this are the ones whose see falling for this are the ones whose see falling for this are the ones whose work isn't serious enough to realize the work isn't serious enough to realize the work isn't serious enough to realize the power of the models. But even more so, power of the models. But even more so, power of the models. But even more so, it's the incredibly talented leaders who it's the incredibly talented leaders who it's the incredibly talented leaders who have largely left behind coding in their have largely left behind coding in their have largely left behind coding in their day-to-day because the thing before and day-to-day because the thing before and day-to-day because the thing before and after the codew writing matters more.

  20. after the codew writing matters more. after the codew writing matters more. They struggled to give up the code in They struggled to give up the code in They struggled to give up the code in the middle, but they did. They won't the middle, but they did. They won't the middle, but they did. They won't give up the things on the other side give up the things on the other side give up the things on the other side yet, which is why they don't see the yet, which is why they don't see the yet, which is why they don't see the benefit. They are testing the models benefit. They are testing the models benefit. They are testing the models against the thing they already stopped against the thing they already stopped against the thing they already stopped doing. And the models have been able to doing. And the models have been able to doing. And the models have been able to do the thing they stopped doing for 6 do the thing they stopped doing for 6 do the thing they stopped doing for 6 months. They're correct there. But the months. They're correct there. But the months. They're correct there. But the moment you let the model go a little moment you let the model go a little moment you let the model go a little further in either direction, you'll further in either direction, you'll further in either direction, you'll suddenly start to see the edges a hell suddenly start to see the edges a hell suddenly start to see the edges a hell of a lot closer. So, as per your of a lot closer. So, as per your of a lot closer. So, as per your request, Zeie, I will stick to making request, Zeie, I will stick to making request, Zeie, I will stick to making videos because otherwise your stupid videos because otherwise your stupid videos because otherwise your stupid take's going to go too far and I need to take's going to go too far and I need to take's going to go too far and I need to make sure the next generation of devs make sure the next generation of devs make sure the next generation of devs who haven't fallen for this [ __ ] don't who haven't fallen for this [ __ ] don't who haven't fallen for this [ __ ] don't because what you're saying sounds good because what you're saying sounds good because what you're saying sounds good and we want to believe it. I want to and we want to believe it. I want to and we want to believe it. I want to believe it. I would love to not have to believe it. I would love to not have to believe it. I would love to not have to spend more money on my models, but I do spend more money on my models, but I do spend more money on my models, but I do because my time is more valuable than my because my time is more valuable than my because my time is more valuable than my [ __ ] posts. And I hope you realize the [ __ ] posts. And I hope you realize the [ __ ] posts. And I hope you realize the same soon, too. I hope you enjoyed this same soon, too. I hope you enjoyed this same soon, too. I hope you enjoyed this video, Zeke. And if anybody else happens video, Zeke. And if anybody else happens video, Zeke. And if anybody else happens to see it, maybe you'll like it, too. to see it, maybe you'll like it, too. to see it, maybe you'll like it, too. Let me know how you feel about this one Let me know how you feel about this one Let me know how you feel about this one and how wide your prompts have been. And and how wide your prompts have been. And and how wide your prompts have been. And if you think I'm crazy for not including if you think I'm crazy for not including if you think I'm crazy for not including file names in my prompts anymore. And file names in my prompts anymore. And file names in my prompts anymore. And until next time, he's nerds.

No summary available yet.

View original episode ↗