← Back
Nate B. Jones July 10, 2026 28m

1.6M agents registered for OpenClaw and did NOTHING.

Read full transcript 21 segments
  1. I'm interested in agents that work, not I'm interested in agents that work, not agents that don't. But one of the agents that don't. But one of the agents that don't. But one of the fundamental problems with finding agents fundamental problems with finding agents fundamental problems with finding agents that work is that we don't know what that work is that we don't know what that work is that we don't know what work looks like whe when it's work looks like whe when it's work looks like whe when it's agent-shaped. We don't know how to agent-shaped. We don't know how to agent-shaped. We don't know how to recognize agent problems when we see recognize agent problems when we see recognize agent problems when we see them. Is it a single agent problem? Is them. Is it a single agent problem? Is them. Is it a single agent problem? Is it a multi- aent problem? Is it not an it a multi- aent problem? Is it not an it a multi- aent problem? Is it not an agent problem at all? It's really hard agent problem at all? It's really hard agent problem at all? It's really hard for people to understand practically in for people to understand practically in for people to understand practically in front of their own desks where work fits front of their own desks where work fits front of their own desks where work fits in these categories. This video solves in these categories. This video solves in these categories. This video solves that. I'm going to walk you through how that. I'm going to walk you through how that. I'm going to walk you through how we know. I'm going to give you the we know. I'm going to give you the we know. I'm going to give you the academic grounding, the practical academic grounding, the practical academic grounding, the practical grounding from Frontier Labs, and I'm grounding from Frontier Labs, and I'm grounding from Frontier Labs, and I'm going to give you a easy one minute going to give you a easy one minute going to give you a easy one minute test. And on top of that, I'm going to test. And on top of that, I'm going to test. And on top of that, I'm going to give you an automated AI skill that you give you an automated AI skill that you give you an automated AI skill that you can jump into and grab and use to get can jump into and grab and use to get can jump into and grab and use to get into the work itself. Because we don't into the work itself. Because we don't into the work itself. Because we don't want to just do this for the sake of want to just do this for the sake of want to just do this for the sake of doing it. You're not sitting there doing it. You're not sitting there doing it. You're not sitting there estimating work for the sake of estimating work for the sake of estimating work for the sake of estimating. That's shaving the yak, estimating. That's shaving the yak, estimating. That's shaving the yak, right? What you're trying to do is right? What you're trying to do is right? What you're trying to do is you're trying to get into the work more you're trying to get into the work more you're trying to get into the work more quickly and pick more accurately. And so quickly and pick more accurately. And so quickly and pick more accurately. And so we're going to use our one minute test we're going to use our one minute test we're going to use our one minute test that I'm going to teach you and we're that I'm going to teach you and we're that I'm going to teach you and we're going to jump straight in. And I I going to jump straight in. And I I going to jump straight in. And I I actually built a tool for this. You're actually built a tool for this. You're actually built a tool for this. You're going to be able to jump straight into a going to be able to jump straight into a going to be able to jump straight into a multi- aent solution or a single agent multi- aent solution or a single agent multi- aent solution or a single agent solution or maybe a chat or even no no solution or maybe a chat or even no no solution or maybe a chat or even no no AI at all. Yes, we're going to talk AI at all. Yes, we're going to talk AI at all. Yes, we're going to talk about when you don't use AI at all. The about when you don't use AI at all. The about when you don't use AI at all. The answer to all this, the reason all this answer to all this, the reason all this answer to all this, the reason all this matters is because we are in a matters is because we are in a matters is because we are in a postopenclaw moment. 1.6 million agents postopenclaw moment. 1.6 million agents postopenclaw moment. 1.6 million agents registered for an agent-driven social registered for an agent-driven social registered for an agent-driven social network at the peak of openclaw. Most of network at the peak of openclaw. Most of network at the peak of openclaw. Most of them did not do a single task. They

  2. them did not do a single task. They them did not do a single task. They didn't do a task because people set up didn't do a task because people set up didn't do a task because people set up OpenClaw, didn't know what to do with it OpenClaw, didn't know what to do with it OpenClaw, didn't know what to do with it next. We live in a world where we have next. We live in a world where we have next. We live in a world where we have intelligence and we don't know how to intelligence and we don't know how to intelligence and we don't know how to use it and we don't have an idea of how use it and we don't have an idea of how use it and we don't have an idea of how to match our tasks to the agents with to match our tasks to the agents with to match our tasks to the agents with confidence. This video solves that. So, confidence. This video solves that. So, confidence. This video solves that. So, let's jump in. So, let me make it let's jump in. So, let me make it let's jump in. So, let me make it concrete. Three tasks. Not my tasks, concrete. Three tasks. Not my tasks, concrete. Three tasks. Not my tasks, right? Your tasks. You have all three of right? Your tasks. You have all three of right? Your tasks. You have all three of these on your desk. Now, I bet you have these on your desk. Now, I bet you have these on your desk. Now, I bet you have a scheduling task, right? You have to a scheduling task, right? You have to a scheduling task, right? You have to find a slot for a meeting. You have to find a slot for a meeting. You have to find a slot for a meeting. You have to book a class. You have to fit an book a class. You have to fit an book a class. You have to fit an appointment around a calendar that's appointment around a calendar that's appointment around a calendar that's really busy. Task two, you have a pile really busy. Task two, you have a pile really busy. Task two, you have a pile of something that you're not going to of something that you're not going to of something that you're not going to read. You have to, right? Maybe it's read. You have to, right? Maybe it's read. You have to, right? Maybe it's your inbox. Maybe it's a contracts your inbox. Maybe it's a contracts your inbox. Maybe it's a contracts folder. Maybe it's a share drive, right? folder. Maybe it's a share drive, right? folder. Maybe it's a share drive, right? It's a big pile of something. Maybe it's It's a big pile of something. Maybe it's It's a big pile of something. Maybe it's a bunch of meeting notes that you always a bunch of meeting notes that you always a bunch of meeting notes that you always meant to organize. Pick a pile. meant to organize. Pick a pile. meant to organize. Pick a pile. Everybody has one. And then number Everybody has one. And then number Everybody has one. And then number three, you have a judgment call to make. three, you have a judgment call to make. three, you have a judgment call to make. Which candidate do you hire maybe? Or Which candidate do you hire maybe? Or Which candidate do you hire maybe? Or what do you name something? Or what what do you name something? Or what what do you name something? Or what direction do you take a product? Here's direction do you take a product? Here's direction do you take a product? Here's what I want you to know about these what I want you to know about these what I want you to know about these three tasks. One of them is a 30-cond AI three tasks. One of them is a 30-cond AI three tasks. One of them is a 30-cond AI job. One of them needs a whole team of job. One of them needs a whole team of job. One of them needs a whole team of agents to do the job properly. And and agents to do the job properly. And and agents to do the job properly. And and the token bill, by the way, is smaller the token bill, by the way, is smaller the token bill, by the way, is smaller than you think for that. And I'll show than you think for that. And I'll show than you think for that. And I'll show you how. And one of them, no matter how you how. And one of them, no matter how you how. And one of them, no matter how much AI you use, is going to need your much AI you use, is going to need your much AI you use, is going to need your involvement. The reason I know this is involvement. The reason I know this is involvement. The reason I know this is that I actually spent the money finding that I actually spent the money finding that I actually spent the money finding out, like I actually ran all three of out, like I actually ran all three of out, like I actually ran all three of these tests, and you're going to see me these tests, and you're going to see me these tests, and you're going to see me run them on screen as we go through this run them on screen as we go through this run them on screen as we go through this video together. So, everybody bought video together. So, everybody bought video together. So, everybody bought thinking, but no one knows what to point thinking, but no one knows what to point thinking, but no one knows what to point that thinking at. This is one of the that thinking at. This is one of the that thinking at. This is one of the core issues with the AI economy. If we

  3. core issues with the AI economy. If we core issues with the AI economy. If we can't figure out how to use this can't figure out how to use this can't figure out how to use this intelligence, we're wasting it. You'll intelligence, we're wasting it. You'll intelligence, we're wasting it. You'll be able to figure out, is it a chat be able to figure out, is it a chat be able to figure out, is it a chat task? Is it a one agent task? Is it a task? Is it a one agent task? Is it a task? Is it a one agent task? Is it a multi- aent task? Or whether you multi- aent task? Or whether you multi- aent task? Or whether you shouldn't bother at all. Right? And the shouldn't bother at all. Right? And the shouldn't bother at all. Right? And the trick is there's only four things you trick is there's only four things you trick is there's only four things you need to estimate to answer that need to estimate to answer that need to estimate to answer that question. And you can do those really question. And you can do those really question. And you can do those really quickly. That's the whole skill. And I'm quickly. That's the whole skill. And I'm quickly. That's the whole skill. And I'm going to run my version of all three of going to run my version of all three of going to run my version of all three of those tasks I just gave you on camera so those tasks I just gave you on camera so those tasks I just gave you on camera so you can watch that test work. And one of you can watch that test work. And one of you can watch that test work. And one of the tasks the agent is going to tell me the tasks the agent is going to tell me the tasks the agent is going to tell me no, which I love. And we'll get into why no, which I love. And we'll get into why no, which I love. And we'll get into why that is. For all of our history before that is. For all of our history before that is. For all of our history before AI came along, if you wanted to apply AI came along, if you wanted to apply AI came along, if you wanted to apply thinking to a problem, you had two thinking to a problem, you had two thinking to a problem, you had two options. You could either hire a brain, options. You could either hire a brain, options. You could either hire a brain, hire someone to do that thinking for hire someone to do that thinking for hire someone to do that thinking for you, or you could wait until you had you, or you could wait until you had you, or you could wait until you had time to think yourself. Thinking came time to think yourself. Thinking came time to think yourself. Thinking came connected to people. People were connected to people. People were connected to people. People were expensive, they were slow to find, they expensive, they were slow to find, they expensive, they were slow to find, they fall asleep, they need food, they need fall asleep, they need food, they need fall asleep, they need food, they need rest, etc. and they also need clear rest, etc. and they also need clear rest, etc. and they also need clear clear direction and clear access to clear direction and clear access to clear direction and clear access to resources and all kinds of things that resources and all kinds of things that resources and all kinds of things that make hiring relatively expensive.

  4. make hiring relatively expensive. make hiring relatively expensive. That era is over. And by the way, that That era is over. And by the way, that That era is over. And by the way, that doesn't mean jobs are over. I've said doesn't mean jobs are over. I've said doesn't mean jobs are over. I've said that and I'll cover that in another that and I'll cover that in another that and I'll cover that in another video. But the era of tying thinking to video. But the era of tying thinking to video. But the era of tying thinking to people is done. Especially post November people is done. Especially post November people is done. Especially post November 2025. Thinking is now metered. It's 2025. Thinking is now metered. It's 2025. Thinking is now metered. It's priced per token. You can buy a little priced per token. You can buy a little priced per token. You can buy a little of it. You can buy a lot of it. You can of it. You can buy a lot of it. You can of it. You can buy a lot of it. You can buy it tonight for a problem you only buy it tonight for a problem you only buy it tonight for a problem you only discovered this afternoon. And nobody, discovered this afternoon. And nobody, discovered this afternoon. And nobody, including me, grew up with instincts for including me, grew up with instincts for including me, grew up with instincts for for any of that, for how we handle that. for any of that, for how we handle that. for any of that, for how we handle that. This is net new to us, right? It's a new This is net new to us, right? It's a new This is net new to us, right? It's a new kind of managerial instinct. Nobody has kind of managerial instinct. Nobody has kind of managerial instinct. Nobody has ever had to ask which task in my week is ever had to ask which task in my week is ever had to ask which task in my week is actually worth 50 bucks of purchase actually worth 50 bucks of purchase actually worth 50 bucks of purchase thought. That question just didn't exist thought. That question just didn't exist thought. That question just didn't exist a year ago. So when people stand in a year ago. So when people stand in a year ago. So when people stand in front of their shiny new agent and say, front of their shiny new agent and say, front of their shiny new agent and say, "What do I even do with this?" That's "What do I even do with this?" That's "What do I even do with this?" That's not a tooling problem, right? It's not a not a tooling problem, right? It's not a not a tooling problem, right? It's not a lack of imagination on our part. It's lack of imagination on our part. It's lack of imagination on our part. It's it's a budgeting question that we don't it's a budgeting question that we don't it's a budgeting question that we don't have the tools to answer because we've have the tools to answer because we've have the tools to answer because we've never had to before. Maybe you watch a never had to before. Maybe you watch a never had to before. Maybe you watch a demo of someone's 23 agents form and and demo of someone's 23 agents form and and demo of someone's 23 agents form and and you feel left behind somehow. I I want you feel left behind somehow. I I want you feel left behind somehow. I I want to give you a word of comfort here. You to give you a word of comfort here. You to give you a word of comfort here. You may not need those 23 agents for the may not need those 23 agents for the may not need those 23 agents for the task. They may be overkill and I'm the task. They may be overkill and I'm the task. They may be overkill and I'm the last person to suggest a multi- aent last person to suggest a multi- aent last person to suggest a multi- aent solution where you don't need one. The solution where you don't need one. The solution where you don't need one. The trick is you might need one. And so what trick is you might need one. And so what trick is you might need one. And so what I realized the most important skill is I realized the most important skill is I realized the most important skill is that we're not learning that we're not that we're not learning that we're not that we're not learning that we're not teaching that I couldn't find anywhere teaching that I couldn't find anywhere teaching that I couldn't find anywhere is how do you know when you need an is how do you know when you need an is how do you know when you need an agent? How do you know when you need agent? How do you know when you need agent? How do you know when you need multiple agents? How do you know when multiple agents? How do you know when multiple agents? How do you know when you need none? Before I jump into the you need none? Before I jump into the you need none? Before I jump into the one minute test for how to get this one minute test for how to get this one minute test for how to get this done, I want to give you the two facts done, I want to give you the two facts done, I want to give you the two facts that I built this test on so you can

  5. that I built this test on so you can that I built this test on so you can understand how to extend this further. understand how to extend this further. understand how to extend this further. So Stanford 2024 researchers took a So Stanford 2024 researchers took a So Stanford 2024 researchers took a cheap coding model and gave it one cheap coding model and gave it one cheap coding model and gave it one attempt at each bug in a standard attempt at each bug in a standard attempt at each bug in a standard benchmark. It fixed 15.9% of them. Not benchmark. It fixed 15.9% of them. Not benchmark. It fixed 15.9% of them. Not impressive even back then, right in impressive even back then, right in impressive even back then, right in 2024. Then they gave the same cheap 2024. Then they gave the same cheap 2024. Then they gave the same cheap model 250 attempts per bug. They jumped model 250 attempts per bug. They jumped model 250 attempts per bug. They jumped to 56% without changing the model, to 56% without changing the model, to 56% without changing the model, without changing the harness, without without changing the harness, without without changing the harness, without changing anything. For context, the best changing anything. For context, the best changing anything. For context, the best single attempt from the best model money single attempt from the best model money single attempt from the best model money could buy at the time in 2024 was 43%. could buy at the time in 2024 was 43%. could buy at the time in 2024 was 43%. So they beat state-of-the-art just by So they beat state-of-the-art just by So they beat state-of-the-art just by trying a little bit more. What's the trying a little bit more. What's the trying a little bit more. What's the takeaway here for us? Why are we quoting takeaway here for us? Why are we quoting takeaway here for us? Why are we quoting a 2024 study for AI in 2026? I'll tell a 2024 study for AI in 2026? I'll tell a 2024 study for AI in 2026? I'll tell you why. Stanford discovered a law, a you why. Stanford discovered a law, a you why. Stanford discovered a law, a pattern in AI token usage that we still pattern in AI token usage that we still pattern in AI token usage that we still see upheld today by agents and that see upheld today by agents and that see upheld today by agents and that informed how I think about teaching all informed how I think about teaching all informed how I think about teaching all of you where tasks are agent applicable of you where tasks are agent applicable of you where tasks are agent applicable and where they're not. Because you see, and where they're not. Because you see, and where they're not. Because you see, when Stanford graphed the curve of when Stanford graphed the curve of when Stanford graphed the curve of improvement, how the AI models got improvement, how the AI models got improvement, how the AI models got better at fixing bugs as they added more better at fixing bugs as they added more better at fixing bugs as they added more turns, they found that the improvement turns, they found that the improvement turns, they found that the improvement followed a law, a smooth, predictable followed a law, a smooth, predictable followed a law, a smooth, predictable curve across four orders of magnitude of curve across four orders of magnitude of curve across four orders of magnitude of attempts. More attempts, more problems attempts. More attempts, more problems attempts. More attempts, more problems solved. It was like clockwork. That's solved. It was like clockwork. That's solved. It was like clockwork. That's academia, right? But in production, this academia, right? But in production, this academia, right? But in production, this holds up as well. Anthropic built a holds up as well. Anthropic built a holds up as well. Anthropic built a research system out of multiple agents research system out of multiple agents research system out of multiple agents and then studied what actually predicted and then studied what actually predicted and then studied what actually predicted whether a run was good or bad. Was it whether a run was good or bad. Was it whether a run was good or bad. Was it prompt wording? It didn't matter as much prompt wording? It didn't matter as much prompt wording? It didn't matter as much as you would think. Not at all really.

  6. as you would think. Not at all really. as you would think. Not at all really. The single biggest factor explaining 80% The single biggest factor explaining 80% The single biggest factor explaining 80% of the difference between a good run of the difference between a good run of the difference between a good run that resulted in a solved problem and a that resulted in a solved problem and a that resulted in a solved problem and a bad run was token spend. Basically, how bad run was token spend. Basically, how bad run was token spend. Basically, how much the system was allowed to think. much the system was allowed to think. much the system was allowed to think. and Anthropic team of agents beat the and Anthropic team of agents beat the and Anthropic team of agents beat the frontier model at the time working alone frontier model at the time working alone frontier model at the time working alone by 90.2%. by 90.2%. by 90.2%. Their explanation, Anthropic's Their explanation, Anthropic's Their explanation, Anthropic's explanation in plain terms is that a explanation in plain terms is that a explanation in plain terms is that a team of agents is how you spend more team of agents is how you spend more team of agents is how you spend more tokens than one agent can usefully hold tokens than one agent can usefully hold tokens than one agent can usefully hold and that has traction against hard and that has traction against hard and that has traction against hard problems. A multi- aent run can cost 10, problems. A multi- aent run can cost 10, problems. A multi- aent run can cost 10, 15, 20, 30 times more than a single 15, 20, 30 times more than a single 15, 20, 30 times more than a single agent run or even more than that. And so agent run or even more than that. And so agent run or even more than that. And so for a long time we were held back for a long time we were held back for a long time we were held back because token spend got expensive and it because token spend got expensive and it because token spend got expensive and it just wasn't practical for individuals to just wasn't practical for individuals to just wasn't practical for individuals to spend on multi- aent teams. That's spend on multi- aent teams. That's spend on multi- aent teams. That's changed. This is why Ringer matters. changed. This is why Ringer matters. changed. This is why Ringer matters. Ringer matters because it brings the Ringer matters because it brings the Ringer matters because it brings the power of multi- aent problem solving power of multi- aent problem solving power of multi- aent problem solving within reach of individuals because within reach of individuals because within reach of individuals because token costs are no longer prohibitive.

  7. token costs are no longer prohibitive. token costs are no longer prohibitive. Because for a long time, budgeting Because for a long time, budgeting Because for a long time, budgeting agents meant capping work, not just agents meant capping work, not just agents meant capping work, not just buying a bigger pile of thinking. And buying a bigger pile of thinking. And buying a bigger pile of thinking. And that's what's changed as we've had that's what's changed as we've had that's what's changed as we've had better open- source models in the last better open- source models in the last better open- source models in the last month or so. Which is exactly why the month or so. Which is exactly why the month or so. Which is exactly why the whole game is knowing which tasks are whole game is knowing which tasks are whole game is knowing which tasks are worth it. Because here's the thing, if worth it. Because here's the thing, if worth it. Because here's the thing, if more tokens reliably meant better more tokens reliably meant better more tokens reliably meant better answers, this would be a very short answers, this would be a very short answers, this would be a very short video. It would just be like spend more. video. It would just be like spend more. video. It would just be like spend more. But there's a catch. And the catch is But there's a catch. And the catch is But there's a catch. And the catch is the half of the Stanford study that the half of the Stanford study that the half of the Stanford study that nobody quotes. Same study, by the way, nobody quotes. Same study, by the way, nobody quotes. Same study, by the way, same cheap model. Remember, still back same cheap model. Remember, still back same cheap model. Remember, still back in 2024, they kept pushing. They pushed in 2024, they kept pushing. They pushed in 2024, they kept pushing. They pushed to a 100 attempts. They pushed to a,000 to a 100 attempts. They pushed to a,000 to a 100 attempts. They pushed to a,000 attempts at this bug solving. They attempts at this bug solving. They attempts at this bug solving. They pushed to 10,000 attempts. And they pushed to 10,000 attempts. And they pushed to 10,000 attempts. And they tracked a simple question. For what tracked a simple question. For what tracked a simple question. For what share of problems does a correct answer share of problems does a correct answer share of problems does a correct answer exist somewhere in the pile of attempts? exist somewhere in the pile of attempts? exist somewhere in the pile of attempts? At 10,000 attempts, over 95% At 10,000 attempts, over 95% At 10,000 attempts, over 95% of the 10,000 attempt runs contained a of the 10,000 attempt runs contained a of the 10,000 attempt runs contained a correct answer. The right answer was correct answer. The right answer was correct answer. The right answer was almost always in there. And now the real almost always in there. And now the real almost always in there. And now the real question comes up. How do you find it?

  8. question comes up. How do you find it? question comes up. How do you find it? How do you find it? Where where do you How do you find it? Where where do you How do you find it? Where where do you get the validation that lets you see get the validation that lets you see get the validation that lets you see that the right answer is there? If that the right answer is there? If that the right answer is there? If you're giving your agent multiple tries, you're giving your agent multiple tries, you're giving your agent multiple tries, is there an automatic checker somewhere, is there an automatic checker somewhere, is there an automatic checker somewhere, a test suite? Stanford tested that too. a test suite? Stanford tested that too. a test suite? Stanford tested that too. And what they found was where there was And what they found was where there was And what they found was where there was an automatic checker, a test suite, an automatic checker, a test suite, an automatic checker, a test suite, something mechanical that could grade something mechanical that could grade something mechanical that could grade each attempt. Yeah, they could find the each attempt. Yeah, they could find the each attempt. Yeah, they could find the answers. Coverage turned straight into answers. Coverage turned straight into answers. Coverage turned straight into results. But where there was no checker results. But where there was no checker results. But where there was no checker and the model had to pick the best and the model had to pick the best and the model had to pick the best answer out of the pile, they tried answer out of the pile, they tried answer out of the pile, they tried majority voting. They tried rewarding majority voting. They tried rewarding majority voting. They tried rewarding models. All of that, everything stalled models. All of that, everything stalled models. All of that, everything stalled out at 100 attempts. In other words, you out at 100 attempts. In other words, you out at 100 attempts. In other words, you need evals. You need external validation need evals. You need external validation need evals. You need external validation in order to scale multi- aent systems in order to scale multi- aent systems in order to scale multi- aent systems because look at that gap, right? The because look at that gap, right? The because look at that gap, right? The right answer exists, but nobody can tell right answer exists, but nobody can tell right answer exists, but nobody can tell which one it is. Every dollar spent past which one it is. Every dollar spent past which one it is. Every dollar spent past that line buys attempts that are that line buys attempts that are that line buys attempts that are generated. The answer is probably in generated. The answer is probably in generated. The answer is probably in there, but they're never found. This is there, but they're never found. This is there, but they're never found. This is a failure of multi- aent systems. And a failure of multi- aent systems. And a failure of multi- aent systems. And that shaded area, that's money. That's that shaded area, that's money. That's that shaded area, that's money. That's money you're spending you're not getting money you're spending you're not getting money you're spending you're not getting back if you don't design your systems back if you don't design your systems back if you don't design your systems correctly. And there's a second caveat correctly. And there's a second caveat correctly. And there's a second caveat on the other side. A single agent cannot on the other side. A single agent cannot on the other side. A single agent cannot absorb unlimited spending either because absorb unlimited spending either because absorb unlimited spending either because everything it reads and does piles into everything it reads and does piles into everything it reads and does piles into a context window. As the context window a context window. As the context window a context window. As the context window fills up, quality tends to drop even fills up, quality tends to drop even fills up, quality tends to drop even with new techniques like autocompaction.

  9. with new techniques like autocompaction. with new techniques like autocompaction. And ultimately, the model will have to And ultimately, the model will have to And ultimately, the model will have to either delegate to other agents or either delegate to other agents or either delegate to other agents or abandon the task. And that's actually abandon the task. And that's actually abandon the task. And that's actually how a lot of these newer agentic models how a lot of these newer agentic models how a lot of these newer agentic models are handling long tasks. In other words, are handling long tasks. In other words, are handling long tasks. In other words, with tools like chat GPT 5.6 Six, you with tools like chat GPT 5.6 Six, you with tools like chat GPT 5.6 Six, you might be inadvertently running multi- might be inadvertently running multi- might be inadvertently running multi- aent solutions if you give it a big aent solutions if you give it a big aent solutions if you give it a big enough task. Part of my goal is to help enough task. Part of my goal is to help enough task. Part of my goal is to help you be intentional about that. One is a you be intentional about that. One is a you be intentional about that. One is a memory constraint, one is an eval memory constraint, one is an eval memory constraint, one is an eval constraint. You have to be able to constraint. You have to be able to constraint. You have to be able to evaluate evaluate evaluate systems and you have to know when your systems and you have to know when your systems and you have to know when your task is big enough that one agent can't task is big enough that one agent can't task is big enough that one agent can't hold it. It's a memory constraint that hold it. It's a memory constraint that hold it. It's a memory constraint that would drive you into a multi-agent would drive you into a multi-agent would drive you into a multi-agent system. So with that in mind, that system. So with that in mind, that system. So with that in mind, that equips us to think more deliberately equips us to think more deliberately equips us to think more deliberately about where we apply multi- aent about where we apply multi- aent about where we apply multi- aent problems. Every team of agents design problems. Every team of agents design problems. Every team of agents design that actually works is an answer to one that actually works is an answer to one that actually works is an answer to one of these two problems. Everything else of these two problems. Everything else of these two problems. Everything else is just more agents. The need to split is just more agents. The need to split is just more agents. The need to split capacity is changing over time as agents capacity is changing over time as agents capacity is changing over time as agents get more capable. And that's one of the get more capable. And that's one of the get more capable. And that's one of the things that has made this problem space things that has made this problem space things that has made this problem space really hard. It's hard to know what to really hard. It's hard to know what to really hard. It's hard to know what to assign to an agent when agents keep assign to an agent when agents keep assign to an agent when agents keep getting smarter. But there's a second getting smarter. But there's a second getting smarter. But there's a second reason to split out work that has reason to split out work that has reason to split out work that has nothing to do with capacity and that nothing to do with capacity and that nothing to do with capacity and that really helps us when we're trying to really helps us when we're trying to really helps us when we're trying to assign multiple agents work. Some work assign multiple agents work. Some work assign multiple agents work. Some work has parts that inherently has to be done has parts that inherently has to be done has parts that inherently has to be done by different minds or different agents.

  10. by different minds or different agents. by different minds or different agents. Not because one agent or one mind lacks Not because one agent or one mind lacks Not because one agent or one mind lacks the skill, but because the parts poison the skill, but because the parts poison the skill, but because the parts poison each other. The auditor who also kept each other. The auditor who also kept each other. The auditor who also kept the books isn't a worse auditor. He's the books isn't a worse auditor. He's the books isn't a worse auditor. He's just not an auditor at all. Right? You just not an auditor at all. Right? You just not an auditor at all. Right? You need to have two different roles in that need to have two different roles in that need to have two different roles in that problem. Peer review only works because problem. Peer review only works because problem. Peer review only works because the reviewer didn't write the paper. the reviewer didn't write the paper. the reviewer didn't write the paper. Your bank will not let the person who Your bank will not let the person who Your bank will not let the person who enters a payment be the person who enters a payment be the person who enters a payment be the person who approves it. Not because they're approves it. Not because they're approves it. Not because they're dishonest, but because for centuries, dishonest, but because for centuries, dishonest, but because for centuries, the way to get reliable work done was to the way to get reliable work done was to the way to get reliable work done was to make sure that you had checks and make sure that you had checks and make sure that you had checks and balances made of multiple minds. And balances made of multiple minds. And balances made of multiple minds. And agents add one thing to this trick that agents add one thing to this trick that agents add one thing to this trick that has never existed before. You cannot has never existed before. You cannot has never existed before. You cannot unknow something. You've read your own unknow something. You've read your own unknow something. You've read your own product page a thousand times. You'll product page a thousand times. You'll product page a thousand times. You'll never see it the way a stranger does, never see it the way a stranger does, never see it the way a stranger does, right? But you can now start a mind that right? But you can now start a mind that right? But you can now start a mind that has never seen it. You can start an has never seen it. You can start an has never seen it. You can start an agent that has never seen it. So you can agent that has never seen it. So you can agent that has never seen it. So you can have fresh eyes on demand for the first have fresh eyes on demand for the first have fresh eyes on demand for the first time in history. Now, where I see this time in history. Now, where I see this time in history. Now, where I see this being most useful is when there are being most useful is when there are being most useful is when there are genuine conflicts of interest you want genuine conflicts of interest you want genuine conflicts of interest you want to balance in an agent system. Things to balance in an agent system. Things to balance in an agent system. Things where you need to review something twice where you need to review something twice where you need to review something twice to ensure it's done correctly. A to ensure it's done correctly. A to ensure it's done correctly. A contract comes to mind, a draft comes to contract comes to mind, a draft comes to contract comes to mind, a draft comes to mind, a plan comes to mind. So, put all mind, a plan comes to mind. So, put all mind, a plan comes to mind. So, put all of this together and you get the agent of this together and you get the agent of this together and you get the agent test. Four things you can estimate about test. Four things you can estimate about test. Four things you can estimate about any task on your desk in about a minute.

  11. any task on your desk in about a minute. any task on your desk in about a minute. First, size. Is the task bigger than First, size. Is the task bigger than First, size. Is the task bigger than what one agent can hold at full quality? what one agent can hold at full quality? what one agent can hold at full quality? Your calendar is not bigger than what Your calendar is not bigger than what Your calendar is not bigger than what one agent can hold, for example, right? one agent can hold, for example, right? one agent can hold, for example, right? It fits in a corner of a context window. It fits in a corner of a context window. It fits in a corner of a context window. It's not a problem. your last quarter of It's not a problem. your last quarter of It's not a problem. your last quarter of email. Well, it depends on how much email. Well, it depends on how much email. Well, it depends on how much email you get, but for a lot of us, that email you get, but for a lot of us, that email you get, but for a lot of us, that fits in a context window, too. Question fits in a context window, too. Question fits in a context window, too. Question two, but a pile of aundred documents, two, but a pile of aundred documents, two, but a pile of aundred documents, that might scale out of the context that might scale out of the context that might scale out of the context window. A thousand documents definitely window. A thousand documents definitely window. A thousand documents definitely would. Independence, can the parts be would. Independence, can the parts be would. Independence, can the parts be done without knowing what the other done without knowing what the other done without knowing what the other parts did? Reading a pile of documents parts did? Reading a pile of documents parts did? Reading a pile of documents actually splits really well because one actually splits really well because one actually splits really well because one reader agent can read any given document reader agent can read any given document reader agent can read any given document and they never need to talk. Coding and they never need to talk. Coding and they never need to talk. Coding sometimes splits and sometimes doesn't. sometimes splits and sometimes doesn't. sometimes splits and sometimes doesn't. It depends on how you tell the agent to It depends on how you tell the agent to It depends on how you tell the agent to organize files. If agents are able to organize files. If agents are able to organize files. If agents are able to organize code into independent parts, organize code into independent parts, organize code into independent parts, then multiple agents can work on coding then multiple agents can work on coding then multiple agents can work on coding problems and it can be extremely problems and it can be extremely problems and it can be extremely effective. Three, separation of effective. Three, separation of effective. Three, separation of concerns. Do any parts of this task need concerns. Do any parts of this task need concerns. Do any parts of this task need to be done by different minds? A real to be done by different minds? A real to be done by different minds? A real critic who didn't write the draft, as an critic who didn't write the draft, as an critic who didn't write the draft, as an example, an overview written by someone example, an overview written by someone example, an overview written by someone who didn't do the reading. output that's who didn't do the reading. output that's who didn't do the reading. output that's kept apart from the inputs that you kept apart from the inputs that you kept apart from the inputs that you drive and analyze because you need a drive and analyze because you need a drive and analyze because you need a separate frame for that. If you're separate frame for that. If you're separate frame for that. If you're thinking about that, you're thinking thinking about that, you're thinking thinking about that, you're thinking about a team of agents. Question four, about a team of agents. Question four, about a team of agents. Question four, and this is a critical one, and this is a critical one, and this is a critical one, checkability. Remember I talked about checkability. Remember I talked about checkability. Remember I talked about that verification evals. Is checking an that verification evals. Is checking an that verification evals. Is checking an answer a whole lot cheaper than answer a whole lot cheaper than answer a whole lot cheaper than producing one? A test suite, an exit producing one? A test suite, an exit producing one? A test suite, an exit code, a source document that you can

  12. code, a source document that you can code, a source document that you can point at, something where you can glance point at, something where you can glance point at, something where you can glance at it and say this is right or this is at it and say this is right or this is at it and say this is right or this is wrong. If checking is almost free, wrong. If checking is almost free, wrong. If checking is almost free, however you do it, then every extra however you do it, then every extra however you do it, then every extra attempt, including if checkability is attempt, including if checkability is attempt, including if checkability is expensive, on the other hand, if it expensive, on the other hand, if it expensive, on the other hand, if it takes a while to check, you're going to takes a while to check, you're going to takes a while to check, you're going to top out on the value of your multi- aent top out on the value of your multi- aent top out on the value of your multi- aent systems relatively quickly. And the systems relatively quickly. And the systems relatively quickly. And the Stanford research suggests that about a Stanford research suggests that about a Stanford research suggests that about a 100 tries, you're just not going to get 100 tries, you're just not going to get 100 tries, you're just not going to get a lot more value out of it. If you put a lot more value out of it. If you put a lot more value out of it. If you put all these four together, you get an all these four together, you get an all these four together, you get an overall verdict. The problem might be overall verdict. The problem might be overall verdict. The problem might be small, in which case it's just a chat small, in which case it's just a chat small, in which case it's just a chat back and forth. It might fit in a back and forth. It might fit in a back and forth. It might fit in a context window, in which case it's an context window, in which case it's an context window, in which case it's an agent with a goal. It's genuinely agent with a goal. It's genuinely agent with a goal. It's genuinely useful. It works alone. It checks its useful. It works alone. It checks its useful. It works alone. It checks its own work and it gets the job done. Or it own work and it gets the job done. Or it own work and it gets the job done. Or it may be bigger than one perspective or may be bigger than one perspective or may be bigger than one perspective or have parts of the task that need have parts of the task that need have parts of the task that need separate minds or perspectives to work separate minds or perspectives to work separate minds or perspectives to work well. That's a team of agents. There's well. That's a team of agents. There's well. That's a team of agents. There's also a fourth option, which is that also a fourth option, which is that also a fourth option, which is that there's also a fourth option, which is there's also a fourth option, which is there's also a fourth option, which is the one that saves you the most money. the one that saves you the most money. the one that saves you the most money. Ironically, maybe you don't need the AI Ironically, maybe you don't need the AI Ironically, maybe you don't need the AI at all. Maybe it's a judgment call that at all. Maybe it's a judgment call that at all. Maybe it's a judgment call that you need to sit with and make on your you need to sit with and make on your you need to sit with and make on your own and that still needs to happen in own and that still needs to happen in own and that still needs to happen in the age of AI. Okay, card one, the the age of AI. Okay, card one, the the age of AI. Okay, card one, the scheduling thing. I think you can guess scheduling thing. I think you can guess scheduling thing. I think you can guess the answer here. My version was find me the answer here. My version was find me the answer here. My version was find me a gym slot this week that fits around my a gym slot this week that fits around my a gym slot this week that fits around my meetings. That is absolutely a single meetings. That is absolutely a single meetings. That is absolutely a single agent task. It is not a multi- aent agent task. It is not a multi- aent agent task. It is not a multi- aent task. It's a relatively quick 5-minute task. It's a relatively quick 5-minute task. It's a relatively quick 5-minute task. You can feed it to Claude. You can task. You can feed it to Claude. You can task. You can feed it to Claude. You can feed it to Codeex. It's not going to be feed it to Codeex. It's not going to be feed it to Codeex. It's not going to be a problem for today's models. This tier a problem for today's models. This tier a problem for today's models. This tier is solved. If this impressed you, you is solved. If this impressed you, you is solved. If this impressed you, you should start to raise the bar for what should start to raise the bar for what should start to raise the bar for what you think AI can do because this has you think AI can do because this has you think AI can do because this has been an easy one for a bit now. My been an easy one for a bit now. My been an easy one for a bit now. My version is a complicated document version is a complicated document version is a complicated document review. I have about 40 different tools

  13. review. I have about 40 different tools review. I have about 40 different tools I use to run my media business. They all I use to run my media business. They all I use to run my media business. They all have different renewal dates. They all have different renewal dates. They all have different renewal dates. They all have different contracts. They all have have different contracts. They all have have different contracts. They all have different emails that they're sending me different emails that they're sending me different emails that they're sending me all the time. And what I need is one all the time. And what I need is one all the time. And what I need is one single dashboard that shows me my single dashboard that shows me my single dashboard that shows me my renewal dates across all of them. And renewal dates across all of them. And renewal dates across all of them. And that also gives me a sense of where I'm that also gives me a sense of where I'm that also gives me a sense of where I'm actually using these tools versus where actually using these tools versus where actually using these tools versus where I'm not. And that in turn gives me an I'm not. And that in turn gives me an I'm not. And that in turn gives me an intelligent way to assess, do I want to intelligent way to assess, do I want to intelligent way to assess, do I want to build an in-home solution to some of build an in-home solution to some of build an in-home solution to some of these or do I want to depend on ones these or do I want to depend on ones these or do I want to depend on ones that I'm using all the time and say no, that I'm using all the time and say no, that I'm using all the time and say no, these are worth the money. It's hard to these are worth the money. It's hard to these are worth the money. It's hard to make that decision intentionally unless make that decision intentionally unless make that decision intentionally unless I know when are the renewal dates coming I know when are the renewal dates coming I know when are the renewal dates coming up, how am I actually using it, that the up, how am I actually using it, that the up, how am I actually using it, that the tool, etc., etc. So there's there's tool, etc., etc. So there's there's tool, etc., etc. So there's there's easily thousands of pages of documents easily thousands of pages of documents easily thousands of pages of documents in this pile. Whether you count the in this pile. Whether you count the in this pile. Whether you count the emails, whether you count the contracts emails, whether you count the contracts emails, whether you count the contracts that I'm agreeing to, whether you count that I'm agreeing to, whether you count that I'm agreeing to, whether you count the logs of usage that I'm piling up as the logs of usage that I'm piling up as the logs of usage that I'm piling up as I use these tools. So when you look at I use these tools. So when you look at I use these tools. So when you look at all of that together, what you need is a all of that together, what you need is a all of that together, what you need is a team of agents to solve that problem for team of agents to solve that problem for team of agents to solve that problem for you to actually go through the different you to actually go through the different you to actually go through the different tools that you use to pop up. Okay, this tools that you use to pop up. Okay, this tools that you use to pop up. Okay, this is the renewal date. this is your actual is the renewal date. this is your actual is the renewal date. this is your actual usage that we're seeing from these usage that we're seeing from these usage that we're seeing from these tools. These are the tokens you're using tools. These are the tokens you're using tools. These are the tokens you're using or if it's not an AI tool, it's or if it's not an AI tool, it's or if it's not an AI tool, it's something else, right? This is the something else, right? This is the something else, right? This is the number of times you've logged in, etc., number of times you've logged in, etc., number of times you've logged in, etc., etc. And then to give me a sense of etc. And then to give me a sense of etc. And then to give me a sense of which are the tools that are open for which are the tools that are open for which are the tools that are open for rebuilding, right? If you look at it by rebuilding, right? If you look at it by rebuilding, right? If you look at it by complexity, you look at it by usage, you complexity, you look at it by usage, you complexity, you look at it by usage, you look at it by cost, you start to get a look at it by cost, you start to get a look at it by cost, you start to get a matrix that gives you business decisions matrix that gives you business decisions matrix that gives you business decisions you can make about your tooling. That's you can make about your tooling. That's you can make about your tooling. That's a complicated task. That's much more

  14. a complicated task. That's much more a complicated task. That's much more than an agent can hold. It is something than an agent can hold. It is something than an agent can hold. It is something I can evaluate fairly easily because I can evaluate fairly easily because I can evaluate fairly easily because when I look at it, I can say, "Oh, yeah, when I look at it, I can say, "Oh, yeah, when I look at it, I can say, "Oh, yeah, I do use that tool that much." You know, I do use that tool that much." You know, I do use that tool that much." You know, I I I do not use this tool that much. I I I I do not use this tool that much. I I I I do not use this tool that much. I can look at it and say, "Oh, yeah, I do can look at it and say, "Oh, yeah, I do can look at it and say, "Oh, yeah, I do use superhuman a ton." Right? That's not use superhuman a ton." Right? That's not use superhuman a ton." Right? That's not a surprise. Or maybe another SAS tool a surprise. Or maybe another SAS tool a surprise. Or maybe another SAS tool comes up and I'm like, I haven't used comes up and I'm like, I haven't used comes up and I'm like, I haven't used that in in ages. What is this? Like, I I that in in ages. What is this? Like, I I that in in ages. What is this? Like, I I don't use it at all. Or maybe there's don't use it at all. Or maybe there's don't use it at all. Or maybe there's one that I look at, it's a candidate for one that I look at, it's a candidate for one that I look at, it's a candidate for disruption, at least for me, where I disruption, at least for me, where I disruption, at least for me, where I want to build it internally. And I think want to build it internally. And I think want to build it internally. And I think that that's very much a per business that that's very much a per business that that's very much a per business kind of decision. If I can use agents to kind of decision. If I can use agents to kind of decision. If I can use agents to do that work, one, I never would have do that work, one, I never would have do that work, one, I never would have had time to do that myself. Two, it had time to do that myself. Two, it had time to do that myself. Two, it helps me make better decisions that helps me make better decisions that helps me make better decisions that focus my leverage as a business. And focus my leverage as a business. And focus my leverage as a business. And three, I can actually make sure that I'm three, I can actually make sure that I'm three, I can actually make sure that I'm maximizing where my dollars are going, maximizing where my dollars are going, maximizing where my dollars are going, whether those are dollars for SAS whether those are dollars for SAS whether those are dollars for SAS contracts or whether those are dollars contracts or whether those are dollars contracts or whether those are dollars for agents, because the agent run isn't for agents, because the agent run isn't for agents, because the agent run isn't free. And I can feel good that I'm free. And I can feel good that I'm free. And I can feel good that I'm actually putting the agents against a actually putting the agents against a actually putting the agents against a problem that has real ROI. Now, let's problem that has real ROI. Now, let's problem that has real ROI. Now, let's spend 60 seconds on the machinery of how spend 60 seconds on the machinery of how spend 60 seconds on the machinery of how this works because this is where the two this works because this is where the two this works because this is where the two limits that I talked about earlier show limits that I talked about earlier show limits that I talked about earlier show up as actual design in the multi-agent up as actual design in the multi-agent up as actual design in the multi-agent harness I put together. And I went deep harness I put together. And I went deep harness I put together. And I went deep on this in Wednesday's video. So, I'm on this in Wednesday's video. So, I'm on this in Wednesday's video. So, I'm just giving you this shape here. Every just giving you this shape here. Every just giving you this shape here. Every task gets a spec. It's written once by task gets a spec. It's written once by task gets a spec. It's written once by the strongest model, which then never the strongest model, which then never the strongest model, which then never touches the work again. Every finished touches the work again. Every finished touches the work again. Every finished task does get a check, and the check is task does get a check, and the check is task does get a check, and the check is mechanical, right? The source has to be mechanical, right? The source has to be mechanical, right? The source has to be attached and match the task or the entry attached and match the task or the entry attached and match the task or the entry is rejected. The agent's opinion of its is rejected. The agent's opinion of its is rejected. The agent's opinion of its own work is not evidence. A failed task own work is not evidence. A failed task own work is not evidence. A failed task gets a retry with the failure included

  15. gets a retry with the failure included gets a retry with the failure included and every result feeds a running and every result feeds a running and every result feeds a running scorecard that I can keep an eye on so scorecard that I can keep an eye on so scorecard that I can keep an eye on so it's easy for me to understand how this it's easy for me to understand how this it's easy for me to understand how this particular agent run is working. That's particular agent run is working. That's particular agent run is working. That's how Ringer works. The idea is that you how Ringer works. The idea is that you how Ringer works. The idea is that you take this whole complicated process of take this whole complicated process of take this whole complicated process of getting validation that allows you to getting validation that allows you to getting validation that allows you to use these multi- aent systems use these multi- aent systems use these multi- aent systems effectively and you turn it into effectively and you turn it into effectively and you turn it into something that is as easy as watching a something that is as easy as watching a something that is as easy as watching a dashboard stream while your agents solve dashboard stream while your agents solve dashboard stream while your agents solve a hard problem. And this setup saves a a hard problem. And this setup saves a a hard problem. And this setup saves a ton on tokens and token costs. One of ton on tokens and token costs. One of ton on tokens and token costs. One of the things that I called out in my video the things that I called out in my video the things that I called out in my video on Wednesday is that I was able to on Wednesday is that I was able to on Wednesday is that I was able to reduce Fable 5 costs by about 10x and reduce Fable 5 costs by about 10x and reduce Fable 5 costs by about 10x and keep the brain power of Fable 5 because keep the brain power of Fable 5 because keep the brain power of Fable 5 because I allowed Fable to be the brains behind I allowed Fable to be the brains behind I allowed Fable to be the brains behind the system and organize and make the system and organize and make the system and organize and make judgment calls on a multi- aent system. judgment calls on a multi- aent system. judgment calls on a multi- aent system. But all of the tokens for doing the work But all of the tokens for doing the work But all of the tokens for doing the work got farmed out to much much cheaper got farmed out to much much cheaper got farmed out to much much cheaper worker agents. So here you can see the worker agents. So here you can see the worker agents. So here you can see the result of the run. You can see tools result of the run. You can see tools result of the run. You can see tools ranked by cost against usage. You can ranked by cost against usage. You can ranked by cost against usage. You can see renewals with their notice see renewals with their notice see renewals with their notice deadlines. You can see any risky terms deadlines. You can see any risky terms deadlines. You can see any risky terms in the contracts. And you can see a in the contracts. And you can see a in the contracts. And you can see a recommendation per tool, whether that's recommendation per tool, whether that's recommendation per tool, whether that's keep it or negotiate it or cancel it or keep it or negotiate it or cancel it or keep it or negotiate it or cancel it or even build your own with sources even build your own with sources even build your own with sources attached. Now, here we come to the part attached. Now, here we come to the part attached. Now, here we come to the part that most AI demos skip. It's not done that most AI demos skip. It's not done that most AI demos skip. It's not done when I say it's done and show you. It's when I say it's done and show you. It's when I say it's done and show you. It's done when I show you the bill. How much done when I show you the bill. How much done when I show you the bill. How much did this cost me to do all of this work?

  16. did this cost me to do all of this work? did this cost me to do all of this work? And Wednesday's video is how this works. And Wednesday's video is how this works. And Wednesday's video is how this works. You use an expensive model to plan and You use an expensive model to plan and You use an expensive model to plan and judge like Fable and a lot of cheap judge like Fable and a lot of cheap judge like Fable and a lot of cheap models to do all the coding and burn all models to do all the coding and burn all models to do all the coding and burn all the tokens and do the research and the tokens and do the research and the tokens and do the research and actually execute the multi- aent actually execute the multi- aent actually execute the multi- aent workflows. This is what makes it workflows. This is what makes it workflows. This is what makes it possible for individuals to go through possible for individuals to go through possible for individuals to go through the process I described for work on the process I described for work on the process I described for work on their desk for all of us, right? And to their desk for all of us, right? And to their desk for all of us, right? And to say this is a multi- aent problem. I'm say this is a multi- aent problem. I'm say this is a multi- aent problem. I'm going to apply multiple agents to this going to apply multiple agents to this going to apply multiple agents to this task and feel good that I'm not burning task and feel good that I'm not burning task and feel good that I'm not burning a lot of money. By the way, this did not a lot of money. By the way, this did not a lot of money. By the way, this did not take long to set up like under an hour. take long to set up like under an hour. take long to set up like under an hour. Like this is not a hard setup task. And Like this is not a hard setup task. And Like this is not a hard setup task. And if you have something with thousands of if you have something with thousands of if you have something with thousands of documents and then another task after documents and then another task after documents and then another task after that and another task after that, that and another task after that, that and another task after that, spending under an hour to set it up, spending under an hour to set it up, spending under an hour to set it up, that's great. Like that's not a problem. that's great. Like that's not a problem. that's great. Like that's not a problem. Now, you may have other kinds of piles. Now, you may have other kinds of piles. Now, you may have other kinds of piles. I'm assuming you do. Maybe it's a I'm assuming you do. Maybe it's a I'm assuming you do. Maybe it's a project handoff. Maybe someone's leaving project handoff. Maybe someone's leaving project handoff. Maybe someone's leaving and and everything lives in meeting and and everything lives in meeting and and everything lives in meeting notes and and chat threads and a notes and and chat threads and a notes and and chat threads and a half-finished stack of documents in a half-finished stack of documents in a half-finished stack of documents in a folder. And the new person has to come folder. And the new person has to come folder. And the new person has to come in and they're going to spend a week in and they're going to spend a week in and they're going to spend a week digging through and getting oriented. digging through and getting oriented. digging through and getting oriented. and you want to speed that up. Hey, same and you want to speed that up. Hey, same and you want to speed that up. Hey, same shape. Sounds like you need to apply a shape. Sounds like you need to apply a shape. Sounds like you need to apply a multiple agent solution and you can multiple agent solution and you can multiple agent solution and you can actually get a briefing for your new actually get a briefing for your new actually get a briefing for your new person that gets them set up and saves person that gets them set up and saves person that gets them set up and saves them dozens of hours. Maybe it's the them dozens of hours. Maybe it's the them dozens of hours. Maybe it's the research archive at the end of a research archive at the end of a research archive at the end of a project. Maybe it's an inbox quarter and project. Maybe it's an inbox quarter and project. Maybe it's an inbox quarter and you have thousands and tens of thousands you have thousands and tens of thousands you have thousands and tens of thousands of emails and you need to sort through of emails and you need to sort through of emails and you need to sort through them all. Maybe it's actually building a them all. Maybe it's actually building a them all. Maybe it's actually building a cold lead follow-up system and you're cold lead follow-up system and you're cold lead follow-up system and you're trying to get your cold pipeline going trying to get your cold pipeline going trying to get your cold pipeline going for sales. You have piles regardless.

  17. for sales. You have piles regardless. for sales. You have piles regardless. Recognize where you have piles of work Recognize where you have piles of work Recognize where you have piles of work and look at them as multi- aent and look at them as multi- aent and look at them as multi- aent susceptible. Other examples in our susceptible. Other examples in our susceptible. Other examples in our personal lives would include cases where personal lives would include cases where personal lives would include cases where we have lots of medical records, cases we have lots of medical records, cases we have lots of medical records, cases where we have lots of bank statements where we have lots of bank statements where we have lots of bank statements and we do want to do analysis of our of and we do want to do analysis of our of and we do want to do analysis of our of our spending patterns of our financial our spending patterns of our financial our spending patterns of our financial health. Now, these want careful setup, health. Now, these want careful setup, health. Now, these want careful setup, right? You want to have exports landing right? You want to have exports landing right? You want to have exports landing in a folder you can control agents that in a folder you can control agents that in a folder you can control agents that only ever read that folder. You ideally only ever read that folder. You ideally only ever read that folder. You ideally want to run it on a machine you own. so want to run it on a machine you own. so want to run it on a machine you own. so your financial data and your health data your financial data and your health data your financial data and your health data don't leak out. So you can see the shape don't leak out. So you can see the shape don't leak out. So you can see the shape of the task, but if it's certain of the task, but if it's certain of the task, but if it's certain sensitive information, you may want to sensitive information, you may want to sensitive information, you may want to look at running it locally on a Mac look at running it locally on a Mac look at running it locally on a Mac Mini. And I have got videos on how to Mini. And I have got videos on how to Mini. And I have got videos on how to set that up and have your own stack as set that up and have your own stack as set that up and have your own stack as well and you can dig into that also. well and you can dig into that also. well and you can dig into that also. Okay, last card. Which candidate do you Okay, last card. Which candidate do you Okay, last card. Which candidate do you hire? Judgment call type problems. What hire? Judgment call type problems. What hire? Judgment call type problems. What do you name a product? What direction do you name a product? What direction do you name a product? What direction does the business take? These are tasks does the business take? These are tasks does the business take? These are tasks that I hear people saying they hand to that I hear people saying they hand to that I hear people saying they hand to AI. Let's not do that. And in fact, I've AI. Let's not do that. And in fact, I've AI. Let's not do that. And in fact, I've heard people say, "I don't do that, but heard people say, "I don't do that, but heard people say, "I don't do that, but I let AI help me make judgment calls on I let AI help me make judgment calls on I let AI help me make judgment calls on candidates." I've let AI help me make candidates." I've let AI help me make candidates." I've let AI help me make judgment calls on which product judgment calls on which product judgment calls on which product direction to go in. I understand that we direction to go in. I understand that we direction to go in. I understand that we want AI to support us. I'm all for want AI to support us. I'm all for want AI to support us. I'm all for research to support us. I'm all for research to support us. I'm all for research to support us. I'm all for opinions from AI to support us. But if opinions from AI to support us. But if opinions from AI to support us. But if you don't have a strong instinct around you don't have a strong instinct around you don't have a strong instinct around what is correct and you're not willing what is correct and you're not willing what is correct and you're not willing to apply your human judgment, this is a to apply your human judgment, this is a to apply your human judgment, this is a situation where you're going to make a situation where you're going to make a situation where you're going to make a mistake because no frontier model is

  18. mistake because no frontier model is mistake because no frontier model is going to be able to beat an expert at going to be able to beat an expert at going to be able to beat an expert at the thing they are most expert in. And the thing they are most expert in. And the thing they are most expert in. And I've actually talked with people who are I've actually talked with people who are I've actually talked with people who are experts in their field at product at experts in their field at product at experts in their field at product at engineering at at investing at at engineering at at investing at at engineering at at investing at at running businesses and they are using running businesses and they are using running businesses and they are using Fable 5 or 5.6 from OpenAI and they say Fable 5 or 5.6 from OpenAI and they say Fable 5 or 5.6 from OpenAI and they say I am better at my job and the decisions I am better at my job and the decisions I am better at my job and the decisions I make because I sort of run chats with I make because I sort of run chats with I make because I sort of run chats with Fable 5. I run chats with chat GPT 5.6 Fable 5. I run chats with chat GPT 5.6 Fable 5. I run chats with chat GPT 5.6 and I get a a wall I can bounce ideas and I get a a wall I can bounce ideas and I get a a wall I can bounce ideas off of. But the instincts of these off of. But the instincts of these off of. But the instincts of these models are not worldclass. models are not worldclass. models are not worldclass. They are not strong enough that you can They are not strong enough that you can They are not strong enough that you can find the jean sequa. You can find that find the jean sequa. You can find that find the jean sequa. You can find that that thing that is unspeakable or that thing that is unspeakable or that thing that is unspeakable or unthinkable that says this is the unthinkable that says this is the unthinkable that says this is the candidate that I am going with and candidate that I am going with and candidate that I am going with and between two candidates that are equally between two candidates that are equally between two candidates that are equally qualified. I know that this person has a qualified. I know that this person has a qualified. I know that this person has a quality of character that I've seen come quality of character that I've seen come quality of character that I've seen come through in interviews that I want on my through in interviews that I want on my through in interviews that I want on my team. AI isn't good at that kind of team. AI isn't good at that kind of team. AI isn't good at that kind of thing. that needs human judgment. And thing. that needs human judgment. And thing. that needs human judgment. And this is where I said the cheapest thing this is where I said the cheapest thing this is where I said the cheapest thing sometimes is to sit down, put the AI to sometimes is to sit down, put the AI to sometimes is to sit down, put the AI to the side, and type out your own answer the side, and type out your own answer the side, and type out your own answer to use your human judgment. So, here's to use your human judgment. So, here's to use your human judgment. So, here's where I'm going to leave you. First, I'm where I'm going to leave you. First, I'm where I'm going to leave you. First, I'm going to give you a skill, a tool that going to give you a skill, a tool that going to give you a skill, a tool that you can use to determine whether you can use to determine whether you can use to determine whether something is single agent, multi-agent, something is single agent, multi-agent, something is single agent, multi-agent, or whether it's something you should do or whether it's something you should do or whether it's something you should do yourself. And you can develop that yourself. And you can develop that yourself. And you can develop that instinct on your own. I gave you the instinct on your own. I gave you the instinct on your own. I gave you the exact questions, but if you want a skill

  19. exact questions, but if you want a skill exact questions, but if you want a skill to do that, I got that. Second, I want to do that, I got that. Second, I want to do that, I got that. Second, I want you to remember that the tools I'm you to remember that the tools I'm you to remember that the tools I'm showing you in this video, they're going showing you in this video, they're going showing you in this video, they're going to be replaced, right? The spreadsheet to be replaced, right? The spreadsheet to be replaced, right? The spreadsheet prompt will evolve. The single goal prompt will evolve. The single goal prompt will evolve. The single goal agent will evolve. Eventually, today's agent will evolve. Eventually, today's agent will evolve. Eventually, today's team setups are going to evolve, too. team setups are going to evolve, too. team setups are going to evolve, too. But what won't change is the power of But what won't change is the power of But what won't change is the power of these estimates and these questions I've these estimates and these questions I've these estimates and these questions I've given you. We will still be asking about given you. We will still be asking about given you. We will still be asking about the size of the task, the independence the size of the task, the independence the size of the task, the independence of the task, whether we have separation of the task, whether we have separation of the task, whether we have separation of concerns in the task, whether we have of concerns in the task, whether we have of concerns in the task, whether we have verifiability because these describe the verifiability because these describe the verifiability because these describe the work, not the evolving tools. And that's work, not the evolving tools. And that's work, not the evolving tools. And that's why I chose them. In a market where why I chose them. In a market where why I chose them. In a market where everything is disposable, this test is everything is disposable, this test is everything is disposable, this test is like a buy it for life purchase, right? like a buy it for life purchase, right? like a buy it for life purchase, right? So you can name your pile. You can name So you can name your pile. You can name So you can name your pile. You can name the handoff. You can name the folder the handoff. You can name the folder the handoff. You can name the folder that nobody opens. You can run your that nobody opens. You can run your that nobody opens. You can run your estimates against it and you can figure estimates against it and you can figure estimates against it and you can figure out is this an agent task? Is it a out is this an agent task? Is it a out is this an agent task? Is it a multi- aent task? And then you can get multi- aent task? And then you can get multi- aent task? And then you can get going right away. And that's what I going right away. And that's what I going right away. And that's what I built for you. I built a tool for you built for you. I built a tool for you built for you. I built a tool for you that lets you estimate all of this in a that lets you estimate all of this in a that lets you estimate all of this in a minute or less and then get started with minute or less and then get started with minute or less and then get started with multi-agent if that's what you need to multi-agent if that's what you need to multi-agent if that's what you need to do or recommends a single agent setup do or recommends a single agent setup do or recommends a single agent setup and you can go right into codeex or claw and you can go right into codeex or claw and you can go right into codeex or claw from there or your tool of choice. And from there or your tool of choice. And from there or your tool of choice. And why did I build this? Because nobody out why did I build this? Because nobody out why did I build this? Because nobody out there will tell you whether your task is there will tell you whether your task is there will tell you whether your task is agent-shaped. I built the thing that agent-shaped. I built the thing that agent-shaped. I built the thing that will. It's live right now. You just will. It's live right now. You just will. It's live right now. You just describe your task. You set the sliders.

  20. describe your task. You set the sliders. describe your task. You set the sliders. Uh so you have four estimates plus the Uh so you have four estimates plus the Uh so you have four estimates plus the two money dials, right? how often does two money dials, right? how often does two money dials, right? how often does it come back as a cost and what is a it come back as a cost and what is a it come back as a cost and what is a good answer worth to you and it's going good answer worth to you and it's going good answer worth to you and it's going to give you a verdict, right? Uh you can to give you a verdict, right? Uh you can to give you a verdict, right? Uh you can you can chat, you can use an agent, a you can chat, you can use an agent, a you can chat, you can use an agent, a team or you don't even have to bother. team or you don't even have to bother. team or you don't even have to bother. And every verdict is going to come with And every verdict is going to come with And every verdict is going to come with a next step attached because a verdict a next step attached because a verdict a next step attached because a verdict without a way forward is just more without a way forward is just more without a way forward is just more homework for you to do, right? You need homework for you to do, right? You need homework for you to do, right? You need something that helps you to actually use something that helps you to actually use something that helps you to actually use that pattern and pick until your that pattern and pick until your that pattern and pick until your instincts really develop because a instincts really develop because a instincts really develop because a verdict without a door to go forward is verdict without a door to go forward is verdict without a door to go forward is just more work for you, right? You want just more work for you, right? You want just more work for you, right? You want something that gets you into the work. something that gets you into the work. something that gets you into the work. You want something that says, "Okay, I'm You want something that says, "Okay, I'm You want something that says, "Okay, I'm going to get into Ringer and do this going to get into Ringer and do this going to get into Ringer and do this task, and it's just a click away. I want task, and it's just a click away. I want task, and it's just a click away. I want to get into chat GPT and do this task. to get into chat GPT and do this task. to get into chat GPT and do this task. It's a click away. Oh, it's human It's a click away. Oh, it's human It's a click away. Oh, it's human judgment." Uh, yeah, that's a good judgment." Uh, yeah, that's a good judgment." Uh, yeah, that's a good reminder. Which, by the way, that may be reminder. Which, by the way, that may be reminder. Which, by the way, that may be the most valuable thing about this tool the most valuable thing about this tool the most valuable thing about this tool is I find a lot of people when they're is I find a lot of people when they're is I find a lot of people when they're so excited about AI overindex on what so excited about AI overindex on what so excited about AI overindex on what they delegate to AI. And this tool, I they delegate to AI. And this tool, I they delegate to AI. And this tool, I find, is a helpful reminder of when to find, is a helpful reminder of when to find, is a helpful reminder of when to apply human judgment. Do the minute of apply human judgment. Do the minute of apply human judgment. Do the minute of thinking yourself first. Check the tool. thinking yourself first. Check the tool. thinking yourself first. Check the tool. If you agree, then you're in alignment If you agree, then you're in alignment If you agree, then you're in alignment and move with confidence. If you and move with confidence. If you and move with confidence. If you disagree, that's an interesting case.

  21. disagree, that's an interesting case. disagree, that's an interesting case. And I want you to think about that And I want you to think about that And I want you to think about that disagreement between the tool and your disagreement between the tool and your disagreement between the tool and your own instinct because there's probably own instinct because there's probably own instinct because there's probably something here for you to learn from. something here for you to learn from. something here for you to learn from. There's probably something the tool is There's probably something the tool is There's probably something the tool is seeing that you may not be recognizing seeing that you may not be recognizing seeing that you may not be recognizing from a size and complexity perspective. from a size and complexity perspective. from a size and complexity perspective. You can get the tool, you can get the You can get the tool, you can get the You can get the tool, you can get the one-click startup for Ringer, and you one-click startup for Ringer, and you one-click startup for Ringer, and you can get all of it at the link below. can get all of it at the link below. can get all of it at the link below. Have fun.

Summary

The video addresses the fundamental problem of identifying when and how to use AI agents effectively, particularly in a "postOpenClaw" era where many agents remainTasked. It references the overwhelming number of registered agents from OpenClaw as an example of underutilization due to unclear task-agent alignment. The practical takeaway is to provide academic and practical grounding, a one-minute test, and an AI skill tool to help users accurately match tasks to the right AI solutions or identify when not to use AI at all, ultimately enabling faster and more accurate task completion.

View original episode ↗