Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori
Read full transcript 15 segments
-
Um, well, thanks a lot for your time. Um, well, thanks a lot for your time. Really appreciate you dropping by and Really appreciate you dropping by and Really appreciate you dropping by and it's always a great honor to speak at it's always a great honor to speak at it's always a great honor to speak at the World's Fair. So, I'll do my best to the World's Fair. So, I'll do my best to the World's Fair. So, I'll do my best to uh give you guys some valuable insights uh give you guys some valuable insights uh give you guys some valuable insights and um yeah, hopefully make it worth and um yeah, hopefully make it worth and um yeah, hopefully make it worth your time. So my name is Maximillian your time. So my name is Maximillian your time. So my name is Maximillian Piros and today I'll be talking about Piros and today I'll be talking about Piros and today I'll be talking about mouse power and this is a talk about mouse power and this is a talk about mouse power and this is a talk about measuring agents through mental models. measuring agents through mental models. measuring agents through mental models. But before I get into talking about But before I get into talking about But before I get into talking about measuring agents, I'm going to talk measuring agents, I'm going to talk measuring agents, I'm going to talk through a bit about how I use them every through a bit about how I use them every through a bit about how I use them every day. And it might seem familiar to you, day. And it might seem familiar to you, day. And it might seem familiar to you, but just to level set, we'll go through but just to level set, we'll go through but just to level set, we'll go through it. So I tend to background them like it. So I tend to background them like it. So I tend to background them like I'm sure a lot of you people are as I'm sure a lot of you people are as I'm sure a lot of you people are as well. Um, so while my active attention well. Um, so while my active attention well. Um, so while my active attention is focusing on one thing, like perhaps is focusing on one thing, like perhaps is focusing on one thing, like perhaps giving this talk to you, I still want to giving this talk to you, I still want to giving this talk to you, I still want to make some progress on peripheral tasks. make some progress on peripheral tasks. make some progress on peripheral tasks. So I'll keep my attention focused on So I'll keep my attention focused on So I'll keep my attention focused on giving this talk while my agents can giving this talk while my agents can giving this talk while my agents can help me explore some designs in the help me explore some designs in the help me explore some designs in the background because I think that uh my background because I think that uh my background because I think that uh my slides need a bit of work. So you know, slides need a bit of work. So you know, slides need a bit of work. So you know, I've got my design system already set I've got my design system already set I've got my design system already set up. I've got some guidance uh given to up. I've got some guidance uh given to up. I've got some guidance uh given to my agents. So I'll kick off an agent to my agents. So I'll kick off an agent to my agents. So I'll kick off an agent to just try to explore some different just try to explore some different just try to explore some different directions on the type treatment and the directions on the type treatment and the directions on the type treatment and the layout and um you know just try to get layout and um you know just try to get layout and um you know just try to get as many explorations as possible. But of as many explorations as possible. But of as many explorations as possible. But of course uh one agent's never enough. So I course uh one agent's never enough. So I course uh one agent's never enough. So I like to kick off a bunch in parallel.
-
like to kick off a bunch in parallel. like to kick off a bunch in parallel. You know I've got a lot of slides to get You know I've got a lot of slides to get You know I've got a lot of slides to get through. So I need all my agents on and through. So I need all my agents on and through. So I need all my agents on and exploring it in different directions. exploring it in different directions. exploring it in different directions. And uh hopefully I can get some And uh hopefully I can get some And uh hopefully I can get some interesting things to make my slides a interesting things to make my slides a interesting things to make my slides a bit better. and uh hopefully they can bit better. and uh hopefully they can bit better. and uh hopefully they can finish the job pretty soon because we're finish the job pretty soon because we're finish the job pretty soon because we're obviously kind of up against uh timeline obviously kind of up against uh timeline obviously kind of up against uh timeline the the uh deadline here. So um this is the the uh deadline here. So um this is the the uh deadline here. So um this is generally how I work. I'm sure it's generally how I work. I'm sure it's generally how I work. I'm sure it's probably familiar to a lot of you where probably familiar to a lot of you where probably familiar to a lot of you where we're just trying to kick off agents for we're just trying to kick off agents for we're just trying to kick off agents for as much as possible in parallel because as much as possible in parallel because as much as possible in parallel because it always feels like there's just way it always feels like there's just way it always feels like there's just way more research to do. We want to be as more research to do. We want to be as more research to do. We want to be as thorough as possible. There's way more thorough as possible. There's way more thorough as possible. There's way more design explorations to do. So whenever design explorations to do. So whenever design explorations to do. So whenever our main focus is on one thing, why not our main focus is on one thing, why not our main focus is on one thing, why not kick a bunch of agents off in parallel kick a bunch of agents off in parallel kick a bunch of agents off in parallel and just try to maximize your time and and just try to maximize your time and and just try to maximize your time and it's a lot of fun of course until you it's a lot of fun of course until you it's a lot of fun of course until you get the bill and then you start to get the bill and then you start to get the bill and then you start to wonder was it all worth it, right? Did wonder was it all worth it, right? Did wonder was it all worth it, right? Did did you vibe code too hard? Were you did you vibe code too hard? Were you did you vibe code too hard? Were you token maxing too much? Like could you token maxing too much? Like could you token maxing too much? Like could you have been more efficient in how you have been more efficient in how you have been more efficient in how you approached uh your sequencing your approached uh your sequencing your approached uh your sequencing your agents? And so this is what I'm going to agents? And so this is what I'm going to agents? And so this is what I'm going to get into today. It's it's how do we get into today. It's it's how do we get into today. It's it's how do we value uh the token cost and specifically value uh the token cost and specifically value uh the token cost and specifically how do we help our customers value it.
-
how do we help our customers value it. how do we help our customers value it. So for the past year and a half I've had So for the past year and a half I've had So for the past year and a half I've had the pleasure of working as the founding the pleasure of working as the founding the pleasure of working as the founding designer at a company called UTRI and we designer at a company called UTRI and we designer at a company called UTRI and we focus on computer use models. These are focus on computer use models. These are focus on computer use models. These are models that learn to use a computer like models that learn to use a computer like models that learn to use a computer like a human would. And the use case for them a human would. And the use case for them a human would. And the use case for them is when you can't get information from is when you can't get information from is when you can't get information from an API or an MCP, why not just send an an API or an MCP, why not just send an an API or an MCP, why not just send an agent out to use a computer like a human agent out to use a computer like a human agent out to use a computer like a human would? And then we can extract all types would? And then we can extract all types would? And then we can extract all types of data and manipulate it in ways that of data and manipulate it in ways that of data and manipulate it in ways that uh let us access all the stuff that uh let us access all the stuff that uh let us access all the stuff that wasn't uh accessible previously. So wasn't uh accessible previously. So wasn't uh accessible previously. So obviously less efficient than APIs and obviously less efficient than APIs and obviously less efficient than APIs and MCPs, but as a last resort, just have an MCPs, but as a last resort, just have an MCPs, but as a last resort, just have an agent go use the computer and try to get agent go use the computer and try to get agent go use the computer and try to get the information. Uh here's the uttory the information. Uh here's the uttory the information. Uh here's the uttory agent using the utori website. Um it's agent using the utori website. Um it's agent using the utori website. Um it's checking out its own benchmark. So kind checking out its own benchmark. So kind checking out its own benchmark. So kind of it's admiring itself in a way. So um of it's admiring itself in a way. So um of it's admiring itself in a way. So um yeah, it gets it gets a bit weird. Uh yeah, it gets it gets a bit weird. Uh yeah, it gets it gets a bit weird. Uh like and a lot of what I do as a like and a lot of what I do as a like and a lot of what I do as a founding designer there is talk to founding designer there is talk to founding designer there is talk to customers, try to understand how can we customers, try to understand how can we customers, try to understand how can we make agents as intuitive as possible. make agents as intuitive as possible. make agents as intuitive as possible. How do we figure out the mental models How do we figure out the mental models How do we figure out the mental models they're using to value uh the use cases they're using to value uh the use cases they're using to value uh the use cases they want to send out agents for? And um they want to send out agents for? And um they want to send out agents for? And um a lot of them do seem pretty confused so a lot of them do seem pretty confused so a lot of them do seem pretty confused so far. A lot of people are excited about far. A lot of people are excited about far. A lot of people are excited about agents, but the a phrase that comes up agents, but the a phrase that comes up agents, but the a phrase that comes up quite often is that they feel like quite often is that they feel like quite often is that they feel like they're just scratching the surface. Um, they're just scratching the surface. Um, they're just scratching the surface. Um, it seems like it's not quite intuitive it seems like it's not quite intuitive it seems like it's not quite intuitive how we can best use them yet. And so in how we can best use them yet. And so in how we can best use them yet. And so in a lot of my uh customer discussions, a lot of my uh customer discussions, a lot of my uh customer discussions, it's it always comes down to a question it's it always comes down to a question it's it always comes down to a question of like what is the best way to use you of like what is the best way to use you of like what is the best way to use you uh to use agents? What are the best use uh to use agents? What are the best use uh to use agents? What are the best use cases for them? And how do I think about cases for them? And how do I think about cases for them? And how do I think about the trade-offs with regards to token the trade-offs with regards to token the trade-offs with regards to token costs relative to value? So, I think costs relative to value? So, I think costs relative to value? So, I think we're still kind of building this muscle
-
we're still kind of building this muscle we're still kind of building this muscle today. And this leads me to the thesis today. And this leads me to the thesis today. And this leads me to the thesis of the talk, which is that I think of the talk, which is that I think of the talk, which is that I think agents have a measurement problem. And agents have a measurement problem. And agents have a measurement problem. And uh, as an example, here's me at work uh, as an example, here's me at work uh, as an example, here's me at work trying to measure some agents, and one trying to measure some agents, and one trying to measure some agents, and one of my co-workers took this photo and of my co-workers took this photo and of my co-workers took this photo and told me I looked like I was trying to told me I looked like I was trying to told me I looked like I was trying to solve the mystery of Pepe Silva. So, uh, solve the mystery of Pepe Silva. So, uh, solve the mystery of Pepe Silva. So, uh, as you can see, it's it's it's not a as you can see, it's it's it's not a as you can see, it's it's it's not a it's not an easy task to to measure it's not an easy task to to measure it's not an easy task to to measure agents. But I'm sure some of you are agents. But I'm sure some of you are agents. But I'm sure some of you are saying, "Hold on a sec. like what is saying, "Hold on a sec. like what is saying, "Hold on a sec. like what is this guy talking about? I've got a fleet this guy talking about? I've got a fleet this guy talking about? I've got a fleet of agents working for me right now. of agents working for me right now. of agents working for me right now. We're building our next million-dollar We're building our next million-dollar We're building our next million-dollar app as we speak and I'm having a totally app as we speak and I'm having a totally app as we speak and I'm having a totally fine time uh measuring my agents. Uh to fine time uh measuring my agents. Uh to fine time uh measuring my agents. Uh to which I will agree with you. Uh, but which I will agree with you. Uh, but which I will agree with you. Uh, but then I will point you to the mandatory then I will point you to the mandatory then I will point you to the mandatory Upton Sinclair quote to remind us all Upton Sinclair quote to remind us all Upton Sinclair quote to remind us all that everybody in this room is very that everybody in this room is very that everybody in this room is very biased and we're early adopters and biased and we're early adopters and biased and we're early adopters and we're very excited to explore this new we're very excited to explore this new we're very excited to explore this new technology, but it doesn't mean that we technology, but it doesn't mean that we technology, but it doesn't mean that we represent the people that ultimately represent the people that ultimately represent the people that ultimately we're going to be trying to help adopt we're going to be trying to help adopt we're going to be trying to help adopt this technology. And so, you know, I this technology. And so, you know, I this technology. And so, you know, I think it's important to remind ourselves think it's important to remind ourselves think it's important to remind ourselves that in some way or another, we probably that in some way or another, we probably that in some way or another, we probably are selling tokens, whether it's are selling tokens, whether it's are selling tokens, whether it's indirectly or directly. And so when we indirectly or directly. And so when we indirectly or directly. And so when we think about our own token usage, uh is think about our own token usage, uh is think about our own token usage, uh is it really representative of all the it really representative of all the it really representative of all the people out there who have never touched people out there who have never touched people out there who have never touched an agent yet? Uh some people are still an agent yet? Uh some people are still an agent yet? Uh some people are still copy and pasting into chatbt. I may be copy and pasting into chatbt. I may be copy and pasting into chatbt. I may be married to one of these people and married to one of these people and married to one of these people and despite how much I tried to get her to despite how much I tried to get her to despite how much I tried to get her to try out agents, she's not let me set her try out agents, she's not let me set her try out agents, she's not let me set her up with it yet. And so, um, as a up with it yet. And so, um, as a up with it yet. And so, um, as a reminder, uh, when we think about reminder, uh, when we think about reminder, uh, when we think about helping people adopt agents, you know, helping people adopt agents, you know, helping people adopt agents, you know, all the people across the world that we all the people across the world that we all the people across the world that we think could get as much excitement and think could get as much excitement and think could get as much excitement and value as as we do when we run off
-
value as as we do when we run off value as as we do when we run off parallel agents, let's just remember parallel agents, let's just remember parallel agents, let's just remember this quote. [snorts] this quote. [snorts] this quote. [snorts] And um, so it really boils down to the And um, so it really boils down to the And um, so it really boils down to the age-old problem of a new technology. And age-old problem of a new technology. And age-old problem of a new technology. And of course, there's tons of history we of course, there's tons of history we of course, there's tons of history we can go to to study how people solve this can go to to study how people solve this can go to to study how people solve this in the past. We have this really in the past. We have this really in the past. We have this really exciting new thing, but we haven't quite exciting new thing, but we haven't quite exciting new thing, but we haven't quite figured out the right ways to figured out the right ways to figured out the right ways to communicate it. And so for this talk, communicate it. And so for this talk, communicate it. And so for this talk, I'll go back to the 1700s and we can I'll go back to the 1700s and we can I'll go back to the 1700s and we can take some notes from when James Watt was take some notes from when James Watt was take some notes from when James Watt was trying to sell steam engines. And at the trying to sell steam engines. And at the trying to sell steam engines. And at the time he decided that a great use case time he decided that a great use case time he decided that a great use case for his steam engines was trying to for his steam engines was trying to for his steam engines was trying to replace a horse gin. And these are the replace a horse gin. And these are the replace a horse gin. And these are the was the power source of a mill at the was the power source of a mill at the was the power source of a mill at the time. So when you're for uh let's say a time. So when you're for uh let's say a time. So when you're for uh let's say a brewery and you need some power source brewery and you need some power source brewery and you need some power source to to grind your barley or whatever, I to to grind your barley or whatever, I to to grind your barley or whatever, I don't know. I'm not like a big brewery don't know. I'm not like a big brewery don't know. I'm not like a big brewery guy, so I don't know exactly how it's guy, so I don't know exactly how it's guy, so I don't know exactly how it's made, but you need a power source. And made, but you need a power source. And made, but you need a power source. And the power source at the time that was the power source at the time that was the power source at the time that was common was you hooked a horse up to a common was you hooked a horse up to a common was you hooked a horse up to a rotary arm and the horse walked in a rotary arm and the horse walked in a rotary arm and the horse walked in a circle and that's how you generated your circle and that's how you generated your circle and that's how you generated your power. Uh, and seems crazy today maybe, power. Uh, and seems crazy today maybe, power. Uh, and seems crazy today maybe, but um at the time was common place. And but um at the time was common place. And but um at the time was common place. And Watt thought, you know, it would be much Watt thought, you know, it would be much Watt thought, you know, it would be much better than a horse is like a very better than a horse is like a very better than a horse is like a very efficient machine. Although he u efficient machine. Although he u efficient machine. Although he u rightfully acknowledged that one of the rightfully acknowledged that one of the rightfully acknowledged that one of the big barriers to adopting it would be big barriers to adopting it would be big barriers to adopting it would be this cognitive dissonance of trying to this cognitive dissonance of trying to this cognitive dissonance of trying to tell people who kind of think in horses tell people who kind of think in horses tell people who kind of think in horses how do you adapt to this to this uh cold how do you adapt to this to this uh cold how do you adapt to this to this uh cold machine that's kind of intimidating and machine that's kind of intimidating and machine that's kind of intimidating and scary and perhaps uh somebody's going to scary and perhaps uh somebody's going to scary and perhaps uh somebody's going to say it's going to solve all your say it's going to solve all your say it's going to solve all your problems but you you can't quite see the problems but you you can't quite see the problems but you you can't quite see the vision yet. So perhaps that sounds vision yet. So perhaps that sounds vision yet. So perhaps that sounds familiar to any of us working in agents familiar to any of us working in agents familiar to any of us working in agents today. And Watt's solution was that he
-
today. And Watt's solution was that he today. And Watt's solution was that he needed to understand um the mental model needed to understand um the mental model needed to understand um the mental model of these people and specifically to of these people and specifically to of these people and specifically to create a metric that would help him uh create a metric that would help him uh create a metric that would help him uh give some baseline of the relative give some baseline of the relative give some baseline of the relative improvement in efficiency. And so he improvement in efficiency. And so he improvement in efficiency. And so he literally studied uh horse jins and literally studied uh horse jins and literally studied uh horse jins and tried to get some kind of armchair tried to get some kind of armchair tried to get some kind of armchair measurements of of how is like what are measurements of of how is like what are measurements of of how is like what are the mechanics and the average um the mechanics and the average um the mechanics and the average um performance of it and eventually came to performance of it and eventually came to performance of it and eventually came to a metric called horsepower which may a metric called horsepower which may a metric called horsepower which may sound familiar. And uh he used this sound familiar. And uh he used this sound familiar. And uh he used this measure to you know this was to quantify measure to you know this was to quantify measure to you know this was to quantify the the general power that the horses the the general power that the horses the the general power that the horses were were um creating at the time and were were um creating at the time and were were um creating at the time and then he could use as a basis to show the then he could use as a basis to show the then he could use as a basis to show the multiplier of efficiency that a steam multiplier of efficiency that a steam multiplier of efficiency that a steam engine could provide. And uh this metric engine could provide. And uh this metric engine could provide. And uh this metric was uh not very scientific at the time. was uh not very scientific at the time. was uh not very scientific at the time. It was not necessarily even accurate you It was not necessarily even accurate you It was not necessarily even accurate you could say. Uh but the main thing it did could say. Uh but the main thing it did could say. Uh but the main thing it did was it communicated an increase in was it communicated an increase in was it communicated an increase in value. And so this let people who who value. And so this let people who who value. And so this let people who who love horses uh let them kind of love horses uh let them kind of love horses uh let them kind of calibrate their um the the efficiency calibrate their um the the efficiency calibrate their um the the efficiency gains that they could get by by uh gains that they could get by by uh gains that they could get by by uh attempting to adopt a steam engine. So attempting to adopt a steam engine. So attempting to adopt a steam engine. So not even necessarily um what you would not even necessarily um what you would not even necessarily um what you would get when you use it, but what would get get when you use it, but what would get get when you use it, but what would get you over the limit of trying it out in you over the limit of trying it out in you over the limit of trying it out in the first place.
-
the first place. the first place. And uh you know it's it's it's a pretty And uh you know it's it's it's a pretty And uh you know it's it's it's a pretty big feat because like although um he had big feat because like although um he had big feat because like although um he had efficiency on his side with regards to efficiency on his side with regards to efficiency on his side with regards to this metric um you know let's be honest this metric um you know let's be honest this metric um you know let's be honest regardless of how efficient this was um regardless of how efficient this was um regardless of how efficient this was um horses just have great vibes. So, like horses just have great vibes. So, like horses just have great vibes. So, like it's kind of hard to beat the vibes of it's kind of hard to beat the vibes of it's kind of hard to beat the vibes of horses. And so, he knew he had to kind horses. And so, he knew he had to kind horses. And so, he knew he had to kind of overcome the emotion uh and actually of overcome the emotion uh and actually of overcome the emotion uh and actually speak to to something that gave him an speak to to something that gave him an speak to to something that gave him an ability to calculate the ROI. ability to calculate the ROI. ability to calculate the ROI. And oh, sorry, skipped something. And And oh, sorry, skipped something. And And oh, sorry, skipped something. And so, yeah, the le the lesson being um if so, yeah, the le the lesson being um if so, yeah, the le the lesson being um if we're not able to give something that is we're not able to give something that is we're not able to give something that is a tangible ROI um for our customers, a tangible ROI um for our customers, a tangible ROI um for our customers, then it's very hard for us to comm then it's very hard for us to comm then it's very hard for us to comm communicate value. And um I think we communicate value. And um I think we communicate value. And um I think we only need to look to our own industry to only need to look to our own industry to only need to look to our own industry to see all the examples where other people see all the examples where other people see all the examples where other people in the technology sector are failing to in the technology sector are failing to in the technology sector are failing to calculate good ROIs as well. And so we calculate good ROIs as well. And so we calculate good ROIs as well. And so we might in this room think this is might in this room think this is might in this room think this is somewhat of a solved problem. Uh but if somewhat of a solved problem. Uh but if somewhat of a solved problem. Uh but if you look to the other engineers in the you look to the other engineers in the you look to the other engineers in the world who are perhaps not as AI pilled, world who are perhaps not as AI pilled, world who are perhaps not as AI pilled, uh they're theoretically very smart and uh they're theoretically very smart and uh they're theoretically very smart and should be able to figure out how to should be able to figure out how to should be able to figure out how to calculate this pretty well. But then you calculate this pretty well. But then you calculate this pretty well. But then you get these scenarios where people are get these scenarios where people are get these scenarios where people are blowing through their entire um token blowing through their entire um token blowing through their entire um token budget for a year and they're blowing budget for a year and they're blowing budget for a year and they're blowing through it in in a quarter or they're through it in in a quarter or they're through it in in a quarter or they're like dealing with token leaderboards and like dealing with token leaderboards and like dealing with token leaderboards and such. And so obviously the incentives such. And so obviously the incentives such. And so obviously the incentives haven't quite aligned and we haven't haven't quite aligned and we haven't haven't quite aligned and we haven't perhaps got the right measure of value perhaps got the right measure of value perhaps got the right measure of value in terms of the technology sector in terms of the technology sector in terms of the technology sector itself. And so how then do we end up itself. And so how then do we end up itself. And so how then do we end up scaling past past that and talk to scaling past past that and talk to scaling past past that and talk to people who have no idea what we're people who have no idea what we're people who have no idea what we're talking about but still try to provide talking about but still try to provide talking about but still try to provide them um a measure of like increased them um a measure of like increased them um a measure of like increased efficiency with agents.
-
efficiency with agents. efficiency with agents. And so right now I think we're kind of And so right now I think we're kind of And so right now I think we're kind of in this doom loop where we're we're in this doom loop where we're we're in this doom loop where we're we're overspending and we're underusing. Uh overspending and we're underusing. Uh overspending and we're underusing. Uh this is a term I borrowed from ramp um this is a term I borrowed from ramp um this is a term I borrowed from ramp um and they have a great blog post on this. and they have a great blog post on this. and they have a great blog post on this. And so it's kind of this vicious cycle And so it's kind of this vicious cycle And so it's kind of this vicious cycle where we're just token maxing and where we're just token maxing and where we're just token maxing and ourselves into austerity and then kind ourselves into austerity and then kind ourselves into austerity and then kind of dropping out of the loop until we get of dropping out of the loop until we get of dropping out of the loop until we get more FOMO to to get activated enough to more FOMO to to get activated enough to more FOMO to to get activated enough to try it again. And so I think we have to try it again. And so I think we have to try it again. And so I think we have to break this loop and I think the way we break this loop and I think the way we break this loop and I think the way we do that is by getting better measures of do that is by getting better measures of do that is by getting better measures of that will communicate value. that will communicate value. that will communicate value. Some people are obviously on the right Some people are obviously on the right Some people are obviously on the right track. There was this chart floating track. There was this chart floating track. There was this chart floating around on X recently that the Coinbase around on X recently that the Coinbase around on X recently that the Coinbase Coinbase CEO posted where they had Coinbase CEO posted where they had Coinbase CEO posted where they had internally started changing the defaults internally started changing the defaults internally started changing the defaults of what models they will start with and of what models they will start with and of what models they will start with and trying to only save the frontier models trying to only save the frontier models trying to only save the frontier models for the hardest tasks and as a result for the hardest tasks and as a result for the hardest tasks and as a result saw some good um saw AI spend start to saw some good um saw AI spend start to saw some good um saw AI spend start to diverge from token usage and uh this is diverge from token usage and uh this is diverge from token usage and uh this is a good start. Uh RAM also as I mentioned a good start. Uh RAM also as I mentioned a good start. Uh RAM also as I mentioned has a great blog post about this. uh but has a great blog post about this. uh but has a great blog post about this. uh but I think the problem is still that it's I think the problem is still that it's I think the problem is still that it's too focused on tokens and tokens are of too focused on tokens and tokens are of too focused on tokens and tokens are of course uh useful as a measurement of an course uh useful as a measurement of an course uh useful as a measurement of an internal system but at at the end of the internal system but at at the end of the internal system but at at the end of the day they're just an output and so the day they're just an output and so the day they're just an output and so the tokens need to then be traced very tokens need to then be traced very tokens need to then be traced very cleanly to an outcome. So, how many uh cleanly to an outcome. So, how many uh cleanly to an outcome. So, how many uh bug uh how many bugs did the tokens uh bug uh how many bugs did the tokens uh bug uh how many bugs did the tokens uh how sorry, how many um bugs squashed did how sorry, how many um bugs squashed did how sorry, how many um bugs squashed did the tokens that we bought? Um sorry, the tokens that we bought? Um sorry, the tokens that we bought? Um sorry, totally butchered that. Um how many uh totally butchered that. Um how many uh totally butchered that. Um how many uh bugs got squashed with the to with our bugs got squashed with the to with our bugs got squashed with the to with our token spend? How many uh support token spend? How many uh support token spend? How many uh support requests got closed, etc. So, clean requests got closed, etc. So, clean requests got closed, etc. So, clean outcomes and then cleanly tying those to
-
outcomes and then cleanly tying those to outcomes and then cleanly tying those to to progress on our objectives. And so, to progress on our objectives. And so, to progress on our objectives. And so, without a very tight measure of ROI, without a very tight measure of ROI, without a very tight measure of ROI, this becomes very hard to do. this becomes very hard to do. this becomes very hard to do. And I think I'll take this further and And I think I'll take this further and And I think I'll take this further and um say that it need not even be the um say that it need not even be the um say that it need not even be the broader um technology industry we're broader um technology industry we're broader um technology industry we're that's encountering this problem, but that's encountering this problem, but that's encountering this problem, but also many of us in this room perhaps also many of us in this room perhaps also many of us in this room perhaps are. And uh although we're all probably are. And uh although we're all probably are. And uh although we're all probably enjoying uh coding with uh with various enjoying uh coding with uh with various enjoying uh coding with uh with various agents and feeling like it it's it does agents and feeling like it it's it does agents and feeling like it it's it does feel like there's something there in feel like there's something there in feel like there's something there in terms of the increase in ability and terms of the increase in ability and terms of the increase in ability and efficiency. Um the problem of course is efficiency. Um the problem of course is efficiency. Um the problem of course is that we're all kind of dying by thousand that we're all kind of dying by thousand that we're all kind of dying by thousand poll requests. And so uh even Enthropic poll requests. And so uh even Enthropic poll requests. And so uh even Enthropic who has uh some people on the team have who has uh some people on the team have who has uh some people on the team have claimed to have solved coding uh they claimed to have solved coding uh they claimed to have solved coding uh they have also admitted that they've not have also admitted that they've not have also admitted that they've not solved code review. And so as a result solved code review. And so as a result solved code review. And so as a result um the the bottleneck is now shifted to um the the bottleneck is now shifted to um the the bottleneck is now shifted to human review where the efficiency gains human review where the efficiency gains human review where the efficiency gains from coding agents aren't quite or from coding agents aren't quite or from coding agents aren't quite or aren't seen yet because we spend most of aren't seen yet because we spend most of aren't seen yet because we spend most of the time reviewing the code and we've the time reviewing the code and we've the time reviewing the code and we've not figured out how to scale that in not figured out how to scale that in not figured out how to scale that in tandem with uh the generation of the tandem with uh the generation of the tandem with uh the generation of the code itself. And so the bottleneck ends code itself. And so the bottleneck ends code itself. And so the bottleneck ends up shifting to the verification side and up shifting to the verification side and up shifting to the verification side and thus we don't have a way to measure thus we don't have a way to measure thus we don't have a way to measure value at scale and uh to judge quality value at scale and uh to judge quality value at scale and uh to judge quality at at the same speed. And so again going at at the same speed. And so again going at at the same speed. And so again going back to the ROI calculations, we back to the ROI calculations, we back to the ROI calculations, we generate all this code but how do we generate all this code but how do we generate all this code but how do we know uh we don't know that enough of it know uh we don't know that enough of it know uh we don't know that enough of it is good to justify the spend.
-
is good to justify the spend. is good to justify the spend. And of course maybe uh code review was And of course maybe uh code review was And of course maybe uh code review was always flawed. Uh but it's just that always flawed. Uh but it's just that always flawed. Uh but it's just that agents are now exposing it for uh the or agents are now exposing it for uh the or agents are now exposing it for uh the or exposing the actual problem. Um, but I exposing the actual problem. Um, but I exposing the actual problem. Um, but I like this quote by Noah Hine who from a like this quote by Noah Hine who from a like this quote by Noah Hine who from a a post about how to solve code review a post about how to solve code review a post about how to solve code review where he's mentioning specifically that where he's mentioning specifically that where he's mentioning specifically that the assumptions underneath code review the assumptions underneath code review the assumptions underneath code review are what's now being what needs to be are what's now being what needs to be are what's now being what needs to be revisited. So we have to check our revisited. So we have to check our revisited. So we have to check our priors to try to figure out a new basis priors to try to figure out a new basis priors to try to figure out a new basis for um how we can code review in the age for um how we can code review in the age for um how we can code review in the age of agents. And I'm not going to go into of agents. And I'm not going to go into of agents. And I'm not going to go into how to solve code review. I think that's how to solve code review. I think that's how to solve code review. I think that's definitely better uh a talk that's definitely better uh a talk that's definitely better uh a talk that's better given by somebody else and um is better given by somebody else and um is better given by somebody else and um is totally different subject. But what I totally different subject. But what I totally different subject. But what I think is important for this talk is why think is important for this talk is why think is important for this talk is why does code review feel like it is does code review feel like it is does code review feel like it is solvable. And I think that Noah is solvable. And I think that Noah is solvable. And I think that Noah is hitting on something important here hitting on something important here hitting on something important here which is that as a as a culture uh code which is that as a as a culture uh code which is that as a as a culture uh code review has a very good uh convergence on review has a very good uh convergence on review has a very good uh convergence on shared assumptions and that lets you um shared assumptions and that lets you um shared assumptions and that lets you um that lets you uh measure things at scale that lets you uh measure things at scale that lets you uh measure things at scale when we can all kind of converge on the when we can all kind of converge on the when we can all kind of converge on the measurement and it becomes somewhat of measurement and it becomes somewhat of measurement and it becomes somewhat of of a clear rubric. And so the task at of a clear rubric. And so the task at of a clear rubric. And so the task at hand now is we have to adopt uh we have hand now is we have to adopt uh we have hand now is we have to adopt uh we have to sorry adapt those assumptions for the to sorry adapt those assumptions for the to sorry adapt those assumptions for the agentic age.
-
And so uh we can we need to go um if if And so uh we can we need to go um if if we're able to do that then we can go we're able to do that then we can go we're able to do that then we can go from execution at the speed of compute from execution at the speed of compute from execution at the speed of compute to measurement at the speed of compute. to measurement at the speed of compute. to measurement at the speed of compute. And of course the measurements need to And of course the measurements need to And of course the measurements need to fit the mental models of the customers fit the mental models of the customers fit the mental models of the customers using it. And um I think the lesson here using it. And um I think the lesson here using it. And um I think the lesson here being that if you're going to um think being that if you're going to um think being that if you're going to um think of how to build an agent for something, of how to build an agent for something, of how to build an agent for something, you also have to think about how do you you also have to think about how do you you also have to think about how do you help the customers build or build or at help the customers build or build or at help the customers build or build or at least create a method for verifying that least create a method for verifying that least create a method for verifying that the output is good. And so it's not the output is good. And so it's not the output is good. And so it's not enough to build it. We also have to help enough to build it. We also have to help enough to build it. We also have to help them we al we also have to help them get them we al we also have to help them get them we al we also have to help them get to clear ROI calculations to justify to clear ROI calculations to justify to clear ROI calculations to justify their spend. And so um this brings me to their spend. And so um this brings me to their spend. And so um this brings me to the idea of mouse power which could be the idea of mouse power which could be the idea of mouse power which could be the equivalent of horsepower for the the equivalent of horsepower for the the equivalent of horsepower for the agentic age. Just as James Walt was able agentic age. Just as James Walt was able agentic age. Just as James Walt was able to show a measure of efficiency relative to show a measure of efficiency relative to show a measure of efficiency relative to the horses in in the gins uh in the to the horses in in the gins uh in the to the horses in in the gins uh in the horse gins uh that were the source of horse gins uh that were the source of horse gins uh that were the source of power at the time. We perhaps can also power at the time. We perhaps can also power at the time. We perhaps can also figure out how do we create a baseline figure out how do we create a baseline figure out how do we create a baseline of efficiency for the way we use of efficiency for the way we use of efficiency for the way we use computers today and can then demonstrate computers today and can then demonstrate computers today and can then demonstrate how much uh better or perhaps more how much uh better or perhaps more how much uh better or perhaps more performant on certain vectors an agent performant on certain vectors an agent performant on certain vectors an agent could be at that task. Um, and of course could be at that task. Um, and of course could be at that task. Um, and of course it's not uh as easy perhaps as easy a it's not uh as easy perhaps as easy a it's not uh as easy perhaps as easy a task as he had back then where he could task as he had back then where he could task as he had back then where he could just study the horsegen because it's not just study the horsegen because it's not just study the horsegen because it's not as if we can create some method to as if we can create some method to as if we can create some method to measure our cursor movements and like measure our cursor movements and like measure our cursor movements and like figure out the delta of how much more figure out the delta of how much more figure out the delta of how much more efficient an agent could move them and efficient an agent could move them and efficient an agent could move them and thus we can say yeah agents are this thus we can say yeah agents are this thus we can say yeah agents are this much more performant than humans at much more performant than humans at much more performant than humans at these tasks. Uh, trust me, I've I've these tasks. Uh, trust me, I've I've these tasks. Uh, trust me, I've I've tried I had Claude vibe code me this tried I had Claude vibe code me this tried I had Claude vibe code me this measurement device and I thought maybe measurement device and I thought maybe measurement device and I thought maybe if I can figure out the movement uh like
-
if I can figure out the movement uh like if I can figure out the movement uh like the potential movement across the screen the potential movement across the screen the potential movement across the screen and measured how fast it went, I could and measured how fast it went, I could and measured how fast it went, I could get some clean measure of mouse power. get some clean measure of mouse power. get some clean measure of mouse power. Uh, but of course this is only joking. Uh, but of course this is only joking. Uh, but of course this is only joking. Um this is of course um like a fool's Um this is of course um like a fool's Um this is of course um like a fool's errand because information space is just errand because information space is just errand because information space is just way too high dimensional and so I think way too high dimensional and so I think way too high dimensional and so I think mouse power is of is never going to be a mouse power is of is never going to be a mouse power is of is never going to be a metric of course but it's more so an metric of course but it's more so an metric of course but it's more so an idea which the idea being if you're idea which the idea being if you're idea which the idea being if you're going to sell somebody an agent you also going to sell somebody an agent you also going to sell somebody an agent you also have to help them with the with the have to help them with the with the have to help them with the with the rubric of how do we actually verify that rubric of how do we actually verify that rubric of how do we actually verify that this agent is doing good work and thus this agent is doing good work and thus this agent is doing good work and thus we can uh have a good measure of saying we can uh have a good measure of saying we can uh have a good measure of saying that these tokens are worth it. Um, so that these tokens are worth it. Um, so that these tokens are worth it. Um, so how to do that of course is is really up how to do that of course is is really up how to do that of course is is really up to you and I won't be able to tell you to you and I won't be able to tell you to you and I won't be able to tell you how do you I don't have any good how do you I don't have any good how do you I don't have any good frameworks for how do you figure out the frameworks for how do you figure out the frameworks for how do you figure out the right measurements to to help provide right measurements to to help provide right measurements to to help provide anybody you're building an agent for. Uh anybody you're building an agent for. Uh anybody you're building an agent for. Uh but what I can do is give a princip uh but what I can do is give a princip uh but what I can do is give a princip uh give an idea that I've been kicking give an idea that I've been kicking give an idea that I've been kicking around which is based um in information around which is based um in information around which is based um in information theory. So going back to Claude theory. So going back to Claude theory. So going back to Claude Shannon's ideas about measuring entropy Shannon's ideas about measuring entropy Shannon's ideas about measuring entropy in information uh entropy being the in information uh entropy being the in information uh entropy being the uncertainty of a probability uncertainty of a probability uncertainty of a probability distribution and of course very much the distribution and of course very much the distribution and of course very much the basis of how we train agents today basis of how we train agents today basis of how we train agents today things like cross entropy and such uh things like cross entropy and such uh things like cross entropy and such uh being a big factor in determining how being a big factor in determining how being a big factor in determining how capable an agent is. Um I think that capable an agent is. Um I think that capable an agent is. Um I think that entropy is an interesting idea to think entropy is an interesting idea to think entropy is an interesting idea to think through with regards to not just the through with regards to not just the through with regards to not just the performance of an agent but also the performance of an agent but also the performance of an agent but also the task that we're setting them out to to task that we're setting them out to to task that we're setting them out to to perform on. And so uh I put together perform on. And so uh I put together perform on. And so uh I put together this matrix which it maps on the x-ax this matrix which it maps on the x-ax this matrix which it maps on the x-ax axis the uncertainty in the steps it axis the uncertainty in the steps it axis the uncertainty in the steps it takes to perform a task. And so when takes to perform a task. And so when takes to perform a task. And so when we're thinking of building an agent I
-
we're thinking of building an agent I we're thinking of building an agent I think it's not enough to just think what think it's not enough to just think what think it's not enough to just think what would be a valuable task for the agent would be a valuable task for the agent would be a valuable task for the agent to do but also thinking about how um how to do but also thinking about how um how to do but also thinking about how um how much uncertainty are in the steps to much uncertainty are in the steps to much uncertainty are in the steps to perform that task itself. So an example perform that task itself. So an example perform that task itself. So an example would be uh booking a flight has much would be uh booking a flight has much would be uh booking a flight has much less uncertainty than let's say painting less uncertainty than let's say painting less uncertainty than let's say painting a masterpiece, right? Because you know a masterpiece, right? Because you know a masterpiece, right? Because you know there's certain information that has to there's certain information that has to there's certain information that has to be that has to happen in the flight be that has to happen in the flight be that has to happen in the flight purchase. There has to be a departing purchase. There has to be a departing purchase. There has to be a departing destination, arriving destination. destination, arriving destination. destination, arriving destination. There's going to be a seat chosen. It There's going to be a seat chosen. It There's going to be a seat chosen. It might be by the person. It might just be might be by the person. It might just be might be by the person. It might just be random. Uh but these things have to random. Uh but these things have to random. Uh but these things have to happen for that task to be completed. happen for that task to be completed. happen for that task to be completed. And on the other hand, there is the task And on the other hand, there is the task And on the other hand, there is the task of like painting a masterpiece, right? of like painting a masterpiece, right? of like painting a masterpiece, right? And who knows what the steps are to And who knows what the steps are to And who knows what the steps are to that. uh and maybe you can get an agent that. uh and maybe you can get an agent that. uh and maybe you can get an agent to do it but it would be very hard to to do it but it would be very hard to to do it but it would be very hard to figure out how we can actually create a figure out how we can actually create a figure out how we can actually create a a reg relatively um predictable pathway a reg relatively um predictable pathway a reg relatively um predictable pathway to that. Uh but then on the other access to that. Uh but then on the other access to that. Uh but then on the other access uh is the the uncertainty in the uh is the the uncertainty in the uh is the the uncertainty in the acceptance criteria itself. So not just acceptance criteria itself. So not just acceptance criteria itself. So not just can the agent perform the task but can can the agent perform the task but can can the agent perform the task but can we help somebody actually or is is there we help somebody actually or is is there we help somebody actually or is is there actually a clean rubric for how it's actually a clean rubric for how it's actually a clean rubric for how it's graded. And so thinking about ideas on graded. And so thinking about ideas on graded. And so thinking about ideas on on these two axes and where they on these two axes and where they on these two axes and where they intersect perhaps gives us a better intersect perhaps gives us a better intersect perhaps gives us a better guide for how to build agents and we can guide for how to build agents and we can guide for how to build agents and we can run through a few examples. Um so if we run through a few examples. Um so if we run through a few examples. Um so if we look at the at the left side your right look at the at the left side your right look at the at the left side your right side um yes um no your left as well. Um side um yes um no your left as well. Um side um yes um no your left as well. Um then I know the last speaker was also uh then I know the last speaker was also uh then I know the last speaker was also uh confused by um so uh you know on the confused by um so uh you know on the confused by um so uh you know on the left side uh when uncertainty in the left side uh when uncertainty in the left side uh when uncertainty in the task steps are low then it's a very it's task steps are low then it's a very it's task steps are low then it's a very it's a very predictable outcome or it's a
-
a very predictable outcome or it's a a very predictable outcome or it's a very predictable pathway to achieve that very predictable pathway to achieve that very predictable pathway to achieve that goal and so then you know why would you goal and so then you know why would you goal and so then you know why would you waste tokens just write a script on the waste tokens just write a script on the waste tokens just write a script on the other side when uh the t the steps to do other side when uh the t the steps to do other side when uh the t the steps to do perform the task are very high in in perform the task are very high in in perform the task are very high in in uncertainty then you you have very uncertainty then you you have very uncertainty then you you have very unpredictable information And so it's unpredictable information And so it's unpredictable information And so it's probably at risk of being out of probably at risk of being out of probably at risk of being out of distribution in pre-training and distribution in pre-training and distribution in pre-training and probably has very sparse rewards for probably has very sparse rewards for probably has very sparse rewards for reinforcement learning. And so perhaps reinforcement learning. And so perhaps reinforcement learning. And so perhaps it's not a good task for an agent it's not a good task for an agent it's not a good task for an agent because it's just much harder to figure because it's just much harder to figure because it's just much harder to figure out how to actually model that data. And out how to actually model that data. And out how to actually model that data. And uh so obviously in the middle is is um uh so obviously in the middle is is um uh so obviously in the middle is is um is so I think I'm out of time, but I'm is so I think I'm out of time, but I'm is so I think I'm out of time, but I'm not getting kicked off yet. U so I'll not getting kicked off yet. U so I'll not getting kicked off yet. U so I'll just finish this up quickly. Um so yeah just finish this up quickly. Um so yeah just finish this up quickly. Um so yeah in the middle is is probably the sweet in the middle is is probably the sweet in the middle is is probably the sweet spot but then on the other axis uh spot but then on the other axis uh spot but then on the other axis uh what's the uncertainty in verifying that what's the uncertainty in verifying that what's the uncertainty in verifying that this actually valuable. So when you have this actually valuable. So when you have this actually valuable. So when you have high uncertainty in the acceptance high uncertainty in the acceptance high uncertainty in the acceptance criteria you pretty much at a spot where criteria you pretty much at a spot where criteria you pretty much at a spot where verification is indistinguishable from verification is indistinguishable from verification is indistinguishable from execution. So why would you build an execution. So why would you build an execution. So why would you build an agent for something that to verify it agent for something that to verify it agent for something that to verify it was useful a person pretty much has to was useful a person pretty much has to was useful a person pretty much has to do the work again. So like waste of do the work again. So like waste of do the work again. So like waste of tokens obviously and then it leaves that tokens obviously and then it leaves that tokens obviously and then it leaves that that middle area where you have this that middle area where you have this that middle area where you have this interesting intersection of tasks that interesting intersection of tasks that interesting intersection of tasks that are um they're not too uncertain in that are um they're not too uncertain in that are um they're not too uncertain in that they or they they have um a degree of they or they they have um a degree of they or they they have um a degree of uncertainty where they're not great.
-
uncertainty where they're not great. uncertainty where they're not great. They're not just a a script or they're They're not just a a script or they're They're not just a a script or they're not out of distribution for training, not out of distribution for training, not out of distribution for training, but they have enough uncertainty to be but they have enough uncertainty to be but they have enough uncertainty to be interesting, but at at the same time, interesting, but at at the same time, interesting, but at at the same time, they also have a property of being they also have a property of being they also have a property of being relatively easy to uh validate the uh to relatively easy to uh validate the uh to relatively easy to uh validate the uh to check the value of them. And so they check the value of them. And so they check the value of them. And so they become in this place where they kind of become in this place where they kind of become in this place where they kind of become the shape of an MP style problem, become the shape of an MP style problem, become the shape of an MP style problem, which means they're easier to verify which means they're easier to verify which means they're easier to verify than to execute. And the reason I say than to execute. And the reason I say than to execute. And the reason I say that is because if you can figure out a that is because if you can figure out a that is because if you can figure out a pretty repeatable p pattern for pretty repeatable p pattern for pretty repeatable p pattern for verifying their their work, you can verifying their their work, you can verifying their their work, you can actually just throw agents at that actually just throw agents at that actually just throw agents at that problem as well. And so of course you problem as well. And so of course you problem as well. And so of course you don't just build the agent, you perhaps don't just build the agent, you perhaps don't just build the agent, you perhaps build the agent that verifies the work build the agent that verifies the work build the agent that verifies the work of the agent. Um and so yeah, this is uh of the agent. Um and so yeah, this is uh of the agent. Um and so yeah, this is uh perhaps uh this is a thought starter perhaps uh this is a thought starter perhaps uh this is a thought starter mostly kind of h kind of um still in the mostly kind of h kind of um still in the mostly kind of h kind of um still in the works. So happy to hear any thoughts on works. So happy to hear any thoughts on works. So happy to hear any thoughts on it. But if uh with this guidance, I hope it. But if uh with this guidance, I hope it. But if uh with this guidance, I hope uh when you're building your next agent, uh when you're building your next agent, uh when you're building your next agent, you can also figure out how to also you can also figure out how to also you can also figure out how to also build its mouth power. And thanks very build its mouth power. And thanks very build its mouth power. And thanks very much.
No summary available yet.
View original episode ↗