Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust
Read full transcript 18 segments
-
Hello everyone. My name is Amaya Hello everyone. My name is Amaya Bhavadkar and I am the field CTO at Bhavadkar and I am the field CTO at Bhavadkar and I am the field CTO at Brain Trust. Uh Brain Trust is a eval Brain Trust. Uh Brain Trust is a eval Brain Trust. Uh Brain Trust is a eval observability platform that helps AI observability platform that helps AI observability platform that helps AI teams build and improve their AI with teams build and improve their AI with teams build and improve their AI with confidence. confidence. confidence. So, So, So, I'm sure all of you, if not, you know, I I'm sure all of you, if not, you know, I I'm sure all of you, if not, you know, I I'm sure everyone here has built some I'm sure everyone here has built some I'm sure everyone here has built some application over the last couple of application over the last couple of application over the last couple of years that has a model at the center of years that has a model at the center of years that has a model at the center of it, right? Some sort of a chatbot or a it, right? Some sort of a chatbot or a it, right? Some sort of a chatbot or a AI agent or some system that's doing AI agent or some system that's doing AI agent or some system that's doing batch processing using AI at the heart batch processing using AI at the heart batch processing using AI at the heart of it. of it. of it. And I'm sure all of you over that time And I'm sure all of you over that time And I'm sure all of you over that time span have done significant uh changes to span have done significant uh changes to span have done significant uh changes to that application. You have either that application. You have either that application. You have either rewritten that application entirely or rewritten that application entirely or rewritten that application entirely or you have like done some pretty complex you have like done some pretty complex you have like done some pretty complex surgery on your application and the way surgery on your application and the way surgery on your application and the way it looks now compared to how it looked it looks now compared to how it looked it looks now compared to how it looked when you started is likely very very when you started is likely very very when you started is likely very very different. And I think everyone's different. And I think everyone's different. And I think everyone's probably uh experienced the same pattern probably uh experienced the same pattern probably uh experienced the same pattern which is like how building a demo with which is like how building a demo with which is like how building a demo with AI is really easy but making it AI is really easy but making it AI is really easy but making it production quality is really hard.
-
production quality is really hard. production quality is really hard. The same way when you're evolving your The same way when you're evolving your The same way when you're evolving your AI application and making significant AI application and making significant AI application and making significant changes to it, it can be very very changes to it, it can be very very changes to it, it can be very very challenging. Right? And uh the challenge challenging. Right? And uh the challenge challenging. Right? And uh the challenge is not because you built it the wrong is not because you built it the wrong is not because you built it the wrong way. The challenge is because the system way. The challenge is because the system way. The challenge is because the system around you is evolving and changing so around you is evolving and changing so around you is evolving and changing so dynamically, so rapidly. You know, the dynamically, so rapidly. You know, the dynamically, so rapidly. You know, the models are changing, the way your users models are changing, the way your users models are changing, the way your users use your application changes, the data use your application changes, the data use your application changes, the data that your application works with that your application works with that your application works with changes. And all of those things require changes. And all of those things require changes. And all of those things require you to continually make changes to your you to continually make changes to your you to continually make changes to your applications. applications. applications. And so if you look at, you know, the And so if you look at, you know, the And so if you look at, you know, the rate at which the models have evolved rate at which the models have evolved rate at which the models have evolved over the last couple of years, it's over the last couple of years, it's over the last couple of years, it's truly astonishing. like every few months truly astonishing. like every few months truly astonishing. like every few months there's a new release and that locks there's a new release and that locks there's a new release and that locks unlocks a you know a ton of new unlocks a you know a ton of new unlocks a you know a ton of new capabilities a ton of new features that capabilities a ton of new features that capabilities a ton of new features that were not present in the previous were not present in the previous were not present in the previous generation of the models right we have generation of the models right we have generation of the models right we have started seeing like models that got started seeing like models that got started seeing like models that got really good at working with tools models really good at working with tools models really good at working with tools models getting really good at handling very getting really good at handling very getting really good at handling very long context long context long context uh we started seeing models generate uh we started seeing models generate uh we started seeing models generate code that can be reliably and safely code that can be reliably and safely code that can be reliably and safely executed in uh sandboxes. We've seen executed in uh sandboxes. We've seen executed in uh sandboxes. We've seen memory systems becoming very memory systems becoming very memory systems becoming very sophisticated and practical. And so each sophisticated and practical. And so each sophisticated and practical. And so each of those was not a minor upgrade. It was of those was not a minor upgrade. It was of those was not a minor upgrade. It was not an incremental change to the not an incremental change to the not an incremental change to the previous state-of-the-art. It was a step previous state-of-the-art. It was a step previous state-of-the-art. It was a step function change, right? And so now we function change, right? And so now we function change, right? And so now we are moving from this era of like re uh
-
are moving from this era of like re uh are moving from this era of like re uh sort of iterating on improving our sort of iterating on improving our sort of iterating on improving our applications to replplatforming our applications to replplatforming our applications to replplatforming our applications because everything is applications because everything is applications because everything is changing so dramatically. So why can't you just drop in a new So why can't you just drop in a new model and ask expect your system to model and ask expect your system to model and ask expect your system to work? Um well the models the previous work? Um well the models the previous work? Um well the models the previous system that you built was built with system that you built was built with system that you built was built with some assumptions around the existing some assumptions around the existing some assumptions around the existing limitations and the constraints that the limitations and the constraints that the limitations and the constraints that the models had. Right? Your previous systems models had. Right? Your previous systems models had. Right? Your previous systems were built to account for the fact that were built to account for the fact that were built to account for the fact that your models weren't really as good at your models weren't really as good at your models weren't really as good at tool calling. for example, and so your tool calling. for example, and so your tool calling. for example, and so your system implemented a bunch of logic to system implemented a bunch of logic to system implemented a bunch of logic to make it work with those limitations. And make it work with those limitations. And make it work with those limitations. And so when you drop in a new model, right, so when you drop in a new model, right, so when you drop in a new model, right, uh you are not able to tap into the new uh you are not able to tap into the new uh you are not able to tap into the new capabilities, the new state-of-the-art capabilities, the new state-of-the-art capabilities, the new state-of-the-art without really restructuring your without really restructuring your without really restructuring your systems in a pretty dramatic way, right? systems in a pretty dramatic way, right? systems in a pretty dramatic way, right? And so in order to capture that kind of And so in order to capture that kind of And so in order to capture that kind of capability, the new unlock, you have to capability, the new unlock, you have to capability, the new unlock, you have to rearchitect.
-
And so as you rearchitect right um what And so as you rearchitect right um what happens is u models evolve. So you've happens is u models evolve. So you've happens is u models evolve. So you've got to go in and change your application got to go in and change your application got to go in and change your application architecture do a lot of work on on architecture do a lot of work on on architecture do a lot of work on on getting it to work with the new models. getting it to work with the new models. getting it to work with the new models. But that means that you also now have to But that means that you also now have to But that means that you also now have to update your evals. The way you ensure update your evals. The way you ensure update your evals. The way you ensure that your system is going to operate that your system is going to operate that your system is going to operate reliably, right? Because every new reliably, right? Because every new reliably, right? Because every new um uh unlock is potentially also giving um uh unlock is potentially also giving um uh unlock is potentially also giving you new surface area where things can go you new surface area where things can go you new surface area where things can go wrong. And so your evals now have to wrong. And so your evals now have to wrong. And so your evals now have to adapt and evolve to your new adapt and evolve to your new adapt and evolve to your new architecture. And so you know architecture. And so you know architecture. And so you know architecture follows model updates and architecture follows model updates and architecture follows model updates and your evals have to follow your your evals have to follow your your evals have to follow your architecture. architecture. architecture. So as I talked through the various So as I talked through the various So as I talked through the various generations of the AI systems generations of the AI systems generations of the AI systems architectures and how you do uh you know architectures and how you do uh you know architectures and how you do uh you know what that architecture is and how the what that architecture is and how the what that architecture is and how the eval eval eval to evolve with those architectural to evolve with those architectural to evolve with those architectural changes. I want to ground it in a real changes. I want to ground it in a real changes. I want to ground it in a real example and so what I want to talk about example and so what I want to talk about example and so what I want to talk about is um on the subsequent slides I'll is um on the subsequent slides I'll is um on the subsequent slides I'll share a bunch of notional evals but I share a bunch of notional evals but I share a bunch of notional evals but I want them to be grounded in a real want them to be grounded in a real want them to be grounded in a real agent. In this case we are going to look agent. In this case we are going to look agent. In this case we are going to look at this S sur agent. the SR agent is at this S sur agent. the SR agent is at this S sur agent. the SR agent is able to not only read information but able to not only read information but able to not only read information but it's able to take actions on and update it's able to take actions on and update it's able to take actions on and update systems so it can you know roll back a systems so it can you know roll back a systems so it can you know roll back a deployment or uh escalate it to a human deployment or uh escalate it to a human deployment or uh escalate it to a human page someone uh so it has access to read page someone uh so it has access to read page someone uh so it has access to read tools and write tools so let's see how tools and write tools so let's see how tools and write tools so let's see how this system would have evolved through this system would have evolved through this system would have evolved through the various generations of AI the various generations of AI the various generations of AI architectures
-
architectures architectures so let's start with the simplest case so let's start with the simplest case so let's start with the simplest case right this is how a lot of AI right this is how a lot of AI right this is how a lot of AI applications started about 3 years ago. applications started about 3 years ago. applications started about 3 years ago. This is a single prompt, a single model This is a single prompt, a single model This is a single prompt, a single model call. You have one input, one model call. You have one input, one model call. You have one input, one model call, one output. And so the focus of call, one output. And so the focus of call, one output. And so the focus of evaluations was on the final answer evaluations was on the final answer evaluations was on the final answer quality, right? Did you get the correct quality, right? Did you get the correct quality, right? Did you get the correct answer in terms of uh accuracy and answer in terms of uh accuracy and answer in terms of uh accuracy and factuality? Um or did the uh model factuality? Um or did the uh model factuality? Um or did the uh model hallucinate something? Did it make up hallucinate something? Did it make up hallucinate something? Did it make up stuff? or did it reference uh old non u stuff? or did it reference uh old non u stuff? or did it reference uh old non u the the previous knowledge that it had the the previous knowledge that it had the the previous knowledge that it had been trained on and not the latest uh been trained on and not the latest uh been trained on and not the latest uh information related to that subject. Um information related to that subject. Um information related to that subject. Um so in this case um you were really so in this case um you were really so in this case um you were really focusing primarily on the final answer focusing primarily on the final answer focusing primarily on the final answer that was your unit of evaluation. And so that was your unit of evaluation. And so that was your unit of evaluation. And so the approach was you would put together the approach was you would put together the approach was you would put together a golden data set. You will create a a golden data set. You will create a a golden data set. You will create a bunch of various scores that were bunch of various scores that were bunch of various scores that were looking at um encoding your definition looking at um encoding your definition looking at um encoding your definition of what good looks like that then you of what good looks like that then you of what good looks like that then you could evaluate the answers against. And could evaluate the answers against. And could evaluate the answers against. And this was great. This was a good way to this was great. This was a good way to this was great. This was a good way to get started. It was narrow because get started. It was narrow because get started. It was narrow because there's no tool calling. There's no there's no tool calling. There's no there's no tool calling. There's no orchestration. there's no uh retrieval, orchestration. there's no uh retrieval, orchestration. there's no uh retrieval, no other steps. It's just a simple call no other steps. It's just a simple call no other steps. It's just a simple call to the model. Uh but the next iteration to the model. Uh but the next iteration to the model. Uh but the next iteration of this was the chain. This is where you of this was the chain. This is where you of this was the chain. This is where you started doing a set of steps before you started doing a set of steps before you started doing a set of steps before you actually made the model call, right? Uh actually made the model call, right? Uh actually made the model call, right? Uh the typical rag application looked like the typical rag application looked like the typical rag application looked like it took the user input. It parsed some
-
it took the user input. It parsed some it took the user input. It parsed some information from the user output input. information from the user output input. information from the user output input. It then used that to retrieve It then used that to retrieve It then used that to retrieve information, then generate the context information, then generate the context information, then generate the context and then hand it over to the model. And and then hand it over to the model. And and then hand it over to the model. And then the model synthesizes reasons on then the model synthesizes reasons on then the model synthesizes reasons on that information, synthesizes an answer that information, synthesizes an answer that information, synthesizes an answer and you evaluate the answer. But there's and you evaluate the answer. But there's and you evaluate the answer. But there's a number of other places where things a number of other places where things a number of other places where things could go wrong. Yeah, your parser could could go wrong. Yeah, your parser could could go wrong. Yeah, your parser could extract the wrong information. It could extract the wrong information. It could extract the wrong information. It could retrieve the wrong context. The model retrieve the wrong context. The model retrieve the wrong context. The model could struggle with the context. Like in could struggle with the context. Like in could struggle with the context. Like in the early days, even though the model the early days, even though the model the early days, even though the model windows were the context window sizes windows were the context window sizes windows were the context window sizes were increasing, the models struggled to were increasing, the models struggled to were increasing, the models struggled to um reason over large context. So context um reason over large context. So context um reason over large context. So context stuffing could be an issue for the model stuffing could be an issue for the model stuffing could be an issue for the model performance. And so now you had multiple performance. And so now you had multiple performance. And so now you had multiple uh areas of failure. And so you needed uh areas of failure. And so you needed uh areas of failure. And so you needed to eval. But this was kind of very um what I But this was kind of very um what I would call very limited like it did would call very limited like it did would call very limited like it did things a very specific way all the time, things a very specific way all the time, things a very specific way all the time, right?
-
right? right? And so in late mid late 2023 early 24 And so in late mid late 2023 early 24 And so in late mid late 2023 early 24 the React paper became really popular. the React paper became really popular. the React paper became really popular. And so folks were looking at building um And so folks were looking at building um And so folks were looking at building um model um in a loop running a model in a model um in a loop running a model in a model um in a loop running a model in a loop where it could uh reason and act uh loop where it could uh reason and act uh loop where it could uh reason and act uh in a step-wise way. So the model could in a step-wise way. So the model could in a step-wise way. So the model could make tool calls. It could then make tool calls. It could then make tool calls. It could then understand what the tool calls returned understand what the tool calls returned understand what the tool calls returned uh reason on that data and then figure uh reason on that data and then figure uh reason on that data and then figure out what the next step was so that it out what the next step was so that it out what the next step was so that it could then continue to run this in a could then continue to run this in a could then continue to run this in a loop till the user intent was finally loop till the user intent was finally loop till the user intent was finally satisfied or the model ran out of the satisfied or the model ran out of the satisfied or the model ran out of the iteration budget. Right? And so this was iteration budget. Right? And so this was iteration budget. Right? And so this was great because it now gives you gives the great because it now gives you gives the great because it now gives you gives the model the AI system a lot more model the AI system a lot more model the AI system a lot more flexibility. It's not pinned down to flexibility. It's not pinned down to flexibility. It's not pinned down to operating in a very specific workflow. operating in a very specific workflow. operating in a very specific workflow. It now is able to reason on the various It now is able to reason on the various It now is able to reason on the various intents and it's able to self-organize, intents and it's able to self-organize, intents and it's able to self-organize, self-chestrate and complete the user self-chestrate and complete the user self-chestrate and complete the user tasks. Unfortunately, the models of that tasks. Unfortunately, the models of that tasks. Unfortunately, the models of that era were not as robust as they needed to era were not as robust as they needed to era were not as robust as they needed to be. So, you know, models struggled with be. So, you know, models struggled with be. So, you know, models struggled with tool callings. They got the arguments tool callings. They got the arguments tool callings. They got the arguments wrong. The models struggled with wrong. The models struggled with wrong. The models struggled with orchestration. So, they called the wrong orchestration. So, they called the wrong orchestration. So, they called the wrong tools. The models still had challenges tools. The models still had challenges tools. The models still had challenges with reasoning. they weren't necessarily with reasoning. they weren't necessarily with reasoning. they weren't necessarily doing a great job of, you know, dealing doing a great job of, you know, dealing doing a great job of, you know, dealing with long context. So you had things with long context. So you had things with long context. So you had things like context collapse. And so while the like context collapse. And so while the like context collapse. And so while the idea was like really really exciting, idea was like really really exciting, idea was like really really exciting, um, it fell short of delivering on the
-
um, it fell short of delivering on the um, it fell short of delivering on the actual promise. actual promise. actual promise. And so what does what do you do when And so what does what do you do when And so what does what do you do when your model can't be controlled, right? your model can't be controlled, right? your model can't be controlled, right? You take the control and you bake that You take the control and you bake that You take the control and you bake that control into the system that you're control into the system that you're control into the system that you're building around the model. And so teams building around the model. And so teams building around the model. And so teams started moving towards these kind of started moving towards these kind of started moving towards these kind of workflow graphs, right? Um they started workflow graphs, right? Um they started workflow graphs, right? Um they started building the orchestration and the building the orchestration and the building the orchestration and the execution and planning logic into the execution and planning logic into the execution and planning logic into the the system itself either as a graph or the system itself either as a graph or the system itself either as a graph or as a state machine. And so you took as a state machine. And so you took as a state machine. And so you took control of the orchestration while you control of the orchestration while you control of the orchestration while you allowed the models to operate at the allowed the models to operate at the allowed the models to operate at the node level. And that way you got a lot node level. And that way you got a lot node level. And that way you got a lot more uh reliability and predictability more uh reliability and predictability more uh reliability and predictability in how your AI was going to operate in how your AI was going to operate in how your AI was going to operate across those various intents. across those various intents. across those various intents. But then the problem is you are now But then the problem is you are now But then the problem is you are now building a system that is designed to building a system that is designed to building a system that is designed to work for a specific set of intents for a work for a specific set of intents for a work for a specific set of intents for a specific types of use cases. And as you specific types of use cases. And as you specific types of use cases. And as you start hand, you know, the system starts start hand, you know, the system starts start hand, you know, the system starts interacting with with instances that are interacting with with instances that are interacting with with instances that are outside that distribution, the system outside that distribution, the system outside that distribution, the system starts struggling with that, right? You starts struggling with that, right? You starts struggling with that, right? You expect um you know a certain set of expect um you know a certain set of expect um you know a certain set of applications or u user interactions to applications or u user interactions to applications or u user interactions to work well because they can be fulfilled work well because they can be fulfilled work well because they can be fulfilled by the orchestration that you have by the orchestration that you have by the orchestration that you have designed. But when your the user intent designed. But when your the user intent designed. But when your the user intent needs to be requires other things to needs to be requires other things to needs to be requires other things to happen beyond what's specified in the happen beyond what's specified in the happen beyond what's specified in the orchestration, the system can start um
-
orchestration, the system can start um orchestration, the system can start um you know breaking at the seams. And uh you know breaking at the seams. And uh you know breaking at the seams. And uh in order to do that, folks were now in order to do that, folks were now in order to do that, folks were now building a lot more complexity into building a lot more complexity into building a lot more complexity into their orchestration logic. And so you're their orchestration logic. And so you're their orchestration logic. And so you're building these special uh branches and building these special uh branches and building these special uh branches and way you handle special intents in the way you handle special intents in the way you handle special intents in the complex graph that described your complex graph that described your complex graph that described your system. And so what that means is like system. And so what that means is like system. And so what that means is like you had now a ton of different surfaces you had now a ton of different surfaces you had now a ton of different surfaces for failure. So you now had to deal with for failure. So you now had to deal with for failure. So you now had to deal with uh you know uh dealing with uh branch uh you know uh dealing with uh branch uh you know uh dealing with uh branch consistency and branching logic consistency and branching logic consistency and branching logic failures. You had to deal with things failures. You had to deal with things failures. You had to deal with things like the contracts between the nodes not like the contracts between the nodes not like the contracts between the nodes not working out well. Uh you had to deal working out well. Uh you had to deal working out well. Uh you had to deal with the limitations of uh the nodes with the limitations of uh the nodes with the limitations of uh the nodes that were you know built for a specific that were you know built for a specific that were you know built for a specific set of use cases. So you know there were set of use cases. So you know there were set of use cases. So you know there were classifier nodes for example and they classifier nodes for example and they classifier nodes for example and they could make mistakes and so you could now could make mistakes and so you could now could make mistakes and so you could now have a significant amount of um you know have a significant amount of um you know have a significant amount of um you know areas where you could uh where the areas where you could uh where the areas where you could uh where the system could fail. And so your evals now system could fail. And so your evals now system could fail. And so your evals now have to not only look at u you know the have to not only look at u you know the have to not only look at u you know the overall orchestration but they now have overall orchestration but they now have overall orchestration but they now have to you have to have node level evals.
-
to you have to have node level evals. to you have to have node level evals. you have to uh make sure that you have you have to uh make sure that you have you have to uh make sure that you have evalu uh you know how you do retry loops. uh you know how you do retry loops. There's a lot of complex behaviors of There's a lot of complex behaviors of There's a lot of complex behaviors of the system that now need to be evaluated the system that now need to be evaluated the system that now need to be evaluated in addition to all the other things that in addition to all the other things that in addition to all the other things that you were evaluating before. So So the graphs were kind of popular like in the graphs were kind of popular like in the graphs were kind of popular like in in late 24 early 25 and so a lot of in late 24 early 25 and so a lot of in late 24 early 25 and so a lot of systems were now implemented using systems were now implemented using systems were now implemented using certain frameworks and they were now in certain frameworks and they were now in certain frameworks and they were now in production. Uh but then um Anthropic and production. Uh but then um Anthropic and production. Uh but then um Anthropic and OpenAI launched some amazing new model OpenAI launched some amazing new model OpenAI launched some amazing new model capabilities mid late 25 and what that capabilities mid late 25 and what that capabilities mid late 25 and what that was like tool calling became extremely was like tool calling became extremely was like tool calling became extremely reliable. We started uh seeing uh much reliable. We started uh seeing uh much reliable. We started uh seeing uh much better orchestration control. Uh the better orchestration control. Uh the better orchestration control. Uh the models were able to plan a lot more models were able to plan a lot more models were able to plan a lot more effectively accurately. They were able effectively accurately. They were able effectively accurately. They were able to manage long horizon tasks. they were to manage long horizon tasks. they were to manage long horizon tasks. they were able to do a much better job of able to do a much better job of able to do a much better job of introspecting and course correcting. And introspecting and course correcting. And introspecting and course correcting. And so like as things went a little off so like as things went a little off so like as things went a little off track, the models were able to, you track, the models were able to, you track, the models were able to, you know, understand that and bring the know, understand that and bring the know, understand that and bring the execution back on track. And so what execution back on track. And so what execution back on track. And so what that meant was a lot of these u u that meant was a lot of these u u that meant was a lot of these u u graph-based systems were not able to graph-based systems were not able to graph-based systems were not able to take advantage of these new take advantage of these new take advantage of these new capabilities. they were still running
-
capabilities. they were still running capabilities. they were still running into some of those like brittleleness into some of those like brittleleness into some of those like brittleleness issues that the new model state-of-art issues that the new model state-of-art issues that the new model state-of-art had unlocked and so um we started had unlocked and so um we started had unlocked and so um we started looking at building out um the react looking at building out um the react looking at building out um the react loop again that's that started working loop again that's that started working loop again that's that started working and so now you had this new AI systems and so now you had this new AI systems and so now you had this new AI systems that could effectively reliably work in that could effectively reliably work in that could effectively reliably work in a loop they could make those tool calls a loop they could make those tool calls a loop they could make those tool calls they could figure out the next step and they could figure out the next step and they could figure out the next step and then they could u essentially go in and then they could u essentially go in and then they could u essentially go in and um fulfill the user intent. But the way um fulfill the user intent. But the way um fulfill the user intent. But the way they worked was very it had a high they worked was very it had a high they worked was very it had a high degree of variance. So every trajectory degree of variance. So every trajectory degree of variance. So every trajectory for the same input if you ran it a for the same input if you ran it a for the same input if you ran it a couple of times you would see you know couple of times you would see you know couple of times you would see you know dramatically different trajectories dramatically different trajectories dramatically different trajectories while yielding the right answer. And so while yielding the right answer. And so while yielding the right answer. And so now there's a lot of variance that you now there's a lot of variance that you now there's a lot of variance that you have to deal with. So now instead of have to deal with. So now instead of have to deal with. So now instead of just focusing on a specific eval the just focusing on a specific eval the just focusing on a specific eval the unit of eval was no longer just one eval unit of eval was no longer just one eval unit of eval was no longer just one eval now you're looking at doing an analysis now you're looking at doing an analysis now you're looking at doing an analysis of the distribution of the evals you're of the distribution of the evals you're of the distribution of the evals you're taking the same eval you're running it taking the same eval you're running it taking the same eval you're running it multiple times you're running it k times multiple times you're running it k times multiple times you're running it k times and you're ensuring that uh you get a and you're ensuring that uh you get a and you're ensuring that uh you get a statistically relevant signal from that statistically relevant signal from that statistically relevant signal from that eval so now new metrics like uh pass at eval so now new metrics like uh pass at eval so now new metrics like uh pass at k and pass raise to k or pass wedge k k and pass raise to k or pass wedge k k and pass raise to k or pass wedge k these were the new metrics that these were the new metrics that these were the new metrics that certainly started to make a lot of certainly started to make a lot of certainly started to make a lot of sense. pass at K is like if you take the sense. pass at K is like if you take the sense. pass at K is like if you take the same that eval and you run it K times same that eval and you run it K times same that eval and you run it K times does it succeed at least once and that does it succeed at least once and that does it succeed at least once and that is a measure of its capability and pass
-
is a measure of its capability and pass is a measure of its capability and pass wedge K is like if you run that eval wedge K is like if you run that eval wedge K is like if you run that eval multiple times how many times of those K multiple times how many times of those K multiple times how many times of those K instances does it run successfully instances does it run successfully instances does it run successfully that's a measure of its uh reliability that's a measure of its uh reliability that's a measure of its uh reliability and so now you can understand whether and so now you can understand whether and so now you can understand whether your system with a high pass at K uh you your system with a high pass at K uh you your system with a high pass at K uh you know is reliable by seeing seeing how it know is reliable by seeing seeing how it know is reliable by seeing seeing how it you know by measuring the pass wedge K you know by measuring the pass wedge K you know by measuring the pass wedge K metric for example. So [snorts] this metric for example. So [snorts] this metric for example. So [snorts] this gives you a lot more um you know u gives you a lot more um you know u gives you a lot more um you know u understanding of like how your system is understanding of like how your system is understanding of like how your system is working what the failure sources are and working what the failure sources are and working what the failure sources are and how you work on those right and then how you work on those right and then how you work on those right and then more recently what we've seen is um more recently what we've seen is um more recently what we've seen is um there's a big shift from it's your there's a big shift from it's your there's a big shift from it's your system is not just a model running in a system is not just a model running in a system is not just a model running in a loop right it becomes a product system loop right it becomes a product system loop right it becomes a product system it's that there's a model in the loop it's that there's a model in the loop it's that there's a model in the loop that's augmented by a lot of peripheral that's augmented by a lot of peripheral that's augmented by a lot of peripheral components you know you have a memory components you know you have a memory components you know you have a memory system that is able to provide robust system that is able to provide robust system that is able to provide robust memory storage and memory um retrieval memory storage and memory um retrieval memory storage and memory um retrieval capabilities uh within a session cross capabilities uh within a session cross capabilities uh within a session cross sessions. Uh models can tap into this sessions. Uh models can tap into this sessions. Uh models can tap into this memory to you know improve upon their memory to you know improve upon their memory to you know improve upon their runs in subsequent instances by learning runs in subsequent instances by learning runs in subsequent instances by learning from previous runs for example. You've from previous runs for example. You've from previous runs for example. You've got robust code execution uh sandboxes got robust code execution uh sandboxes got robust code execution uh sandboxes now and so you can run model generated now and so you can run model generated now and so you can run model generated code reliably robustly on uh uh during code reliably robustly on uh uh during code reliably robustly on uh uh during uh execution. You've got um MCP and uh execution. You've got um MCP and uh execution. You've got um MCP and skill uh directories that the model can skill uh directories that the model can skill uh directories that the model can now tap into and you can you know weave
-
now tap into and you can you know weave now tap into and you can you know weave in extensibility. You now have things in extensibility. You now have things in extensibility. You now have things like a skills repository or a skill like a skills repository or a skill like a skills repository or a skill systems that can be used to continually systems that can be used to continually systems that can be used to continually augment the the the capabilities of augment the the the capabilities of augment the the the capabilities of models through you know symbolic models through you know symbolic models through you know symbolic instructions. And so uh now you know instructions. And so uh now you know instructions. And so uh now you know like uh these systems are getting pretty like uh these systems are getting pretty like uh these systems are getting pretty complex and as a result uh you know if complex and as a result uh you know if complex and as a result uh you know if you are continuing to to use the eval you are continuing to to use the eval you are continuing to to use the eval from the previous generation you're from the previous generation you're from the previous generation you're going to get sort of a partial coverage going to get sort of a partial coverage going to get sort of a partial coverage of your system. you're not going to see of your system. you're not going to see of your system. you're not going to see uh how your system is fragile in ways uh how your system is fragile in ways uh how your system is fragile in ways because of the unlock because of the new because of the unlock because of the new because of the unlock because of the new surface that you have uh you know uh surface that you have uh you know uh surface that you have uh you know uh unlocked in your new system. unlocked in your new system. unlocked in your new system. So So So what that means is what that means is what that means is um just reflecting back on the pattern um just reflecting back on the pattern um just reflecting back on the pattern is like you know all of these model is like you know all of these model is like you know all of these model innovations resulted in in you know innovations resulted in in you know innovations resulted in in you know corresponding shift in the architectures corresponding shift in the architectures corresponding shift in the architectures and so so you've seen these waves of and so so you've seen these waves of and so so you've seen these waves of architecture and then what's needed is architecture and then what's needed is architecture and then what's needed is like your evals to be congru congruent like your evals to be congru congruent like your evals to be congru congruent with that architecture right uh because with that architecture right uh because with that architecture right uh because ultimately it's the eval that are sort ultimately it's the eval that are sort ultimately it's the eval that are sort of your durable asset that describe how of your durable asset that describe how of your durable asset that describe how your system is supposed to work. And as your system is supposed to work. And as your system is supposed to work. And as you go through these generational you go through these generational you go through these generational shifts, that's a good way to ensure that shifts, that's a good way to ensure that shifts, that's a good way to ensure that you know your system your user users you know your system your user users you know your system your user users experience your system in a way that experience your system in a way that experience your system in a way that things that were working are not broken, things that were working are not broken, things that were working are not broken, but it's unlocked a bunch of new but it's unlocked a bunch of new but it's unlocked a bunch of new capability.
-
capability. capability. And so everyone's seen this, you know, And so everyone's seen this, you know, And so everyone's seen this, you know, diagram of this flywheel. Everyone's diagram of this flywheel. Everyone's diagram of this flywheel. Everyone's sort of like bought into it sort of like bought into it sort of like bought into it conceptually, right? the idea of conceptually, right? the idea of conceptually, right? the idea of harvesting data from production to harvesting data from production to harvesting data from production to inform your eval so that your evals are inform your eval so that your evals are inform your eval so that your evals are reflective of the real world. I think reflective of the real world. I think reflective of the real world. I think that all makes sense, right? And and that all makes sense, right? And and that all makes sense, right? And and this is the way that you know teams that this is the way that you know teams that this is the way that you know teams that are doing a great job at building and are doing a great job at building and are doing a great job at building and shipping and improving their AI systems, shipping and improving their AI systems, shipping and improving their AI systems, they they they follow this workflow they they they follow this workflow they they they follow this workflow pretty religiously. pretty religiously. pretty religiously. Um so I've talked to a lot of teams and Um so I've talked to a lot of teams and Um so I've talked to a lot of teams and I think while there is a general I think while there is a general I think while there is a general acceptance that yeah you need to run acceptance that yeah you need to run acceptance that yeah you need to run that workflow um in practice a lot of that workflow um in practice a lot of that workflow um in practice a lot of teams don't do that their eval are teams don't do that their eval are teams don't do that their eval are somewhat static and even if you're not somewhat static and even if you're not somewhat static and even if you're not changing your AI agent architecture changing your AI agent architecture changing your AI agent architecture you're you know by not really being you're you know by not really being you're you know by not really being disciplined about running that that disciplined about running that that disciplined about running that that workflow that flywheel you are now workflow that flywheel you are now workflow that flywheel you are now getting stagnant evals that are not getting stagnant evals that are not getting stagnant evals that are not being as effective in helping you being as effective in helping you being as effective in helping you measure and improve the quality of your measure and improve the quality of your measure and improve the quality of your AI. And especially as you go through AI. And especially as you go through AI. And especially as you go through this generational shift, it's really this generational shift, it's really this generational shift, it's really important that you need a mechanism to important that you need a mechanism to important that you need a mechanism to not only harvest data from production in not only harvest data from production in not only harvest data from production in a way that shows you failures that a way that shows you failures that a way that shows you failures that you are looking out for because you you are looking out for because you you are looking out for because you defined what good looks like as part of defined what good looks like as part of defined what good looks like as part of your evals. But you also want something your evals. But you also want something your evals. But you also want something to shine a light on the new failure to shine a light on the new failure to shine a light on the new failure types, right? the system is going to types, right? the system is going to types, right? the system is going to fail in new and novel ways in ways that fail in new and novel ways in ways that fail in new and novel ways in ways that you might not have anticipated and you you might not have anticipated and you you might not have anticipated and you now need to start harvesting that data
-
now need to start harvesting that data now need to start harvesting that data in a meaningful way. And you want to do in a meaningful way. And you want to do in a meaningful way. And you want to do this again as as part of the flywheel. this again as as part of the flywheel. this again as as part of the flywheel. And so this is where you need systems to And so this is where you need systems to And so this is where you need systems to come in and u shine a light on things come in and u shine a light on things come in and u shine a light on things that are broken in ways that you had that are broken in ways that you had that are broken in ways that you had anticipated, but also broken in a way in anticipated, but also broken in a way in anticipated, but also broken in a way in ways that you had not anticipated. And ways that you had not anticipated. And ways that you had not anticipated. And this is really important. this is really important. this is really important. So I'm going to quickly talk a little So I'm going to quickly talk a little So I'm going to quickly talk a little bit about like how we do this in brain bit about like how we do this in brain bit about like how we do this in brain trust. So brain trust provides all the trust. So brain trust provides all the trust. So brain trust provides all the components that you need to run this components that you need to run this components that you need to run this flywheel. We've got evals, we've got flywheel. We've got evals, we've got flywheel. We've got evals, we've got observability. We have ways in which you observability. We have ways in which you observability. We have ways in which you can get insights from your production can get insights from your production can get insights from your production data to harvest u new eval cases that data to harvest u new eval cases that data to harvest u new eval cases that you can then pass off to the to the team you can then pass off to the to the team you can then pass off to the to the team that they can then use to hill climb and that they can then use to hill climb and that they can then use to hill climb and improve your AI system. But topics is a improve your AI system. But topics is a improve your AI system. But topics is a really cool feature. What topics does, really cool feature. What topics does, really cool feature. What topics does, it does a cluster analysis on all of it does a cluster analysis on all of it does a cluster analysis on all of your production data. And so the idea your production data. And so the idea your production data. And so the idea over here is now you are able to find over here is now you are able to find over here is now you are able to find new categories of failures that you had new categories of failures that you had new categories of failures that you had not anticipated. So your system is now not anticipated. So your system is now not anticipated. So your system is now able to look at all what's going on in able to look at all what's going on in able to look at all what's going on in production and it's able to now start production and it's able to now start production and it's able to now start surfacing these new failure modes that surfacing these new failure modes that surfacing these new failure modes that tell you here's a new new failure uh you tell you here's a new new failure uh you tell you here's a new new failure uh you know um uh situation that you hadn't know um uh situation that you hadn't know um uh situation that you hadn't thought about and you didn't have any thought about and you didn't have any thought about and you didn't have any guardrails in place. didn't have any guardrails in place. didn't have any guardrails in place. didn't have any eval in place and so now it's really eval in place and so now it's really eval in place and so now it's really easy for teams to expand the set of easy for teams to expand the set of easy for teams to expand the set of their evals to now cover those kind of
-
their evals to now cover those kind of their evals to now cover those kind of new failures. And so this is this is a new failures. And so this is this is a new failures. And so this is this is a pretty exciting uh capability in brain pretty exciting uh capability in brain pretty exciting uh capability in brain trust that enables these teams to trust that enables these teams to trust that enables these teams to continually not only get new failure continually not only get new failure continually not only get new failure examples for known failure modes but examples for known failure modes but examples for known failure modes but more importantly as they make these more importantly as they make these more importantly as they make these systemic architectural changes they're systemic architectural changes they're systemic architectural changes they're able to also understand the new ways in able to also understand the new ways in able to also understand the new ways in which your system is going to fail and which your system is going to fail and which your system is going to fail and build out effective data sets from build out effective data sets from build out effective data sets from production data. So I think the takeaway for today's talk So I think the takeaway for today's talk is that is that is that the models will keep on changing. Uh I I the models will keep on changing. Uh I I the models will keep on changing. Uh I I don't think we're going to see any don't think we're going to see any don't think we're going to see any slowdown. I don't think we have hit a slowdown. I don't think we have hit a slowdown. I don't think we have hit a plateau yet. I think there are lots of plateau yet. I think there are lots of plateau yet. I think there are lots of unlocks that are coming down um this unlocks that are coming down um this unlocks that are coming down um this road. Um and as a result you will be road. Um and as a result you will be road. Um and as a result you will be making significant making significant making significant changes to your AI systems. you know, changes to your AI systems. you know, changes to your AI systems. you know, you'll be doing a lot of surgery on your you'll be doing a lot of surgery on your you'll be doing a lot of surgery on your AI agents in the coming months, years.
-
AI agents in the coming months, years. AI agents in the coming months, years. And so it's really important that you And so it's really important that you And so it's really important that you have a robust workflow system in place have a robust workflow system in place have a robust workflow system in place to ensure that as you make those to ensure that as you make those to ensure that as you make those changes, as you incorporate these new changes, as you incorporate these new changes, as you incorporate these new models into your systems, that your models into your systems, that your models into your systems, that your systems continue to get better at doing systems continue to get better at doing systems continue to get better at doing new things, but also continue to work new things, but also continue to work new things, but also continue to work well for the things that they were doing well for the things that they were doing well for the things that they were doing before. And so building out like a before. And so building out like a before. And so building out like a robust eval robust eval robust eval discipline uh with the right tools and discipline uh with the right tools and discipline uh with the right tools and the right automation and the right the right automation and the right the right automation and the right systems becomes paramount to manage systems becomes paramount to manage systems becomes paramount to manage these generational changes. And so these generational changes. And so these generational changes. And so ultimately what you want is um to really ultimately what you want is um to really ultimately what you want is um to really uh index on that flywheel and make it uh index on that flywheel and make it uh index on that flywheel and make it part of your workflow so that uh you part of your workflow so that uh you part of your workflow so that uh you know the ability to know the ability to know the ability to improve incrementally when the changes improve incrementally when the changes improve incrementally when the changes in the system are incremental and the in the system are incremental and the in the system are incremental and the ability to improve your system in a in ability to improve your system in a in ability to improve your system in a in sort of a step function way are both sort of a step function way are both sort of a step function way are both supported by your evals.
-
supported by your evals. supported by your evals. So with that, So with that, So with that, I want to say thank you. I want to say thank you. I want to say thank you. [applause]
Summary
The main theme is the challenge of evolving AI applications due to rapid changes in models, data, and user behavior, contrasting the ease of building demos with the difficulty of production quality. The topic revolves around AI observability and how platforms like Brain Trust help teams build and improve AI confidently amidst these dynamic changes. The practical takeaway is that continuous adaptation and robust observability are essential for maintaining and enhancing AI applications in a fast-evolving landscape.