FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft
Read full transcript 16 segments
-
Okay. Um, good morning everyone. So, um, Okay. Um, good morning everyone. So, um, I'm Tisha and I have Sushim with me as I'm Tisha and I have Sushim with me as I'm Tisha and I have Sushim with me as my co-presenter. All right. So, we'll be my co-presenter. All right. So, we'll be my co-presenter. All right. So, we'll be talking about the most expensive talking about the most expensive talking about the most expensive question in AI today. I think a lot of question in AI today. I think a lot of question in AI today. I think a lot of you would have come across the scenario you would have come across the scenario you would have come across the scenario that um you know when you opened an AI that um you know when you opened an AI that um you know when you opened an AI bill like through your agent workflows bill like through your agent workflows bill like through your agent workflows um you couldn't actually trace back um you couldn't actually trace back um you couldn't actually trace back where that bill was actually coming from where that bill was actually coming from where that bill was actually coming from right and um and I don't think that's a right and um and I don't think that's a right and um and I don't think that's a problem right now because right now the problem right now because right now the problem right now because right now the industry is valuing you know um token industry is valuing you know um token industry is valuing you know um token maxing that is like spending the most maxing that is like spending the most maxing that is like spending the most amount of tokens for exploration for all amount of tokens for exploration for all amount of tokens for exploration for all of those purposes of those purposes of those purposes And um people are proud to call And um people are proud to call And um people are proud to call themselves token billionaires and um I themselves token billionaires and um I themselves token billionaires and um I think that's all right but this talk is think that's all right but this talk is think that's all right but this talk is you know the shift from token maxing to you know the shift from token maxing to you know the shift from token maxing to value maxing you know how do we get value maxing you know how do we get value maxing you know how do we get there and um we'll talk about it from there and um we'll talk about it from there and um we'll talk about it from this question um who spent all the this question um who spent all the this question um who spent all the tokens and um if anyone spent all the tokens and um if anyone spent all the tokens and um if anyone spent all the tokens there has to be value associated tokens there has to be value associated tokens there has to be value associated with this right and that is um the talk with this right and that is um the talk with this right and that is um the talk about.
-
All right. Now in order to minimize the All right. Now in order to minimize the gap you know from token maxing to value gap you know from token maxing to value gap you know from token maxing to value maxing we'll kind of see we'll observe maxing we'll kind of see we'll observe maxing we'll kind of see we'll observe the patterns which the like the existing the patterns which the like the existing the patterns which the like the existing um u like the past software evolution um u like the past software evolution um u like the past software evolution eras had like for instance when we talk eras had like for instance when we talk eras had like for instance when we talk about the SAS era the interface was UI about the SAS era the interface was UI about the SAS era the interface was UI and the control was in the form of usage and the control was in the form of usage and the control was in the form of usage caps right like or the seat limits or caps right like or the seat limits or caps right like or the seat limits or tier based policies tier based policies tier based policies Now when we moved on to the cloud era, Now when we moved on to the cloud era, Now when we moved on to the cloud era, the control surface again changed. The the control surface again changed. The the control surface again changed. The model became pay as you go and the model became pay as you go and the model became pay as you go and the control moved like in the form of control moved like in the form of control moved like in the form of autoprovisioning and you know autoprovisioning and you know autoprovisioning and you know autoscaling policies. autoscaling policies. autoscaling policies. Now we are in the agentic era right and Now we are in the agentic era right and Now we are in the agentic era right and um now how the cost is calculated here um now how the cost is calculated here um now how the cost is calculated here is in the form of model calls right like is in the form of model calls right like is in the form of model calls right like u how like the code calls your model but u how like the code calls your model but u how like the code calls your model but what we've observed is that there isn't what we've observed is that there isn't what we've observed is that there isn't a proper control plane in place for that a proper control plane in place for that a proper control plane in place for that like we do have control plane in place like we do have control plane in place like we do have control plane in place for you in place as model gateways where for you in place as model gateways where for you in place as model gateways where they're um are hard caps or there is they're um are hard caps or there is they're um are hard caps or there is model routing to downgrade the model but model routing to downgrade the model but model routing to downgrade the model but the part like where the code you know the part like where the code you know the part like where the code you know calls the model that um is what we'll be calls the model that um is what we'll be calls the model that um is what we'll be talking about uh today talking about uh today talking about uh today and um we also you know see um like in
-
and um we also you know see um like in and um we also you know see um like in the last year we've seen a lot of the last year we've seen a lot of the last year we've seen a lot of unbounded consumption happening like um unbounded consumption happening like um unbounded consumption happening like um if you've read the news. There was news if you've read the news. There was news if you've read the news. There was news about the like the uh AI budget for Uber about the like the uh AI budget for Uber about the like the uh AI budget for Uber getting exhausted within 4 months and um getting exhausted within 4 months and um getting exhausted within 4 months and um there were companies who like who ran there were companies who like who ran there were companies who like who ran into you know into you know into you know like hundreds of millions of dollars like hundreds of millions of dollars like hundreds of millions of dollars within just months or days and like within just months or days and like within just months or days and like there were a lot of um like other news there were a lot of um like other news there were a lot of um like other news in place as well where like these in place as well where like these in place as well where like these runaway loops um led to a very like runaway loops um led to a very like runaway loops um led to a very like massive increase in the cost and there massive increase in the cost and there massive increase in the cost and there wasn't proper mechanisms to control it. wasn't proper mechanisms to control it. wasn't proper mechanisms to control it. Um so when we see all of this the first Um so when we see all of this the first Um so when we see all of this the first thing that comes to our mind is is there thing that comes to our mind is is there thing that comes to our mind is is there a tool or is there a product to save us? a tool or is there a product to save us? a tool or is there a product to save us? But uh we'll instead talk about the But uh we'll instead talk about the But uh we'll instead talk about the first principles of how you know we can first principles of how you know we can first principles of how you know we can design a system which is actually true design a system which is actually true design a system which is actually true enough to solve the problem from the enough to solve the problem from the enough to solve the problem from the very root. So for that let's um like very root. So for that let's um like very root. So for that let's um like dive onto the principles. First of all, dive onto the principles. First of all, dive onto the principles. First of all, let's talk about token being the unit of let's talk about token being the unit of let's talk about token being the unit of cost. Right? We are charged in terms of cost. Right? We are charged in terms of cost. Right? We are charged in terms of token. So the now we have to see value token. So the now we have to see value token. So the now we have to see value also in terms of token. Right? Next um also in terms of token. Right? Next um also in terms of token. Right? Next um we all know that cost is created at the we all know that cost is created at the we all know that cost is created at the LLM like the model call boundary. Um so LLM like the model call boundary. Um so LLM like the model call boundary. Um so that is what we'll have to track and if
-
that is what we'll have to track and if that is what we'll have to track and if we don't have proper attribution like if we don't have proper attribution like if we don't have proper attribution like if we don't know what agent want run made we don't know what agent want run made we don't know what agent want run made that particular call we we can't you that particular call we we can't you that particular call we we can't you know control it right we we just know know control it right we we just know know control it right we we just know the like the broad uh picture of what the like the broad uh picture of what the like the broad uh picture of what went wrong but we don't we can't you went wrong but we don't we can't you went wrong but we don't we can't you know trace it back or narrow it down. So know trace it back or narrow it down. So know trace it back or narrow it down. So that is why attribution is a very that is why attribution is a very that is why attribution is a very important element to have and um like important element to have and um like important element to have and um like once you know which particular run or once you know which particular run or once you know which particular run or which particular agent is actually you which particular agent is actually you which particular agent is actually you know attributing to the cost you should know attributing to the cost you should know attributing to the cost you should have proper policies in place to have proper policies in place to have proper policies in place to actually stop it. Like um let's take actually stop it. Like um let's take actually stop it. Like um let's take example that um if you have a you know example that um if you have a you know example that um if you have a you know um a loop which is you know running um um a loop which is you know running um um a loop which is you know running um very excessively and which is not very excessively and which is not very excessively and which is not required or you know if your context is required or you know if your context is required or you know if your context is growing very out of range. you should growing very out of range. you should growing very out of range. you should have in place policies which can um like have in place policies which can um like have in place policies which can um like solve that particular thing there and solve that particular thing there and solve that particular thing there and there instead of halting that and if um there instead of halting that and if um there instead of halting that and if um and as the last resort only a like a and as the last resort only a like a and as the last resort only a like a halting or a should happen from a budget halting or a should happen from a budget halting or a should happen from a budget cap. So these are the first principles.
-
cap. So these are the first principles. cap. So these are the first principles. Now let's see how we can you know define Now let's see how we can you know define Now let's see how we can you know define an ideal uh platform on top of that from an ideal uh platform on top of that from an ideal uh platform on top of that from these principles which we talked about. these principles which we talked about. these principles which we talked about. All right. Uh so one thing which is very All right. Uh so one thing which is very All right. Uh so one thing which is very important that which matters here is important that which matters here is important that which matters here is that u when we talk about um like the um that u when we talk about um like the um that u when we talk about um like the um existing frameworks for token ops or for existing frameworks for token ops or for existing frameworks for token ops or for token management most of them are at the token management most of them are at the token management most of them are at the um like u basically monitor the model uh um like u basically monitor the model uh um like u basically monitor the model uh request. they like they are like model request. they like they are like model request. they like they are like model gateways which will u you know u gateways which will u you know u gateways which will u you know u basically um do like model routing or basically um do like model routing or basically um do like model routing or hard budget capping. But what we need hard budget capping. But what we need hard budget capping. But what we need right now is something which you know um right now is something which you know um right now is something which you know um like monitors you at the run instead. like monitors you at the run instead. like monitors you at the run instead. Like um if you see we need something uh Like um if you see we need something uh Like um if you see we need something uh which can control the loop between like which can control the loop between like which can control the loop between like the agent call between the tool um and the agent call between the tool um and the agent call between the tool um and the agent. something you know which can the agent. something you know which can the agent. something you know which can um uh see or control the the spawning of um uh see or control the the spawning of um uh see or control the the spawning of multiple sub aents happening from a one multiple sub aents happening from a one multiple sub aents happening from a one main agent or um like something which main agent or um like something which main agent or um like something which can control the growing of context. So can control the growing of context. So can control the growing of context. So like that is the need of the right and like that is the need of the right and like that is the need of the right and that is what we need. So for all of this that is what we need. So for all of this that is what we need. So for all of this um we like uh kind of are proposing a um we like uh kind of are proposing a um we like uh kind of are proposing a platform which first of all um has a platform which first of all um has a platform which first of all um has a cumulative budget across like the uh
-
cumulative budget across like the uh cumulative budget across like the uh like the attribution runs which happened like the attribution runs which happened like the attribution runs which happened and then where enforcement actually and then where enforcement actually and then where enforcement actually happens in call path rather than um you happens in call path rather than um you happens in call path rather than um you know a separate thing like for example know a separate thing like for example know a separate thing like for example if something goes wrong if your like if if something goes wrong if your like if if something goes wrong if your like if your context is just growing heavily. your context is just growing heavily. your context is just growing heavily. Then like in place compaction should Then like in place compaction should Then like in place compaction should happen or like in place caching or happen or like in place caching or happen or like in place caching or something like that should happen. And something like that should happen. And something like that should happen. And um after that if like after basically um after that if like after basically um after that if like after basically exhausting the list of all in place exhausting the list of all in place exhausting the list of all in place policies only like uh the budget cap policies only like uh the budget cap policies only like uh the budget cap should happen at the very last. Um so should happen at the very last. Um so should happen at the very last. Um so that is something which we are that is something which we are that is something which we are proposing. But um if you look at the proposing. But um if you look at the proposing. But um if you look at the landscape today, if you see the uh like landscape today, if you see the uh like landscape today, if you see the uh like the uh tools like um this light LLM, the uh tools like um this light LLM, the uh tools like um this light LLM, port key, cloudflare, all of those they port key, cloudflare, all of those they port key, cloudflare, all of those they happen at again the request level right happen at again the request level right happen at again the request level right um like if you see like halting is um like if you see like halting is um like if you see like halting is there, routing is there for some of there, routing is there for some of there, routing is there for some of those but all of this again is at a those but all of this again is at a those but all of this again is at a request and you can't control the cost request and you can't control the cost request and you can't control the cost at the uh request layer uh at the model at the uh request layer uh at the model at the uh request layer uh at the model layer, Right.
-
layer, Right. layer, Right. So this is the missing piece which is So this is the missing piece which is So this is the missing piece which is you know the you know the you know the u basically navigating it at the um you u basically navigating it at the um you u basically navigating it at the um you know the model the agent run layer. know the model the agent run layer. know the model the agent run layer. So for that we have token ops which is So for that we have token ops which is So for that we have token ops which is uh you know a runaway token governance uh you know a runaway token governance uh you know a runaway token governance for AI agents and u this is the uh for AI agents and u this is the uh for AI agents and u this is the uh architecture for that. So first of all architecture for that. So first of all architecture for that. So first of all one thing I would want to highlight is one thing I would want to highlight is one thing I would want to highlight is the like the intentional design decision the like the intentional design decision the like the intentional design decision we took here was an out ofbound plane. we took here was an out ofbound plane. we took here was an out ofbound plane. So it doesn't interfere with your code So it doesn't interfere with your code So it doesn't interfere with your code at all. Um so if you see here that out at all. Um so if you see here that out at all. Um so if you see here that out of the bandound plane has three modules of the bandound plane has three modules of the bandound plane has three modules which I'll be talking about. The first which I'll be talking about. The first which I'll be talking about. The first one being instrumentation. It is a one being instrumentation. It is a one being instrumentation. It is a common observability layer where you common observability layer where you common observability layer where you know you'll u like uh have u like the know you'll u like uh have u like the know you'll u like uh have u like the basic telemetry the open telemetry the basic telemetry the open telemetry the basic telemetry the open telemetry the cost in microns and um like the um like cost in microns and um like the um like cost in microns and um like the um like enrichment layer and basically um the uh enrichment layer and basically um the uh enrichment layer and basically um the uh attribution like what caused that uh attribution like what caused that uh attribution like what caused that uh like particular run like particular run like particular run and then there is um obviously and then there is um obviously and then there is um obviously accounting where accounting where accounting where you'll basically accumulate it in a kind you'll basically accumulate it in a kind you'll basically accumulate it in a kind of a ledger like the total runs which of a ledger like the total runs which of a ledger like the total runs which are happening. And finally we have this are happening. And finally we have this are happening. And finally we have this enforced layer which has uh two main enforced layer which has uh two main enforced layer which has uh two main purposes. one is steering it um through
-
purposes. one is steering it um through purposes. one is steering it um through the policies which we've defined which I the policies which we've defined which I the policies which we've defined which I think will cover later and um then we think will cover later and um then we think will cover later and um then we have halt in place as the you know final have halt in place as the you know final have halt in place as the you know final um like u final thing if um you know um like u final thing if um you know um like u final thing if um you know your budget is getting exhausted your budget is getting exhausted your budget is getting exhausted so yeah that is there now when we again so yeah that is there now when we again so yeah that is there now when we again look at the landscape this kind of will look at the landscape this kind of will look at the landscape this kind of will solve a lot of problems solve a lot of problems solve a lot of problems um which kind of happened uh when we um which kind of happened uh when we um which kind of happened uh when we like look at the previous um tools or like look at the previous um tools or like look at the previous um tools or products there because uh it is at products there because uh it is at products there because uh it is at happening at run and it is you know uh happening at run and it is you know uh happening at run and it is you know uh helping you solve the problem from the helping you solve the problem from the helping you solve the problem from the very root by steering it in place. very root by steering it in place. very root by steering it in place. All right. So uh with this I would like All right. So uh with this I would like All right. So uh with this I would like to hand it over to Sashim for the demo. to hand it over to Sashim for the demo. to hand it over to Sashim for the demo. >> Yeah. >> Oh yeah. Now I think I should be able to >> Oh yeah. Now I think I should be able to everyone in the back can hear me. All everyone in the back can hear me. All everyone in the back can hear me. All right, perfect. So yeah, we have right, perfect. So yeah, we have right, perfect. So yeah, we have established the principles behind token established the principles behind token established the principles behind token ops till now. Right. Now let's shift ops till now. Right. Now let's shift ops till now. Right. Now let's shift gears, talk about the design part of it gears, talk about the design part of it gears, talk about the design part of it and uh maybe get into the code and the and uh maybe get into the code and the and uh maybe get into the code and the eventual demo. Right? So what I have eventual demo. Right? So what I have eventual demo. Right? So what I have behind me on the screen is the like behind me on the screen is the like behind me on the screen is the like bird's eye view of what token ops looks bird's eye view of what token ops looks bird's eye view of what token ops looks like today. It's it's three layers.
-
like today. It's it's three layers. like today. It's it's three layers. We'll go left to right and top to We'll go left to right and top to We'll go left to right and top to bottom. So on the left most you have bottom. So on the left most you have bottom. So on the left most you have your own agent runtime which you're your own agent runtime which you're your own agent runtime which you're trying to instrument and kind of manage trying to instrument and kind of manage trying to instrument and kind of manage the cost for right the middle layer is the cost for right the middle layer is the cost for right the middle layer is what we're calling the bridge that what we're calling the bridge that what we're calling the bridge that basically shuffles data between your basically shuffles data between your basically shuffles data between your agent and the control plane and the agent and the control plane and the agent and the control plane and the control plane is where the mind of the control plane is where the mind of the control plane is where the mind of the system lies right so let's talk about system lies right so let's talk about system lies right so let's talk about the bridge layer very briefly if we uh the bridge layer very briefly if we uh the bridge layer very briefly if we uh go from top to bottom you have the go from top to bottom you have the go from top to bottom you have the attribution on top so what we're trying attribution on top so what we're trying attribution on top so what we're trying to do here is every agent run that you to do here is every agent run that you to do here is every agent run that you do it's attributed to some user do it's attributed to some user do it's attributed to some user dimensions so the idea is everything dimensions so the idea is everything dimensions so the idea is everything that you do every run of the agent is that you do every run of the agent is that you do every run of the agent is accounted to some usability or some accounted to some usability or some accounted to some usability or some usage. This comes in handy later. We'll usage. This comes in handy later. We'll usage. This comes in handy later. We'll talk about it. Uh the second part which talk about it. Uh the second part which talk about it. Uh the second part which is the boundary annotation that you see is the boundary annotation that you see is the boundary annotation that you see this is pretty much the heart and soul this is pretty much the heart and soul this is pretty much the heart and soul of this middle layer. So the idea behind of this middle layer. So the idea behind of this middle layer. So the idea behind the boundary annotation is that you take the boundary annotation is that you take the boundary annotation is that you take any method. It doesn't matter what any method. It doesn't matter what any method. It doesn't matter what framework you're using. You might be framework you're using. You might be framework you're using. You might be using uh let's say lang chain lang using uh let's say lang chain lang using uh let's say lang chain lang whatever. If you have a method you can whatever. If you have a method you can whatever. If you have a method you can annotate it with boundary. What this annotate it with boundary. What this annotate it with boundary. What this annotation is going to do is it's going annotation is going to do is it's going annotation is going to do is it's going to do two things. First it's going to to do two things. First it's going to to do two things. First it's going to track the input and the output and it's track the input and the output and it's track the input and the output and it's going to flight that up to the control going to flight that up to the control going to flight that up to the control layer and record it there as a ledger layer and record it there as a ledger layer and record it there as a ledger entry. Now this will be annotated with entry. Now this will be annotated with entry. Now this will be annotated with the further agent run ID and the other the further agent run ID and the other the further agent run ID and the other attributes and so on. The second thing attributes and so on. The second thing attributes and so on. The second thing the boundary annotation does is it acts the boundary annotation does is it acts the boundary annotation does is it acts as a channel through which the control as a channel through which the control as a channel through which the control plane can push actions down to the plane can push actions down to the plane can push actions down to the agent. This is where the entire agent. This is where the entire agent. This is where the entire intelligence lies. So we do not have a intelligence lies. So we do not have a intelligence lies. So we do not have a single directional highway. We want the single directional highway. We want the single directional highway. We want the control plane to be able to tweak the control plane to be able to tweak the control plane to be able to tweak the behavior of the agent on the fly to behavior of the agent on the fly to behavior of the agent on the fly to ensure that we are able to squeeze in ensure that we are able to squeeze in ensure that we are able to squeeze in more runs inside our budget cap. Right
-
more runs inside our budget cap. Right more runs inside our budget cap. Right now let's say the control plane pushes now let's say the control plane pushes now let's say the control plane pushes down an action. Let's take a small down an action. Let's take a small down an action. Let's take a small example. Let's say you have a rag example. Let's say you have a rag example. Let's say you have a rag retrieval tool which is generating like retrieval tool which is generating like retrieval tool which is generating like 20 chunks every retrieval for every call 20 chunks every retrieval for every call 20 chunks every retrieval for every call and that's eating up eating up your and that's eating up eating up your and that's eating up eating up your budget. And let's say the LLM is not budget. And let's say the LLM is not budget. And let's say the LLM is not even using the chunks that are after even using the chunks that are after even using the chunks that are after five because they are just not relevant, five because they are just not relevant, five because they are just not relevant, right? They're sorted by relevance. So right? They're sorted by relevance. So right? They're sorted by relevance. So let's say the control plane observes let's say the control plane observes let's say the control plane observes this and it wants to limit the output to this and it wants to limit the output to this and it wants to limit the output to just five chunks. So it can push down an just five chunks. So it can push down an just five chunks. So it can push down an action but that action has to be action but that action has to be action but that action has to be received by boundary and then has to be received by boundary and then has to be received by boundary and then has to be executed by something. That is where the executed by something. That is where the executed by something. That is where the third node, the governor node comes in. third node, the governor node comes in. third node, the governor node comes in. The governor knows what actions are The governor knows what actions are The governor knows what actions are allowed on your agent by you as a allowed on your agent by you as a allowed on your agent by you as a developer and it receives those actions developer and it receives those actions developer and it receives those actions from the control plane and knows how to from the control plane and knows how to from the control plane and knows how to apply it in a non-destructive way. So apply it in a non-destructive way. So apply it in a non-destructive way. So that's the first three. The fourth one that's the first three. The fourth one that's the first three. The fourth one wrap uh the wrap complete is essentially wrap uh the wrap complete is essentially wrap uh the wrap complete is essentially just a helper method. So as we know most just a helper method. So as we know most just a helper method. So as we know most of the agent providers or the model of the agent providers or the model of the agent providers or the model providers they provide objects rather providers they provide objects rather providers they provide objects rather than methods for their LMS right. So than methods for their LMS right. So than methods for their LMS right. So wrap complete is just another way of wrap complete is just another way of wrap complete is just another way of applying boundary on objects rather than applying boundary on objects rather than applying boundary on objects rather than methods. Let's shift right to the methods. Let's shift right to the methods. Let's shift right to the control plane. On the control plane the control plane. On the control plane the control plane. On the control plane the first layer is the segment. Now this is first layer is the segment. Now this is first layer is the segment. Now this is where the attribution that we talked where the attribution that we talked where the attribution that we talked about earlier comes into picture. So any about earlier comes into picture. So any about earlier comes into picture. So any dimensions that you float from the dimensions that you float from the dimensions that you float from the attribution layer. Let's say you have a attribution layer. Let's say you have a attribution layer. Let's say you have a preview agent that you share with preview agent that you share with preview agent that you share with everyone in this room and your agent is everyone in this room and your agent is everyone in this room and your agent is floating a dimension saying that cohort floating a dimension saying that cohort floating a dimension saying that cohort is AIE 2026 right so you can create a is AIE 2026 right so you can create a is AIE 2026 right so you can create a segment which is a cohort of users which segment which is a cohort of users which segment which is a cohort of users which is based on this tag like dimension is based on this tag like dimension is based on this tag like dimension being AI 2026 right and you can apply being AI 2026 right and you can apply being AI 2026 right and you can apply your budgets at this cohort level so you your budgets at this cohort level so you your budgets at this cohort level so you don't necessarily have to restrict
-
don't necessarily have to restrict don't necessarily have to restrict everything at an agent level or a run everything at an agent level or a run everything at an agent level or a run level you can do you can do rollups you level you can do you can do rollups you level you can do you can do rollups you can do fine grain or coarse grain can do fine grain or coarse grain can do fine grain or coarse grain control right so that's the segmentation control right so that's the segmentation control right so that's the segmentation part of Ledger as I mentioned is just part of Ledger as I mentioned is just part of Ledger as I mentioned is just one agent run all the traces in one one agent run all the traces in one one agent run all the traces in one place. Then you have budgets. Budgets place. Then you have budgets. Budgets place. Then you have budgets. Budgets are basically just the static thresholds are basically just the static thresholds are basically just the static thresholds that work across a time window against a that work across a time window against a that work across a time window against a particular segment or an agent run. And particular segment or an agent run. And particular segment or an agent run. And then you have actions. So on the actions then you have actions. So on the actions then you have actions. So on the actions part we have broadly two flavors. First part we have broadly two flavors. First part we have broadly two flavors. First is the halt type actions which basically is the halt type actions which basically is the halt type actions which basically just kill your agent if it exceeds a just kill your agent if it exceeds a just kill your agent if it exceeds a budget. The second part where we are budget. The second part where we are budget. The second part where we are adding value is the steer type actions. adding value is the steer type actions. adding value is the steer type actions. So here we do not kill the agent. So here we do not kill the agent. So here we do not kill the agent. Instead we try to steer the behavior of Instead we try to steer the behavior of Instead we try to steer the behavior of the agent or the components of the agent the agent or the components of the agent the agent or the components of the agent to try and fit that particular run to try and fit that particular run to try and fit that particular run within the alerted budget. Right? And within the alerted budget. Right? And within the alerted budget. Right? And then the policies layer is where it all then the policies layer is where it all then the policies layer is where it all comes together. You basically uh group comes together. You basically uh group comes together. You basically uh group the budgets the actions and then set the budgets the actions and then set the budgets the actions and then set your policies against certain segments your policies against certain segments your policies against certain segments or agent runs and that is where it or agent runs and that is where it or agent runs and that is where it executes. Right? So moving on uh what executes. Right? So moving on uh what executes. Right? So moving on uh what changes in your code that is the changes in your code that is the changes in your code that is the boundary annotation that we just talked boundary annotation that we just talked boundary annotation that we just talked about. As Disha mentioned earlier this about. As Disha mentioned earlier this about. As Disha mentioned earlier this is all out of band. So you do not have is all out of band. So you do not have is all out of band. So you do not have to change your code. You just have to to change your code. You just have to to change your code. You just have to apply the annotation on the methods that apply the annotation on the methods that apply the annotation on the methods that you have. This boundary annotation will you have. This boundary annotation will you have. This boundary annotation will take care of floating all the take care of floating all the take care of floating all the information up to the control plane. And information up to the control plane. And information up to the control plane. And uh the control plane lies in your own uh the control plane lies in your own uh the control plane lies in your own tenant. So you do not need to worry tenant. So you do not need to worry tenant. So you do not need to worry about any data leaks or anything. Then about any data leaks or anything. Then about any data leaks or anything. Then if I talk about the governor, so for the if I talk about the governor, so for the if I talk about the governor, so for the governor, you just have to create an governor, you just have to create an governor, you just have to create an instance. You just have to pass it your instance. You just have to pass it your instance. You just have to pass it your own configs. These configs will own configs. These configs will own configs. These configs will basically declare what sort of actions basically declare what sort of actions basically declare what sort of actions are allowed for those agents, right? so are allowed for those agents, right? so are allowed for those agents, right? so that your control plane cannot just
-
that your control plane cannot just that your control plane cannot just willingly do any random things on your willingly do any random things on your willingly do any random things on your on your agents. So before we move on to on your agents. So before we move on to on your agents. So before we move on to the demo, I'll just briefly touch upon the demo, I'll just briefly touch upon the demo, I'll just briefly touch upon the uh test that we're going to use the uh test that we're going to use the uh test that we're going to use today. So it's a simple two agent today. So it's a simple two agent today. So it's a simple two agent workflow. We have a research agent which workflow. We have a research agent which workflow. We have a research agent which has access to a search tool. Uh you give has access to a search tool. Uh you give has access to a search tool. Uh you give it a question. It's allowed to look up it a question. It's allowed to look up it a question. It's allowed to look up on the web as many times as it wants. on the web as many times as it wants. on the web as many times as it wants. And once it knows that it has all the And once it knows that it has all the And once it knows that it has all the data, it passes the findings on to the data, it passes the findings on to the data, it passes the findings on to the second agent which is a summarizer which second agent which is a summarizer which second agent which is a summarizer which creates creates a research report. creates creates a research report. creates creates a research report. Right? So with that out of the way, Right? So with that out of the way, Right? So with that out of the way, let's just quickly walk over to the let's just quickly walk over to the let's just quickly walk over to the demo. So for the demo, we have three demo. So for the demo, we have three demo. So for the demo, we have three different scenarios that we're going to different scenarios that we're going to different scenarios that we're going to talk about. For the first one, we're talk about. For the first one, we're talk about. For the first one, we're going to run the token ops in what we going to run the token ops in what we going to run the token ops in what we call preview mode. So in preview mode, call preview mode. So in preview mode, call preview mode. So in preview mode, what happens is that all the policies what happens is that all the policies what happens is that all the policies run as is, but the enforcement doesn't run as is, but the enforcement doesn't run as is, but the enforcement doesn't happen. So if you see we ran a happen. So if you see we ran a happen. So if you see we ran a particular run over here which completed particular run over here which completed particular run over here which completed but we did not see any sort of failures but we did not see any sort of failures but we did not see any sort of failures there. The policies executed but the there. The policies executed but the there. The policies executed but the actions that were associated with those actions that were associated with those actions that were associated with those policies were not allowed to be policies were not allowed to be policies were not allowed to be executed. So we're just going to load executed. So we're just going to load executed. So we're just going to load the dashboard screen here.
-
the dashboard screen here. the dashboard screen here. Yeah. So this is the governance output. Yeah. So this is the governance output. Yeah. So this is the governance output. Governance is off. The run completed. Governance is off. The run completed. Governance is off. The run completed. But in the dashboard you can see the But in the dashboard you can see the But in the dashboard you can see the policies have executed. So you can see policies have executed. So you can see policies have executed. So you can see the cost budget, the cost guard and so the cost budget, the cost guard and so the cost budget, the cost guard and so on. Right? So this was the first on. Right? So this was the first on. Right? So this was the first scenario. For the second scenario, what scenario. For the second scenario, what scenario. For the second scenario, what we're going to do is we're going to turn we're going to do is we're going to turn we're going to do is we're going to turn on the governance. Now while that is on the governance. Now while that is on the governance. Now while that is happening, I just want to touch upon why happening, I just want to touch upon why happening, I just want to touch upon why this is important. So if you want to this is important. So if you want to this is important. So if you want to like include this product into your like include this product into your like include this product into your production agents, you want to have a production agents, you want to have a production agents, you want to have a safe environment or a safe way to safe environment or a safe way to safe environment or a safe way to firstly put it in your production firstly put it in your production firstly put it in your production environment, test the guardrails, tweak environment, test the guardrails, tweak environment, test the guardrails, tweak the guardrail, see what's the policies the guardrail, see what's the policies the guardrail, see what's the policies are doing and then finalize the are doing and then finalize the are doing and then finalize the thresholds. Right? So this is the second thresholds. Right? So this is the second thresholds. Right? So this is the second one where we have now enforced the one where we have now enforced the one where we have now enforced the governance and you can see in the governance and you can see in the governance and you can see in the dashboard that the pre-all cost cap has dashboard that the pre-all cost cap has dashboard that the pre-all cost cap has exceeded. So you had a budget allotted exceeded. So you had a budget allotted exceeded. So you had a budget allotted for this run but the agent exceeded the for this run but the agent exceeded the for this run but the agent exceeded the budget and it was killed immediately. So budget and it was killed immediately. So budget and it was killed immediately. So that's the simple circuit breaker sort that's the simple circuit breaker sort that's the simple circuit breaker sort of a methodology. So this is the halt of a methodology. So this is the halt of a methodology. So this is the halt behavior. And now let's see the steer behavior. And now let's see the steer behavior. And now let's see the steer behavior which is the which is where we behavior which is the which is where we behavior which is the which is where we are trying to add value to this entire are trying to add value to this entire are trying to add value to this entire cost management scenario. So this time cost management scenario. So this time cost management scenario. So this time we're going to run the third the second we're going to run the third the second we're going to run the third the second prompt. The budget allotted for this one prompt. The budget allotted for this one prompt. The budget allotted for this one is slightly higher but it's still not is slightly higher but it's still not is slightly higher but it's still not high enough for the agent to complete in high enough for the agent to complete in high enough for the agent to complete in time. So what instead happens is there time. So what instead happens is there time. So what instead happens is there is something called cost guard which is something called cost guard which is something called cost guard which kicks in. This cost guard it takes into kicks in. This cost guard it takes into kicks in. This cost guard it takes into account two things. First how much of account two things. First how much of account two things. First how much of your allotted budget have you consumed?
-
your allotted budget have you consumed? your allotted budget have you consumed? Second what is the velocity at which Second what is the velocity at which Second what is the velocity at which you're consuming tokens. [music] Now you're consuming tokens. [music] Now you're consuming tokens. [music] Now based on these two things if it predicts based on these two things if it predicts based on these two things if it predicts that you're going to run out of your that you're going to run out of your that you're going to run out of your tokens or your allotted budget by the tokens or your allotted budget by the tokens or your allotted budget by the end of the run it's going to inject end of the run it's going to inject end of the run it's going to inject something into your system instructions something into your system instructions something into your system instructions that something could be as simple as hey that something could be as simple as hey that something could be as simple as hey you're running out of budget so make you're running out of budget so make you're running out of budget so make sure that the LM outputs are more sure that the LM outputs are more sure that the LM outputs are more succinct or more summarized right so succinct or more summarized right so succinct or more summarized right so that is the way we are doing the that is the way we are doing the that is the way we are doing the steering now the this was a very simple steering now the this was a very simple steering now the this was a very simple test bench to show you like how this test bench to show you like how this test bench to show you like how this works on a like working code we have works on a like working code we have works on a like working code we have also benchmarked it on a couple of open also benchmarked it on a couple of open also benchmarked it on a couple of open source repos. So we have benchmarked it source repos. So we have benchmarked it source repos. So we have benchmarked it on browser use as well as metagp. Uh we on browser use as well as metagp. Uh we on browser use as well as metagp. Uh we ran it across multiple iterations across ran it across multiple iterations across ran it across multiple iterations across stress tests across simple scenarios stress tests across simple scenarios stress tests across simple scenarios hard scenarios and everything. And the hard scenarios and everything. And the hard scenarios and everything. And the results we see are the average spend results we see are the average spend results we see are the average spend goes down by almost 78% with token ops goes down by almost 78% with token ops goes down by almost 78% with token ops enabled with the full policy suit that enabled with the full policy suit that enabled with the full policy suit that we have today. On the completion part we have today. On the completion part we have today. On the completion part when we compare it with throttling just when we compare it with throttling just when we compare it with throttling just simple throttling your simple throttling simple throttling your simple throttling simple throttling your simple throttling is going to kill your agent runs no is going to kill your agent runs no is going to kill your agent runs no matter what. Right? So with the reduced matter what. Right? So with the reduced matter what. Right? So with the reduced average spend what you get is you get an average spend what you get is you get an average spend what you get is you get an uplift in that completion percentage uplift in that completion percentage uplift in that completion percentage from 67% to roughly 96%. So that is the from 67% to roughly 96%. So that is the from 67% to roughly 96%. So that is the value ad that token ops is doing here.
-
value ad that token ops is doing here. value ad that token ops is doing here. Now this is the policy catalog that we Now this is the policy catalog that we Now this is the policy catalog that we run this benchmark against. This is what run this benchmark against. This is what run this benchmark against. This is what we support today. We kind of researched we support today. We kind of researched we support today. We kind of researched what are the different failure modes what are the different failure modes what are the different failure modes that are there today out in the wild and that are there today out in the wild and that are there today out in the wild and tried to cover most of them here. So you tried to cover most of them here. So you tried to cover most of them here. So you have things across spend management, you have things across spend management, you have things across spend management, you have things across context management have things across context management have things across context management like context compaction, tool output like context compaction, tool output like context compaction, tool output reduction, you have things across loop reduction, you have things across loop reduction, you have things across loop detection and progress detection and detection and progress detection and detection and progress detection and stuff like that. So this is the entire stuff like that. So this is the entire stuff like that. So this is the entire set of policies that we support. And at set of policies that we support. And at set of policies that we support. And at the bottom you can see the actions. So the bottom you can see the actions. So the bottom you can see the actions. So as I mentioned earlier, we have two as I mentioned earlier, we have two as I mentioned earlier, we have two flavors. You have the uh the halt type flavors. You have the uh the halt type flavors. You have the uh the halt type actions and then the steer type actions. actions and then the steer type actions. actions and then the steer type actions. So for the steer we can do allow, So for the steer we can do allow, So for the steer we can do allow, mutate, inject and so on. And for the mutate, inject and so on. And for the mutate, inject and so on. And for the halt, it can be a simple kill. But this halt, it can be a simple kill. But this halt, it can be a simple kill. But this is not the end state that we envision is not the end state that we envision is not the end state that we envision for this. The end state is we have a lot for this. The end state is we have a lot for this. The end state is we have a lot of data right we have a ledger that is of data right we have a ledger that is of data right we have a ledger that is continuously being updated. So what we continuously being updated. So what we continuously being updated. So what we want to try is we want to try a want to try is we want to try a want to try is we want to try a self-learning module within the token self-learning module within the token self-learning module within the token ops plane within the control plane which ops plane within the control plane which ops plane within the control plane which can look at this ledger and ask this can look at this ledger and ask this can look at this ledger and ask this question hey why or what is the failure question hey why or what is the failure question hey why or what is the failure mode that I'm still not able to catch mode that I'm still not able to catch mode that I'm still not able to catch and then based on that it can do two and then based on that it can do two and then based on that it can do two things one is it can enhance it can things one is it can enhance it can things one is it can enhance it can generate new policies on the fly based generate new policies on the fly based generate new policies on the fly based on the missing or the still uh runaway on the missing or the still uh runaway on the missing or the still uh runaway costs or it can refine the existing costs or it can refine the existing costs or it can refine the existing parameters for the existing policies parameters for the existing policies parameters for the existing policies that are there so that the runaway costs that are there so that the runaway costs that are there so that the runaway costs are managed more effectively in the are managed more effectively in the are managed more effectively in the future. So with that I think uh that is future. So with that I think uh that is future. So with that I think uh that is all we have for you guys today. Thank all we have for you guys today. Thank all we have for you guys today. Thank you so much for your time and you can you so much for your time and you can you so much for your time and you can scan this QR code that's the public scan this QR code that's the public scan this QR code that's the public wiki. We are updating it almost wiki. We are updating it almost wiki. We are updating it almost regularly. So you can scan this and stay regularly. So you can scan this and stay regularly. So you can scan this and stay up to date and uh Tisha and I are around
-
up to date and uh Tisha and I are around up to date and uh Tisha and I are around so if you guys have any questions or if so if you guys have any questions or if so if you guys have any questions or if you want to discuss more about it just you want to discuss more about it just you want to discuss more about it just let us know. That's it. Thank you. let us know. That's it. Thank you. let us know. That's it. Thank you. [applause]
Summary
The main theme of the tech transcript is the shift from "token maxing" to "value maxing" in AI. It references past software evolution eras like SaaS and cloud computing to illustrate how control and costing models have changed, concluding that the current agentic era lacks a proper control plane for model calls, leading to unbounded consumption.