← Back
AI Engineer August 29, 2026 23m

From Tokenmaxxing to Trusted Throughput — Mingsheng Hong, Ironclad

Read full transcript 17 segments
  1. All right, let's get started. Apologies All right, let's get started. Apologies for the delay, but I'm really excited to for the delay, but I'm really excited to for the delay, but I'm really excited to be here. I'm Mingshan, VP of engineering be here. I'm Mingshan, VP of engineering be here. I'm Mingshan, VP of engineering focused on AI at Ironclad. And today focused on AI at Ironclad. And today focused on AI at Ironclad. And today I'll be telling you about something I'll be telling you about something I'll be telling you about something that's probably on top of many of your that's probably on top of many of your that's probably on top of many of your mind. uh how to control and optimize for mind. uh how to control and optimize for mind. uh how to control and optimize for your AI token spend. Can I get a get a your AI token spend. Can I get a get a your AI token spend. Can I get a get a quick show of hands that this is a quick show of hands that this is a quick show of hands that this is a relevant topic? relevant topic? relevant topic? Okay, awesome. I appreciate that. So, we have all heard a few sensational So, we have all heard a few sensational stories from the media. There's an stories from the media. There's an stories from the media. There's an interesting Amazon story where an interesting Amazon story where an interesting Amazon story where an employee just created kind of a employee just created kind of a employee just created kind of a voluntary dashboard and everyone start voluntary dashboard and everyone start voluntary dashboard and everyone start tracking their own AI token usage. I'm tracking their own AI token usage. I'm tracking their own AI token usage. I'm not sure there's explicit encouragement not sure there's explicit encouragement not sure there's explicit encouragement from the leadership, but the effect is from the leadership, but the effect is from the leadership, but the effect is you know engineers some of the engineers you know engineers some of the engineers you know engineers some of the engineers started competing with each other in started competing with each other in started competing with each other in maximizing their token usage and get to maximizing their token usage and get to maximizing their token usage and get to the top of the so-called leaderboard. the top of the so-called leaderboard. the top of the so-called leaderboard. There's a similar story from Meta and There's a similar story from Meta and There's a similar story from Meta and then another even more sensational story then another even more sensational story then another even more sensational story about some companies spending $500 about some companies spending $500 about some companies spending $500 million on cloud oops within a month. So million on cloud oops within a month. So million on cloud oops within a month. So while these may not be happening in your while these may not be happening in your while these may not be happening in your companies today, the threats, the risks companies today, the threats, the risks companies today, the threats, the risks are real. How do we think about the are real. How do we think about the are real. How do we think about the policies? How do we measure the cost?

  2. policies? How do we measure the cost? policies? How do we measure the cost? And how do we control and optimize for And how do we control and optimize for And how do we control and optimize for it? it? it? So one initial learning I want to share So one initial learning I want to share So one initial learning I want to share is it is really important to have is it is really important to have is it is really important to have dashboard that track every team every dashboard that track every team every dashboard that track every team every individual's token usage and cost but individual's token usage and cost but individual's token usage and cost but that should not be positioned as a that should not be positioned as a that should not be positioned as a leaderboard. We think of the the usage leaderboard. We think of the the usage leaderboard. We think of the the usage dashboard more as a smoke detector. If dashboard more as a smoke detector. If dashboard more as a smoke detector. If there are local pockets of teams or there are local pockets of teams or there are local pockets of teams or individuals that don't use much AI token individuals that don't use much AI token individuals that don't use much AI token that might be a signal worth that might be a signal worth that might be a signal worth investigating. But beyond that certainly investigating. But beyond that certainly investigating. But beyond that certainly we don't want to create even indirect we don't want to create even indirect we don't want to create even indirect incentive to maximize the token usage incentive to maximize the token usage incentive to maximize the token usage itself. So how do we think about it then? First So how do we think about it then? First I want to make sure that we position I want to make sure that we position I want to make sure that we position this talk for those of you whose teams this talk for those of you whose teams this talk for those of you whose teams have already gone through the hump of have already gone through the hump of have already gone through the hump of getting AI adopted. If you're still in getting AI adopted. If you're still in getting AI adopted. If you're still in the initial process of provisioning easy the initial process of provisioning easy the initial process of provisioning easy access to your engineers or encouraging access to your engineers or encouraging access to your engineers or encouraging the teams and individuals to adopt, then the teams and individuals to adopt, then the teams and individuals to adopt, then you may not be ready to implement some you may not be ready to implement some you may not be ready to implement some of the ideas for controlling and of the ideas for controlling and of the ideas for controlling and optimizing for cost. But that's okay.

  3. optimizing for cost. But that's okay. optimizing for cost. But that's okay. This could still be a good discussion. This could still be a good discussion. This could still be a good discussion. And frankly, we just got over that hump And frankly, we just got over that hump And frankly, we just got over that hump over the last couple quarters. So this over the last couple quarters. So this over the last couple quarters. So this is a very topical subject that every is a very topical subject that every is a very topical subject that every engineering leader I believe is engineering leader I believe is engineering leader I believe is navigating. So I would love to start navigating. So I would love to start navigating. So I would love to start that dialogue with you all today to that dialogue with you all today to that dialogue with you all today to explore the best practices. Can I get a explore the best practices. Can I get a explore the best practices. Can I get a quick show of hand for those of you quick show of hand for those of you quick show of hand for those of you whose teams have gone over the initial whose teams have gone over the initial whose teams have gone over the initial adoption phase now you are starting to adoption phase now you are starting to adoption phase now you are starting to seriously worry about the cost. Okay, I seriously worry about the cost. Okay, I seriously worry about the cost. Okay, I see roughly half of the hands raised. see roughly half of the hands raised. see roughly half of the hands raised. Thank you. So let's talk about then how Thank you. So let's talk about then how Thank you. So let's talk about then how we can control and how we can optimize we can control and how we can optimize we can control and how we can optimize what we call the trusted throughput as a what we call the trusted throughput as a what we call the trusted throughput as a kind of a proxy metric as a way to kind of a proxy metric as a way to kind of a proxy metric as a way to measure your ROI. But before that, just measure your ROI. But before that, just measure your ROI. But before that, just for those of you who are in the process for those of you who are in the process for those of you who are in the process of still increasing adoption, one lesson of still increasing adoption, one lesson of still increasing adoption, one lesson we learned is to after the kind of the we learned is to after the kind of the we learned is to after the kind of the top down leadership push is to sit down top down leadership push is to sit down top down leadership push is to sit down with the individual teams and uh the with the individual teams and uh the with the individual teams and uh the individuals who may be resistant or individuals who may be resistant or individuals who may be resistant or struggling with adoption, understand struggling with adoption, understand struggling with adoption, understand where they came from. For example, there where they came from. For example, there where they came from. For example, there are some legitimate concerns that I are some legitimate concerns that I are some legitimate concerns that I heard, you know, people say, "Hey, I heard, you know, people say, "Hey, I heard, you know, people say, "Hey, I used to really take pride and joy in used to really take pride and joy in used to really take pride and joy in handcrafting the code and now a lot of handcrafting the code and now a lot of handcrafting the code and now a lot of the joy and the pride got taken away and the joy and the pride got taken away and the joy and the pride got taken away and replaced with me reviewing AI slop code, replaced with me reviewing AI slop code, replaced with me reviewing AI slop code, right? So that doesn't sound like a very right? So that doesn't sound like a very right? So that doesn't sound like a very satisfying professional activity and satisfying professional activity and satisfying professional activity and that's where we need to kind of dig down that's where we need to kind of dig down that's where we need to kind of dig down and understand what are still the kind and understand what are still the kind and understand what are still the kind of the high impact and uh engineering of the high impact and uh engineering of the high impact and uh engineering tasks technical work that we can help

  4. tasks technical work that we can help tasks technical work that we can help our engineers continue to grow our engineers continue to grow our engineers continue to grow themselves in the era of the AI. So I wanted to share with you a bit more So I wanted to share with you a bit more about what we at ironclad does and about what we at ironclad does and about what we at ironclad does and there's an interesting connection there's an interesting connection there's an interesting connection actually within how we think about actually within how we think about actually within how we think about optimizing for engineering AI token optimizing for engineering AI token optimizing for engineering AI token usage. So, ironclad is a legal usage. So, ironclad is a legal usage. So, ironclad is a legal contracting AI companies AI company. We contracting AI companies AI company. We contracting AI companies AI company. We build AI features and native AI products build AI features and native AI products build AI features and native AI products to help lawyers, procurement and other to help lawyers, procurement and other to help lawyers, procurement and other business users move forward new business users move forward new business users move forward new contracts, move them forward faster with contracts, move them forward faster with contracts, move them forward faster with controlled risk. What that means is controlled risk. What that means is controlled risk. What that means is building trust is the number one building trust is the number one building trust is the number one priority with our AI product features priority with our AI product features priority with our AI product features and products. And for the prior speak uh and products. And for the prior speak uh and products. And for the prior speak uh speaker speaker, she did a wonderful job speaker speaker, she did a wonderful job speaker speaker, she did a wonderful job telling you about the importance of telling you about the importance of telling you about the importance of trust and how to build it in their trust and how to build it in their trust and how to build it in their domain. In our ironclad product domain, domain. In our ironclad product domain, domain. In our ironclad product domain, it often means lawyers especially, but it often means lawyers especially, but it often means lawyers especially, but other persona as well taking the time to other persona as well taking the time to other persona as well taking the time to kind of test the water and see if they kind of test the water and see if they kind of test the water and see if they can trust the AI output. For example, can trust the AI output. For example, can trust the AI output. For example, they may feed our conversational search they may feed our conversational search they may feed our conversational search a set of contracts they are firmly a set of contracts they are firmly a set of contracts they are firmly familiar with and they run a search and familiar with and they run a search and familiar with and they run a search and see if the output is towards the see if the output is towards the see if the output is towards the expectation. If so, they may expand on expectation. If so, they may expand on expectation. If so, they may expand on searching for things they don't know searching for things they don't know searching for things they don't know about or apply other workflows using AI about or apply other workflows using AI about or apply other workflows using AI to solve other things like redlinining to solve other things like redlinining to solve other things like redlinining the contract um and you know finding the contract um and you know finding the contract um and you know finding anomalies and so on. And so similarly anomalies and so on. And so similarly anomalies and so on. And so similarly using AI and making sure AI is

  5. using AI and making sure AI is using AI and making sure AI is delivering high engineering value also delivering high engineering value also delivering high engineering value also involves a you know a sequence of steps involves a you know a sequence of steps involves a you know a sequence of steps in gaining trust from the internal in gaining trust from the internal in gaining trust from the internal engineers the leadership as well with as engineers the leadership as well with as engineers the leadership as well with as uh with our customers. So this is the uh with our customers. So this is the uh with our customers. So this is the focus of our talk today and this probably will not come as a and this probably will not come as a surprise here. The goal is not to surprise here. The goal is not to surprise here. The goal is not to minimizing or not even necessarily to minimizing or not even necessarily to minimizing or not even necessarily to reduce token spend. So here we kind of reduce token spend. So here we kind of reduce token spend. So here we kind of use the word it's not about austerity. use the word it's not about austerity. use the word it's not about austerity. It's about further improving the ROI of It's about further improving the ROI of It's about further improving the ROI of the token spend. the token spend. the token spend. So how do we do that? Here we propose um So how do we do that? Here we propose um So how do we do that? Here we propose um a concept we call trusted throughput. So a concept we call trusted throughput. So a concept we call trusted throughput. So the trusted throughput comes from having the trusted throughput comes from having the trusted throughput comes from having the code reviewed and validated the code reviewed and validated the code reviewed and validated internally and ultimately validated in internally and ultimately validated in internally and ultimately validated in customer uh in customer deployments. So how do we go and how do we think So how do we go and how do we think about controlling the cost and uh about controlling the cost and uh about controlling the cost and uh measuring and in turn optimizing the measuring and in turn optimizing the measuring and in turn optimizing the ROI? The first step is I'm pretty ROI? The first step is I'm pretty ROI? The first step is I'm pretty confident that all of you your teams who confident that all of you your teams who confident that all of you your teams who have been adopting AI have been have been adopting AI have been have been adopting AI have been measuring the cost. If you're using a measuring the cost. If you're using a measuring the cost. If you're using a single tool like claw code or codeex single tool like claw code or codeex single tool like claw code or codeex then you tend to get very rich analytics then you tend to get very rich analytics then you tend to get very rich analytics from the vendor's dashboard already. If from the vendor's dashboard already. If from the vendor's dashboard already. If you're like us who use a combination of you're like us who use a combination of you're like us who use a combination of these different coding tools then we these different coding tools then we these different coding tools then we basically use AI to build simple basically use AI to build simple basically use AI to build simple dashboards and pipelines to extract such

  6. dashboards and pipelines to extract such dashboards and pipelines to extract such vendor data. So we can kind of vendor data. So we can kind of vendor data. So we can kind of crossorrelate them. Then we can break it crossorrelate them. Then we can break it crossorrelate them. Then we can break it down, aggregate and then break down by down, aggregate and then break down by down, aggregate and then break down by per team, per individual, what is their per team, per individual, what is their per team, per individual, what is their cost usage across all of these uh tools. cost usage across all of these uh tools. cost usage across all of these uh tools. So that's the first step for measuring So that's the first step for measuring So that's the first step for measuring cost. Now one pitfall I have seen and we cost. Now one pitfall I have seen and we cost. Now one pitfall I have seen and we wanted to caution everybody is to then wanted to caution everybody is to then wanted to caution everybody is to then jump from measuring cost to start jump from measuring cost to start jump from measuring cost to start reducing or minimizing the cost, right? reducing or minimizing the cost, right? reducing or minimizing the cost, right? Cutting cost. We think that is Cutting cost. We think that is Cutting cost. We think that is premature. Instead, the other important premature. Instead, the other important premature. Instead, the other important side of the equation for ROI is to side of the equation for ROI is to side of the equation for ROI is to measure value. How much value are we measure value. How much value are we measure value. How much value are we getting from burning the tokens? Once we getting from burning the tokens? Once we getting from burning the tokens? Once we can measure the cost and value side, we can measure the cost and value side, we can measure the cost and value side, we understand ROI and then to improve ROI, understand ROI and then to improve ROI, understand ROI and then to improve ROI, we want to find and then fix the we want to find and then fix the we want to find and then fix the bottlenecks. In the next couple slides, bottlenecks. In the next couple slides, bottlenecks. In the next couple slides, I'm going to introduce two new I'm going to introduce two new I'm going to introduce two new bottlenecks we identify in this whole bottlenecks we identify in this whole bottlenecks we identify in this whole new software development life cycle new software development life cycle new software development life cycle where code generation now becomes where code generation now becomes where code generation now becomes abundant thanks to AI. But the pressure abundant thanks to AI. But the pressure abundant thanks to AI. But the pressure is now getting pushed down to code is now getting pushed down to code is now getting pushed down to code review and continuous integration CICD review and continuous integration CICD review and continuous integration CICD the merging the code. So we'll talk the merging the code. So we'll talk the merging the code. So we'll talk about that and finally we'll put about that and finally we'll put about that and finally we'll put together these ideas into a pra together these ideas into a pra together these ideas into a pra pragmatic framework of how we think pragmatic framework of how we think pragmatic framework of how we think about optimizing the ROI and thus the about optimizing the ROI and thus the about optimizing the ROI and thus the leverage in using AI.

  7. Okay. So this is kind of just a slide in Okay. So this is kind of just a slide in building or using the vendor dashboard building or using the vendor dashboard building or using the vendor dashboard to measure the cost. And again we want to measure the cost. And again we want to measure the cost. And again we want to caution that here the main goal for to caution that here the main goal for to caution that here the main goal for regularly reviewing the dashboard is to regularly reviewing the dashboard is to regularly reviewing the dashboard is to see a if there's still adoption gap see a if there's still adoption gap see a if there's still adoption gap within individual pockets of teams or within individual pockets of teams or within individual pockets of teams or the individual engineers and b if there the individual engineers and b if there the individual engineers and b if there are any sudden surprises in kind of the are any sudden surprises in kind of the are any sudden surprises in kind of the usage burst and if so understand what's usage burst and if so understand what's usage burst and if so understand what's been happening if they're legitimate and been happening if they're legitimate and been happening if they're legitimate and then also compare teams then also compare teams then also compare teams contextually. So this is important. We contextually. So this is important. We contextually. So this is important. We don't control just the AI usage per se don't control just the AI usage per se don't control just the AI usage per se because for example a platform because for example a platform because for example a platform infrastructure team the way they use AI infrastructure team the way they use AI infrastructure team the way they use AI and the way they get value may be and the way they get value may be and the way they get value may be different from the UI team. So we need different from the UI team. So we need different from the UI team. So we need to take the context into consideration. to take the context into consideration. to take the context into consideration. All of such review analysis is to help All of such review analysis is to help All of such review analysis is to help us extract learnings. So there's a us extract learnings. So there's a us extract learnings. So there's a self-learning loop that we can then feed self-learning loop that we can then feed self-learning loop that we can then feed back into institutional best practices. back into institutional best practices. back into institutional best practices. What we don't want to use the dashboards What we don't want to use the dashboards What we don't want to use the dashboards are to kind of stack rank people, right? are to kind of stack rank people, right? are to kind of stack rank people, right? making it a a leaderboard and somehow making it a a leaderboard and somehow making it a a leaderboard and somehow reward maximization.

  8. reward maximization. reward maximization. There's an interesting analogy I want to There's an interesting analogy I want to There's an interesting analogy I want to draw with uh a traditional edge draw with uh a traditional edge draw with uh a traditional edge productivity metric called lines of productivity metric called lines of productivity metric called lines of code. So I believe all of you will be code. So I believe all of you will be code. So I believe all of you will be tracking that metric but it wouldn't be tracking that metric but it wouldn't be tracking that metric but it wouldn't be wise to use that metric as the key goal wise to use that metric as the key goal wise to use that metric as the key goal to measure engine velocity because if we to measure engine velocity because if we to measure engine velocity because if we want productive and high quality engine want productive and high quality engine want productive and high quality engine work one can argue that removing code is work one can argue that removing code is work one can argue that removing code is even better. So, LOC line of code is an even better. So, LOC line of code is an even better. So, LOC line of code is an important metric but not something we important metric but not something we important metric but not something we want to directly optimize for. Same want to directly optimize for. Same want to directly optimize for. Same thing for the token usage and spend. thing for the token usage and spend. thing for the token usage and spend. So, that that gets us to the notion of So, that that gets us to the notion of So, that that gets us to the notion of trusted throughput. How do we think trusted throughput. How do we think trusted throughput. How do we think about that? How do we define that? about that? How do we define that? about that? How do we define that? First, I want to kind of share the First, I want to kind of share the First, I want to kind of share the quantified uh side of the things. What quantified uh side of the things. What quantified uh side of the things. What are the metrics that kind of we have are the metrics that kind of we have are the metrics that kind of we have been involving in defining and tracking. been involving in defining and tracking. been involving in defining and tracking. So we talked about line of code is So we talked about line of code is So we talked about line of code is clearly not a good way to measure if AI clearly not a good way to measure if AI clearly not a good way to measure if AI is you know generating a lot of value. is you know generating a lot of value. is you know generating a lot of value. So the next evolution can be let's count So the next evolution can be let's count So the next evolution can be let's count the number of open PRs pull requests.

  9. the number of open PRs pull requests. the number of open PRs pull requests. The intuition being engineers are using The intuition being engineers are using The intuition being engineers are using AI to generate a lot more code. So let's AI to generate a lot more code. So let's AI to generate a lot more code. So let's measure the open PR. So clearly we see a measure the open PR. So clearly we see a measure the open PR. So clearly we see a big kind of inflection in the open PR big kind of inflection in the open PR big kind of inflection in the open PR count. But eventually as we as I assume count. But eventually as we as I assume count. But eventually as we as I assume everyone would agree over the time even everyone would agree over the time even everyone would agree over the time even though people may do oneoff you know R&D though people may do oneoff you know R&D though people may do oneoff you know R&D work to try out things without lending work to try out things without lending work to try out things without lending them but eventually we're all measured them but eventually we're all measured them but eventually we're all measured by the code we ship. So therefore we by the code we ship. So therefore we by the code we ship. So therefore we evolved from tracking the open PR count evolved from tracking the open PR count evolved from tracking the open PR count to tracking the merge PR count. So to tracking the merge PR count. So to tracking the merge PR count. So that's an improvement. that's an improvement. that's an improvement. But the next question is not every But the next question is not every But the next question is not every merged PR is equal. There can be a PO merged PR is equal. There can be a PO merged PR is equal. There can be a PO with only 10 lines of code that takes with only 10 lines of code that takes with only 10 lines of code that takes forever that finds and fix a concurrency forever that finds and fix a concurrency forever that finds and fix a concurrency bug or there can be a thousand line kind bug or there can be a thousand line kind bug or there can be a thousand line kind of boilerplate code that just takes a of boilerplate code that just takes a of boilerplate code that just takes a lot of time to then kind of generate and lot of time to then kind of generate and lot of time to then kind of generate and review but otherwise it's not necessary review but otherwise it's not necessary review but otherwise it's not necessary adding as much business value. adding as much business value. adding as much business value. So as such we then started kind of So as such we then started kind of So as such we then started kind of tagging each merged PR with some sort of tagging each merged PR with some sort of tagging each merged PR with some sort of complexity score. There's no traditional complexity score. There's no traditional complexity score. There's no traditional definition of what that means. We looked definition of what that means. We looked definition of what that means. We looked at the literature a bit. So we just took at the literature a bit. So we just took at the literature a bit. So we just took a pragmatic approach of giving AI a a pragmatic approach of giving AI a a pragmatic approach of giving AI a well-crafted prompt and then we feed the well-crafted prompt and then we feed the well-crafted prompt and then we feed the PR into basically one or two M and say PR into basically one or two M and say PR into basically one or two M and say score the complexity based on t-shirt score the complexity based on t-shirt score the complexity based on t-shirt size. So I the idea being if you use AI size. So I the idea being if you use AI size. So I the idea being if you use AI to generate a more complex PR we to generate a more complex PR we to generate a more complex PR we consider that as being more valuable consider that as being more valuable consider that as being more valuable basically that's how we kind of add a basically that's how we kind of add a basically that's how we kind of add a weightage to each merged PR but that's weightage to each merged PR but that's weightage to each merged PR but that's not the end of the journey that's still

  10. not the end of the journey that's still not the end of the journey that's still something we're going to evolve keep something we're going to evolve keep something we're going to evolve keep evolving and I would love to discuss evolving and I would love to discuss evolving and I would love to discuss with everyone on kind of how we end up with everyone on kind of how we end up with everyone on kind of how we end up creating defining a set of metrics that creating defining a set of metrics that creating defining a set of metrics that kind of approximate the value AI is kind of approximate the value AI is kind of approximate the value AI is generating. generating. generating. Now let's look at the qualitative view. Now let's look at the qualitative view. Now let's look at the qualitative view. What we think about the way we would What we think about the way we would What we think about the way we would define trusted throughput is a high define trusted throughput is a high define trusted throughput is a high quality output that's interested by both quality output that's interested by both quality output that's interested by both internal engineering and leadership and internal engineering and leadership and internal engineering and leadership and external customers. We think they come external customers. We think they come external customers. We think they come from three buckets. from three buckets. from three buckets. The first bucket is all of the objective The first bucket is all of the objective The first bucket is all of the objective metrics that we run with checking the metrics that we run with checking the metrics that we run with checking the test coverage whether uh all of the test coverage whether uh all of the test coverage whether uh all of the predefined security checks are passing. predefined security checks are passing. predefined security checks are passing. Do we go through the regular canarying Do we go through the regular canarying Do we go through the regular canarying practice as we roll out features safely practice as we roll out features safely practice as we roll out features safely and so on. In addition, we complement and so on. In addition, we complement and so on. In addition, we complement the subjective objective metrics with the subjective objective metrics with the subjective objective metrics with our subjective human judgment. So that's our subjective human judgment. So that's our subjective human judgment. So that's where the code review, the design review where the code review, the design review where the code review, the design review come in to look at the code quality, come in to look at the code quality, come in to look at the code quality, clarity, maintenance, architecture fit clarity, maintenance, architecture fit clarity, maintenance, architecture fit and so on. And then finally we want to and so on. And then finally we want to and so on. And then finally we want to make sure through all of these internal make sure through all of these internal make sure through all of these internal objective and subjective check when the objective and subjective check when the objective and subjective check when the rubber meets the road how customer rubber meets the road how customer rubber meets the road how customer perceive the changes are there perceive the changes are there perceive the changes are there production fire that lead to ro production fire that lead to ro production fire that lead to ro rollbacks do customers complain have rollbacks do customers complain have rollbacks do customers complain have tickets that talk about usability uh tickets that talk about usability uh tickets that talk about usability uh friction uh bugs and so on. So these are friction uh bugs and so on. So these are friction uh bugs and so on. So these are the three buckets that together form the three buckets that together form the three buckets that together form what we think is trusted throughput from what we think is trusted throughput from what we think is trusted throughput from engineering.

  11. Okay. So now let's talk about from a Okay. So now let's talk about from a software deploy deployment life cycle software deploy deployment life cycle software deploy deployment life cycle perspective where we observe the new perspective where we observe the new perspective where we observe the new bottlenecks are as I mentioned earlier bottlenecks are as I mentioned earlier bottlenecks are as I mentioned earlier AI code generation is making PR creation AI code generation is making PR creation AI code generation is making PR creation abundant. So now the the bottleneck from abundant. So now the the bottleneck from abundant. So now the the bottleneck from kind of the whole life cycle perspective kind of the whole life cycle perspective kind of the whole life cycle perspective gets shifted onto re review and they're gets shifted onto re review and they're gets shifted onto re review and they're subsequently merging the PR. Does that subsequently merging the PR. Does that subsequently merging the PR. Does that resonate? resonate? resonate? I see some heads nodding. So this is I see some heads nodding. So this is I see some heads nodding. So this is where we spend time on figuring out how where we spend time on figuring out how where we spend time on figuring out how we can further improve the review we can further improve the review we can further improve the review process as well as the continuous process as well as the continuous process as well as the continuous integration the CI process. So we will integration the CI process. So we will integration the CI process. So we will dive into these two topics in the next dive into these two topics in the next dive into these two topics in the next couple slides here. I just want to say a couple slides here. I just want to say a couple slides here. I just want to say a potential anti-attern anti-solution is potential anti-attern anti-solution is potential anti-attern anti-solution is that hey if the CI infrastructure gets that hey if the CI infrastructure gets that hey if the CI infrastructure gets overloaded then a workar around by overloaded then a workar around by overloaded then a workar around by engineers to stop splitting PR just engineers to stop splitting PR just engineers to stop splitting PR just start submitting large PR for review and start submitting large PR for review and start submitting large PR for review and submission because if it takes an hour submission because if it takes an hour submission because if it takes an hour to run all of your regression test and to run all of your regression test and to run all of your regression test and submit it I don't want to break my PR submit it I don't want to break my PR submit it I don't want to break my PR into 10 right which might take 10 hours into 10 right which might take 10 hours into 10 right which might take 10 hours however this in our view can be pretty however this in our view can be pretty however this in our view can be pretty risky because it makes the human review risky because it makes the human review risky because it makes the human review overhead higher it also reduce the overhead higher it also reduce the overhead higher it also reduce the quality of the review because the human quality of the review because the human quality of the review because the human attention can be spread thin so that is attention can be spread thin so that is attention can be spread thin so that is an anti-attern I wanted to caution so for code review the key principle we so for code review the key principle we use is to make sure we onboard AI use is to make sure we onboard AI use is to make sure we onboard AI tooling as the first level of defense

  12. tooling as the first level of defense tooling as the first level of defense they don't replace human reviewers but they don't replace human reviewers but they don't replace human reviewers but we want to offload human reviewers as we want to offload human reviewers as we want to offload human reviewers as much as possible let the AI review take much as possible let the AI review take much as possible let the AI review take care of simpler things like coding style care of simpler things like coding style care of simpler things like coding style issues or if there's a missing test issues or if there's a missing test issues or if there's a missing test coverage. So, make sure the author gets coverage. So, make sure the author gets coverage. So, make sure the author gets through all of them before then the through all of them before then the through all of them before then the review gets routed to a human reviewer. review gets routed to a human reviewer. review gets routed to a human reviewer. And this way our human engineers can And this way our human engineers can And this way our human engineers can focus on applying their deep judgment on focus on applying their deep judgment on focus on applying their deep judgment on aspects that are somewhat subjective aspects that are somewhat subjective aspects that are somewhat subjective like if the code is good, if the like if the code is good, if the like if the code is good, if the architecture is sound, if the code uh uh architecture is sound, if the code uh uh architecture is sound, if the code uh uh passes kind of the security uh the passes kind of the security uh the passes kind of the security uh the security design and so on. so that in security design and so on. so that in security design and so on. so that in the end our engineering team can take the end our engineering team can take the end our engineering team can take the final accountability. Now let's look at CI. So I assume all of Now let's look at CI. So I assume all of you deploy some form of CI uh CI/CD and you deploy some form of CI uh CI/CD and you deploy some form of CI uh CI/CD and what we're seeing is thanks to AI now what we're seeing is thanks to AI now what we're seeing is thanks to AI now making it much easier to generate code making it much easier to generate code making it much easier to generate code as splitting code into smaller but more as splitting code into smaller but more as splitting code into smaller but more PRs it puts a lot more pressure on the PRs it puts a lot more pressure on the PRs it puts a lot more pressure on the CI and this is something that uh if we CI and this is something that uh if we CI and this is something that uh if we don't address uh at a company level don't address uh at a company level don't address uh at a company level individual engineers can be struggling individual engineers can be struggling individual engineers can be struggling because that means they have to waste because that means they have to waste because that means they have to waste their human time babysitting the PR to their human time babysitting the PR to their human time babysitting the PR to get merged. If they run into flaky test get merged. If they run into flaky test get merged. If they run into flaky test then they have to manually they hit then they have to manually they hit then they have to manually they hit rerun it's very frustrating or they can rerun it's very frustrating or they can rerun it's very frustrating or they can recruit an AI agent to babysit and kind recruit an AI agent to babysit and kind recruit an AI agent to babysit and kind of do a loop but that in turn waste AI of do a loop but that in turn waste AI of do a loop but that in turn waste AI token as well. So these are not these

  13. token as well. So these are not these token as well. So these are not these are just workarounds not perfect are just workarounds not perfect are just workarounds not perfect solution and also tend to make engineers solution and also tend to make engineers solution and also tend to make engineers feel a little bit lower morale a little feel a little bit lower morale a little feel a little bit lower morale a little bit more frustrated. So what we what we bit more frustrated. So what we what we bit more frustrated. So what we what we are doing is kind of we put more uh are doing is kind of we put more uh are doing is kind of we put more uh developer experience uh platform kind of developer experience uh platform kind of developer experience uh platform kind of engineering to invest into reducing engineering to invest into reducing engineering to invest into reducing removing the flaky test improving the CI removing the flaky test improving the CI removing the flaky test improving the CI infrastructure and the key thing here is infrastructure and the key thing here is infrastructure and the key thing here is to also define and measure the right to also define and measure the right to also define and measure the right metrics for example uh the work clock metrics for example uh the work clock metrics for example uh the work clock time between when a peer is ready to time between when a peer is ready to time between when a peer is ready to submit till when it's submitted right if submit till when it's submitted right if submit till when it's submitted right if a typical CR uh run takes an hour. Does a typical CR uh run takes an hour. Does a typical CR uh run takes an hour. Does the typical PR submission take two or the typical PR submission take two or the typical PR submission take two or three hours? In which case, that's a red three hours? In which case, that's a red three hours? In which case, that's a red flag and also the number of times a PR flag and also the number of times a PR flag and also the number of times a PR needs to get retrieded for passing needs to get retrieded for passing needs to get retrieded for passing through the test. So, these are the key through the test. So, these are the key through the test. So, these are the key metrics that we are using to measure our metrics that we are using to measure our metrics that we are using to measure our developer experiences and the relevant developer experiences and the relevant developer experiences and the relevant team who is focused on improving uh team who is focused on improving uh team who is focused on improving uh these uh the developer experience.

  14. So with all of the analysis and ideas So with all of the analysis and ideas here we uh want to share kind of the a here we uh want to share kind of the a here we uh want to share kind of the a pragmatic framework of how we can then pragmatic framework of how we can then pragmatic framework of how we can then measure and optimize token usage. It has measure and optimize token usage. It has measure and optimize token usage. It has three aspects. The first one is set the three aspects. The first one is set the three aspects. The first one is set the right set of guards across setting the right set of guards across setting the right set of guards across setting the budget and quota tracking usage defining budget and quota tracking usage defining budget and quota tracking usage defining anomalies so that no users uh leaders anomalies so that no users uh leaders anomalies so that no users uh leaders can get notified if something feels can get notified if something feels can get notified if something feels wrong. This is complementaryary to still wrong. This is complementaryary to still wrong. This is complementaryary to still regular human review which can catch regular human review which can catch regular human review which can catch other interesting patterns or learnings other interesting patterns or learnings other interesting patterns or learnings and feedback into the institutional and feedback into the institutional and feedback into the institutional knowledge base. knowledge base. knowledge base. Let me just couple that with the third Let me just couple that with the third Let me just couple that with the third item here which is the learning loop we item here which is the learning loop we item here which is the learning loop we talk about as our leadership work with talk about as our leadership work with talk about as our leadership work with individuals to define these guard rails individuals to define these guard rails individuals to define these guard rails review the metrics and then refine review the metrics and then refine review the metrics and then refine that's how we kind of close the learning that's how we kind of close the learning that's how we kind of close the learning loop. In addition to that, we want to loop. In addition to that, we want to loop. In addition to that, we want to work with our teams, individual work with our teams, individual work with our teams, individual engineers to continue to search for and engineers to continue to search for and engineers to continue to search for and if needed innovate on the best practices if needed innovate on the best practices if needed innovate on the best practices of how to use AI, how to use AI to build of how to use AI, how to use AI to build of how to use AI, how to use AI to build products and also use it internally. For products and also use it internally. For products and also use it internally. For example, example, example, some engineers may be writing an agentic some engineers may be writing an agentic some engineers may be writing an agentic loop as part of the harness when they loop as part of the harness when they loop as part of the harness when they use cloud code. after they generated use cloud code. after they generated use cloud code. after they generated initial PR they go and loop around and initial PR they go and loop around and initial PR they go and loop around and say try and pass the set of tests and say try and pass the set of tests and say try and pass the set of tests and then if some tests don't pass just auto then if some tests don't pass just auto then if some tests don't pass just auto fix the test or the code and retry. One fix the test or the code and retry. One fix the test or the code and retry. One thing to watch out for is to put a limit thing to watch out for is to put a limit thing to watch out for is to put a limit on the number of loop steps to make sure on the number of loop steps to make sure on the number of loop steps to make sure if things go out of control we don't

  15. if things go out of control we don't if things go out of control we don't waste too many tokens on that. Another waste too many tokens on that. Another waste too many tokens on that. Another example is prompt caching. This is example is prompt caching. This is example is prompt caching. This is becoming increasingly more prevalent by becoming increasingly more prevalent by becoming increasingly more prevalent by the commercial uh model vendors where the commercial uh model vendors where the commercial uh model vendors where what they advise is if you send a prompt what they advise is if you send a prompt what they advise is if you send a prompt with the same prefix they could optimize with the same prefix they could optimize with the same prefix they could optimize how they process the prefix of the how they process the prefix of the how they process the prefix of the prompt. What that means then as a user prompt. What that means then as a user prompt. What that means then as a user to those is that we want to encourage to those is that we want to encourage to those is that we want to encourage our users to structure their prompt that our users to structure their prompt that our users to structure their prompt that way. For example, if your prompt way. For example, if your prompt way. For example, if your prompt consists of a system prompt followed by consists of a system prompt followed by consists of a system prompt followed by a user prompt, you want to put the a user prompt, you want to put the a user prompt, you want to put the system prompt that's fixed at the top system prompt that's fixed at the top system prompt that's fixed at the top and the varying content at the bottom. and the varying content at the bottom. and the varying content at the bottom. Context pruning is also important. We Context pruning is also important. We Context pruning is also important. We want to kind of drill it into each want to kind of drill it into each want to kind of drill it into each individual users kind of new kind of individual users kind of new kind of individual users kind of new kind of muscle memory. So they are aware that as muscle memory. So they are aware that as muscle memory. So they are aware that as they build out the context through a they build out the context through a they build out the context through a longer chat session, they would be longer chat session, they would be longer chat session, they would be mindful of summarizing the context and mindful of summarizing the context and mindful of summarizing the context and make sure that the token usage is make sure that the token usage is make sure that the token usage is efficient that way. There are efficient that way. There are efficient that way. There are increasingly more tools like claw code increasingly more tools like claw code increasingly more tools like claw code that will automatically manage and that will automatically manage and that will automatically manage and compact the context for you. And so this compact the context for you. And so this compact the context for you. And so this increases the token usage efficiency but increases the token usage efficiency but increases the token usage efficiency but also increase the quality of AI output.

  16. also increase the quality of AI output. also increase the quality of AI output. There are other ideas we're exploring as There are other ideas we're exploring as There are other ideas we're exploring as well. So I know we're at time so this is So I know we're at time so this is towards the end of the talk. There is towards the end of the talk. There is towards the end of the talk. There is sometimes we also face build versus by sometimes we also face build versus by sometimes we also face build versus by decision. The principle is simple for decision. The principle is simple for decision. The principle is simple for things that are non- differentiating things that are non- differentiating things that are non- differentiating like IDE CI infrastructure we want to like IDE CI infrastructure we want to like IDE CI infrastructure we want to buy. But then for things that are buy. But then for things that are buy. But then for things that are specific to our context like how we specific to our context like how we specific to our context like how we would generate high quality PR for small would generate high quality PR for small would generate high quality PR for small bug fixes versus building a new UI bug fixes versus building a new UI bug fixes versus building a new UI feature for refactoring and so on. We feature for refactoring and so on. We feature for refactoring and so on. We have our internal playbook which is a have our internal playbook which is a have our internal playbook which is a set of well-crafted AI prompts. So we set of well-crafted AI prompts. So we set of well-crafted AI prompts. So we save that and share across our team. So save that and share across our team. So save that and share across our team. So that gets reused and enhanced. So that's that gets reused and enhanced. So that's that gets reused and enhanced. So that's something we must build internally. When something we must build internally. When something we must build internally. When it comes to case to case though, it comes to case to case though, it comes to case to case though, sometimes it's still a bit ambiguous sometimes it's still a bit ambiguous sometimes it's still a bit ambiguous like we're trying to build what we call like we're trying to build what we call like we're trying to build what we call builder agent. That's like a cloud-based builder agent. That's like a cloud-based builder agent. That's like a cloud-based code generation that wrap the cloud code generation that wrap the cloud code generation that wrap the cloud codec and so on. While we know there are codec and so on. While we know there are codec and so on. While we know there are also other vendors out there that we're also other vendors out there that we're also other vendors out there that we're still exploring. So we love to exchange still exploring. So we love to exchange still exploring. So we love to exchange thoughts on that.

  17. thoughts on that. thoughts on that. So then to summarize here are a couple So then to summarize here are a couple So then to summarize here are a couple key lessons as we went through the last key lessons as we went through the last key lessons as we went through the last couple quarters of journey. I wanted to couple quarters of journey. I wanted to couple quarters of journey. I wanted to share so that hopefully you could kind share so that hopefully you could kind share so that hopefully you could kind of accelerate your process there. If I of accelerate your process there. If I of accelerate your process there. If I were to summarize these three things I were to summarize these three things I were to summarize these three things I would it's about learning planning ahead would it's about learning planning ahead would it's about learning planning ahead and learn from other people's stories and learn from other people's stories and learn from other people's stories mistakes. So what that means is think mistakes. So what that means is think mistakes. So what that means is think about build respences by early on as you about build respences by early on as you about build respences by early on as you are encouraging more code gen think are encouraging more code gen think are encouraging more code gen think about how that impact your code review about how that impact your code review about how that impact your code review and CI and how you can address these new and CI and how you can address these new and CI and how you can address these new bottlenecks. And finally, continue to bottlenecks. And finally, continue to bottlenecks. And finally, continue to define and instrument your system to get define and instrument your system to get define and instrument your system to get the right metrics to measure the health the right metrics to measure the health the right metrics to measure the health of your CI system and the whole of your CI system and the whole of your CI system and the whole developer experience in general. developer experience in general. developer experience in general. So that's it for the talk. We believe So that's it for the talk. We believe So that's it for the talk. We believe that this is the golden era of AI where that this is the golden era of AI where that this is the golden era of AI where maximizing token ROI is the key for maximizing token ROI is the key for maximizing token ROI is the key for every team success. And with that, I every team success. And with that, I every team success. And with that, I just want to end with saying we are just want to end with saying we are just want to end with saying we are hiring. I know this is engineering hiring. I know this is engineering hiring. I know this is engineering leadership crowd but if you know of leadership crowd but if you know of leadership crowd but if you know of someone who is interested in building someone who is interested in building someone who is interested in building cutting edge legal contracting AI we cutting edge legal contracting AI we cutting edge legal contracting AI we would love to talk. Thank you.

Summary

The main theme is controlling and optimizing AI token spend, with references to sensational stories from Amazon and Meta highlighting the risks of uncontrolled usage. The practical takeaway is to implement dashboards that track token usage and cost for teams and individuals, but to position them as "smoke detectors" rather than leaderboards to avoid incentivizing excessive consumption.

View original episode ↗