From AI-Assisted to AI-Native: Building a Frontier Development Team — Clare Liguori, AWS
Read full transcript 18 segments
-
>> My name is Claire La Gory and I'm a >> My name is Claire La Gory and I'm a senior principal engineer at AWS. I senior principal engineer at AWS. I senior principal engineer at AWS. I mostly work on Kuro, our agent encoding mostly work on Kuro, our agent encoding mostly work on Kuro, our agent encoding assistant, but today I want to talk assistant, but today I want to talk assistant, but today I want to talk about some of the practices we've been about some of the practices we've been about some of the practices we've been seeing inside of Amazon and Amazon teams seeing inside of Amazon and Amazon teams seeing inside of Amazon and Amazon teams where we've been seeing really exciting where we've been seeing really exciting where we've been seeing really exciting results of productivity increases that results of productivity increases that results of productivity increases that are step function improvements since are step function improvements since are step function improvements since what what we've been seeing with AI so what what we've been seeing with AI so what what we've been seeing with AI so far. far. far. So, I've been working on agentic AI for So, I've been working on agentic AI for So, I've been working on agentic AI for over 3 years now and I've kind of seen over 3 years now and I've kind of seen over 3 years now and I've kind of seen the evolution that's happened in our the evolution that's happened in our the evolution that's happened in our industry when it comes to coding industry when it comes to coding industry when it comes to coding assistance with AI. First, we had this assistance with AI. First, we had this assistance with AI. First, we had this inline code completion helping us to inline code completion helping us to inline code completion helping us to write the next line, maybe the next write the next line, maybe the next write the next line, maybe the next function. We moved on to chat, asking function. We moved on to chat, asking function. We moved on to chat, asking questions about our code. Everybody questions about our code. Everybody questions about our code. Everybody started doing vibe coding sometime last started doing vibe coding sometime last started doing vibe coding sometime last year, but now we're starting to see kind year, but now we're starting to see kind year, but now we're starting to see kind of an early adopter phase of what we've of an early adopter phase of what we've of an early adopter phase of what we've been calling frontier development. been calling frontier development. been calling frontier development. And completely anecdotally, based on my And completely anecdotally, based on my And completely anecdotally, based on my own experience, I've really only felt own experience, I've really only felt own experience, I've really only felt maybe 10 to 20% more productive with all maybe 10 to 20% more productive with all maybe 10 to 20% more productive with all of these phases that have come before.
-
of these phases that have come before. of these phases that have come before. But now inside of Amazon, we've been But now inside of Amazon, we've been But now inside of Amazon, we've been running pilots with different teams running pilots with different teams running pilots with different teams across the company and we've been seeing across the company and we've been seeing across the company and we've been seeing a median of 4.5x productivity a median of 4.5x productivity a median of 4.5x productivity improvement and sometimes more than 10x. improvement and sometimes more than 10x. improvement and sometimes more than 10x. So, something has really changed here So, something has really changed here So, something has really changed here now that we're seeing these step now that we're seeing these step now that we're seeing these step function improvements in productivity. function improvements in productivity. function improvements in productivity. And I like to And I like to And I like to define what we've been calling frontier define what we've been calling frontier define what we've been calling frontier developers inside of Amazon by three developers inside of Amazon by three developers inside of Amazon by three behaviors that I've been seeing. One is behaviors that I've been seeing. One is behaviors that I've been seeing. One is hands-off coding. Frontier developers hands-off coding. Frontier developers hands-off coding. Frontier developers write maybe 1 to 2% of the code that write maybe 1 to 2% of the code that write maybe 1 to 2% of the code that they produce. The rest is agents. they produce. The rest is agents. they produce. The rest is agents. The second is that they interact with The second is that they interact with The second is that they interact with their agents infrequently. They'll aim their agents infrequently. They'll aim their agents infrequently. They'll aim to get their coding assistant to run for to get their coding assistant to run for to get their coding assistant to run for up to hours at a time without their up to hours at a time without their up to hours at a time without their intervention. intervention. intervention. And third is that they minimize idle And third is that they minimize idle And third is that they minimize idle time. time. time. These frontier developers tend to run These frontier developers tend to run These frontier developers tend to run multiple agents in parallel churning multiple agents in parallel churning multiple agents in parallel churning through a backlog of tasks. through a backlog of tasks. through a backlog of tasks. The first time that I saw a frontier The first time that I saw a frontier The first time that I saw a frontier developer team was the Bedrock Mantle developer team was the Bedrock Mantle developer team was the Bedrock Mantle team. Bedrock is our model hosting team. Bedrock is our model hosting team. Bedrock is our model hosting service.
-
service. service. Hosts LLMs like Claude and GPT. And Hosts LLMs like Claude and GPT. And Hosts LLMs like Claude and GPT. And sometime last year we knew or I say we sometime last year we knew or I say we sometime last year we knew or I say we but the Bedrock team but the Bedrock team but the Bedrock team knew that they were going to need to knew that they were going to need to knew that they were going to need to build a new inference data plane. But build a new inference data plane. But build a new inference data plane. But they had estimated it at 30 people over they had estimated it at 30 people over they had estimated it at 30 people over 18 months. This is a big big service and 18 months. This is a big big service and 18 months. This is a big big service and it was going to take time to build the it was going to take time to build the it was going to take time to build the new one, migrate customers over, migrate new one, migrate customers over, migrate new one, migrate customers over, migrate models over. They decided to take a step models over. They decided to take a step models over. They decided to take a step back. They took six people and they back. They took six people and they back. They took six people and they built [snorts] it in 76 days with Kiro. built [snorts] it in 76 days with Kiro. built [snorts] it in 76 days with Kiro. So this was a huge achievement. This was So this was a huge achievement. This was So this was a huge achievement. This was the first time we've we'd seen anything the first time we've we'd seen anything the first time we've we'd seen anything of the kind inside of Amazon. So this of the kind inside of Amazon. So this of the kind inside of Amazon. So this was truly the pathfinder team that was truly the pathfinder team that was truly the pathfinder team that proved that it was possible to get up to proved that it was possible to get up to proved that it was possible to get up to 20X improvement. Now they looked at 20X improvement. Now they looked at 20X improvement. Now they looked at commits and I'll talk about a couple of commits and I'll talk about a couple of commits and I'll talk about a couple of other ways that we are uh measuring other ways that we are uh measuring other ways that we are uh measuring productivity improvements. productivity improvements. productivity improvements. But there was one problem with this But there was one problem with this But there was one problem with this story which was that yes, it was built story which was that yes, it was built story which was that yes, it was built with six people. It was built with some with six people. It was built with some with six people. It was built with some of the top engineers literally in the of the top engineers literally in the of the top engineers literally in the company including two distinguished company including two distinguished company including two distinguished engineers. So this was not just any team engineers. So this was not just any team engineers. So this was not just any team of six people. These were experts in of six people. These were experts in of six people. These were experts in distributed systems, experts at LLMs and distributed systems, experts at LLMs and distributed systems, experts at LLMs and their architecture.
-
their architecture. their architecture. So this the story was amazing and it So this the story was amazing and it So this the story was amazing and it kind of spread like wildfire across kind of spread like wildfire across kind of spread like wildfire across Amazon, but it was also very Amazon, but it was also very Amazon, but it was also very unachievable for a lot of teams. There unachievable for a lot of teams. There unachievable for a lot of teams. There were a lot of questions about can this were a lot of questions about can this were a lot of questions about can this actually be reproduced on another team? actually be reproduced on another team? actually be reproduced on another team? So, another experiment that I want to So, another experiment that I want to So, another experiment that I want to talk about is an experimental sprint talk about is an experimental sprint talk about is an experimental sprint that was done in the Prime Video that was done in the Prime Video that was done in the Prime Video organization. organization. organization. They took a 10-day sprint and they did They took a 10-day sprint and they did They took a 10-day sprint and they did an experiment where they put, again, six an experiment where they put, again, six an experiment where they put, again, six engineers in a room and they let them go engineers in a room and they let them go engineers in a room and they let them go wild with Kiro. wild with Kiro. wild with Kiro. Uh they brought down the project Uh they brought down the project Uh they brought down the project delivery time estimate from what was delivery time estimate from what was delivery time estimate from what was going to be 90 weeks down to 24 based on going to be 90 weeks down to 24 based on going to be 90 weeks down to 24 based on all of the progress they had made in all of the progress they had made in all of the progress they had made in this 10-day sprint. And they they looked this 10-day sprint. And they they looked this 10-day sprint. And they they looked at their commit history and they looked at their commit history and they looked at their commit history and they looked at what did they used to do prior to at what did they used to do prior to at what did they used to do prior to this 10-day sprint and how many commits this 10-day sprint and how many commits this 10-day sprint and how many commits did they produce just in this 10 days. did they produce just in this 10 days. did they produce just in this 10 days. And so, this sprint really proved that And so, this sprint really proved that And so, this sprint really proved that we can achieve, again, at least we can achieve, again, at least we can achieve, again, at least something close to what the Bedrock something close to what the Bedrock something close to what the Bedrock Mantle team had uh had achieved with a Mantle team had uh had achieved with a Mantle team had uh had achieved with a different set of engineers.
-
different set of engineers. different set of engineers. But again, there was a challenge with But again, there was a challenge with But again, there was a challenge with this story, which was it was six this story, which was it was six this story, which was it was six engineers in a room, but they had no engineers in a room, but they had no engineers in a room, but they had no on-call duties, limited meetings, very on-call duties, limited meetings, very on-call duties, limited meetings, very few distractions, which we all know are few distractions, which we all know are few distractions, which we all know are regular in the lives of an engineer. regular in the lives of an engineer. regular in the lives of an engineer. And the senior engineer on the team had And the senior engineer on the team had And the senior engineer on the team had spent the previous 3 weeks creating very spent the previous 3 weeks creating very spent the previous 3 weeks creating very detailed, small, well-scoped tasks with detailed, small, well-scoped tasks with detailed, small, well-scoped tasks with detailed requirements for these detailed requirements for these detailed requirements for these six [clears throat] engineers to just go six [clears throat] engineers to just go six [clears throat] engineers to just go churn on for those 2 weeks. churn on for those 2 weeks. churn on for those 2 weeks. So, this was again not necessarily real So, this was again not necessarily real So, this was again not necessarily real life. This was a structured sprint, uh a life. This was a structured sprint, uh a life. This was a structured sprint, uh a a point in time that they were able to a point in time that they were able to a point in time that they were able to achieve this, but again, the question is achieve this, but again, the question is achieve this, but again, the question is is this achievable on real teams on is this achievable on real teams on is this achievable on real teams on day-to-day day-to-day day-to-day work? work? work? So, Amazon stores which encompasses So, Amazon stores which encompasses So, Amazon stores which encompasses amazon.com, all of our retail websites, amazon.com, all of our retail websites, amazon.com, all of our retail websites, as well as our physical stores, as well as our physical stores, as well as our physical stores, did a more structured pilot. They did a more structured pilot. They did a more structured pilot. They watched 50 teams that were totally watched 50 teams that were totally watched 50 teams that were totally normal normal distribution of um early normal normal distribution of um early normal normal distribution of um early career folks, mid-career, senior career folks, mid-career, senior career folks, mid-career, senior engineers, and that worked on existing engineers, and that worked on existing engineers, and that worked on existing systems. Nothing green field like the systems. Nothing green field like the systems. Nothing green field like the mantle team got to build from the ground mantle team got to build from the ground mantle team got to build from the ground up, but existing systems with existing up, but existing systems with existing up, but existing systems with existing code bases.
-
code bases. code bases. And they they watched them for the And they they watched them for the And they they watched them for the better part of last year, and they found better part of last year, and they found better part of last year, and they found something super interesting. something super interesting. something super interesting. They found that there was a big They found that there was a big They found that there was a big difference in the productivity gains difference in the productivity gains difference in the productivity gains that they saw between half of the teams that they saw between half of the teams that they saw between half of the teams and the other half. and the other half. and the other half. And in this case, they used a And in this case, they used a And in this case, they used a productivity metric of deployment productivity metric of deployment productivity metric of deployment velocity to production. So, not just velocity to production. So, not just velocity to production. So, not just commits, how many commits are they commits, how many commits are they commits, how many commits are they producing, but how quickly are we producing, but how quickly are we producing, but how quickly are we getting changes out to customers? How getting changes out to customers? How getting changes out to customers? How how quickly are we able to ship things? how quickly are we able to ship things? how quickly are we able to ship things? And they saw that for half of the teams, And they saw that for half of the teams, And they saw that for half of the teams, they achieved less than 3x increase. they achieved less than 3x increase. they achieved less than 3x increase. And what they found that was the And what they found that was the And what they found that was the difference between seeing less than 3x difference between seeing less than 3x difference between seeing less than 3x productivity increase, these teams that productivity increase, these teams that productivity increase, these teams that saw a median of 4.5x, and and in some saw a median of 4.5x, and and in some saw a median of 4.5x, and and in some cases more than 10, cases more than 10, cases more than 10, was how they used the tools. 90% of was how they used the tools. 90% of was how they used the tools. 90% of these teams used Kiro, among other these teams used Kiro, among other these teams used Kiro, among other internal tools that we have, and what internal tools that we have, and what internal tools that we have, and what they found was it wasn't about the they found was it wasn't about the they found was it wasn't about the tools, it was about the way that they tools, it was about the way that they tools, it was about the way that they worked.
-
worked. worked. The teams that achieved step function The teams that achieved step function The teams that achieved step function improvements improvements improvements intentionally changed the way that they intentionally changed the way that they intentionally changed the way that they worked, and the other simply kind of worked, and the other simply kind of worked, and the other simply kind of sprinkled Kiro and some of the other sprinkled Kiro and some of the other sprinkled Kiro and some of the other tools that we have on top of their tools that we have on top of their tools that we have on top of their existing way of working. And for me at existing way of working. And for me at existing way of working. And for me at least, this was the big aha moment. That least, this was the big aha moment. That least, this was the big aha moment. That why I hadn't been feeling potentially why I hadn't been feeling potentially why I hadn't been feeling potentially the massive gains that productive that the massive gains that productive that the massive gains that productive that in in productivity that AI has promised. in in productivity that AI has promised. in in productivity that AI has promised. It's about changing the way that we It's about changing the way that we It's about changing the way that we work. work. work. So, across this pilot, they went and So, across this pilot, they went and So, across this pilot, they went and interviewed uh the teams that were interviewed uh the teams that were interviewed uh the teams that were involved in the pilot as well as some of involved in the pilot as well as some of involved in the pilot as well as some of these other teams on the Bedrock mantel these other teams on the Bedrock mantel these other teams on the Bedrock mantel team, on uh Prime Video, and they found team, on uh Prime Video, and they found team, on uh Prime Video, and they found five habits. And and I use the word five habits. And and I use the word five habits. And and I use the word habits very specifically because again, habits very specifically because again, habits very specifically because again, it's not about that one sprint. It's it's not about that one sprint. It's it's not about that one sprint. It's about doing this day-to-day. And it And about doing this day-to-day. And it And about doing this day-to-day. And it And what they found when they interviewed what they found when they interviewed what they found when they interviewed with these teams was that it really was with these teams was that it really was with these teams was that it really was habits that they had to build habits that they had to build habits that they had to build day-to-day. When we change our way of day-to-day. When we change our way of day-to-day. When we change our way of working, it's it's hard to build these working, it's it's hard to build these working, it's it's hard to build these habits. It takes time to build these habits. It takes time to build these habits. It takes time to build these habits.
-
habits. habits. So, let's go through each of these one So, let's go through each of these one So, let's go through each of these one by one. by one. by one. Habit number one is investing in agent Habit number one is investing in agent Habit number one is investing in agent context. We have a lot of stuff in our context. We have a lot of stuff in our context. We have a lot of stuff in our head. We tend to transfer all of that head. We tend to transfer all of that head. We tend to transfer all of that stuff in our head to other people stuff in our head to other people stuff in our head to other people through Slack conversations, through through Slack conversations, through through Slack conversations, through onboarding, mentors, things like that, onboarding, mentors, things like that, onboarding, mentors, things like that, through code reviews, through through code reviews, through through code reviews, through stand-ups and sprint planning, and they stand-ups and sprint planning, and they stand-ups and sprint planning, and they had to write all of that down. And the had to write all of that down. And the had to write all of that down. And the habit that they built was every time the habit that they built was every time the habit that they built was every time the agent makes a mistake or does something agent makes a mistake or does something agent makes a mistake or does something not the way that you would have done it, not the way that you would have done it, not the way that you would have done it, what am I missing in my skills files? what am I missing in my skills files? what am I missing in my skills files? What am I missing in my steering files What am I missing in my steering files What am I missing in my steering files that the agent needed? that the agent needed? that the agent needed? But then, as we know, across last year, But then, as we know, across last year, But then, as we know, across last year, we saw leaps and bounds in models' we saw leaps and bounds in models' we saw leaps and bounds in models' abilities and their behaviors. abilities and their behaviors. abilities and their behaviors. Uh the Sonnet 3.7 in the middle of last Uh the Sonnet 3.7 in the middle of last Uh the Sonnet 3.7 in the middle of last year had a lot of quirks that we had to year had a lot of quirks that we had to year had a lot of quirks that we had to put a lot of do nots in our uh in our put a lot of do nots in our uh in our put a lot of do nots in our uh in our steering files, and now we don't have to steering files, and now we don't have to steering files, and now we don't have to do that as much with Opus 4.5 as of last do that as much with Opus 4.5 as of last do that as much with Opus 4.5 as of last November, and then we've had 6 months November, and then we've had 6 months November, and then we've had 6 months more than 6 months of improvement since more than 6 months of improvement since more than 6 months of improvement since then then then uh with all of the new versions of uh with all of the new versions of uh with all of the new versions of models that have come out since then.
-
models that have come out since then. models that have come out since then. And so, the question, the new habit, And so, the question, the new habit, And so, the question, the new habit, again, is do I still need this in my again, is do I still need this in my again, is do I still need this in my steering files or is this just bloating steering files or is this just bloating steering files or is this just bloating context? context? context? The second one is slowing down to speed The second one is slowing down to speed The second one is slowing down to speed up. In almost every team that was up. In almost every team that was up. In almost every team that was interviewed, they reported that their interviewed, they reported that their interviewed, they reported that their productivity actually went down as they productivity actually went down as they productivity actually went down as they intentionally adopted a new way of intentionally adopted a new way of intentionally adopted a new way of working. working. working. That's counterintuitive, right? You have That's counterintuitive, right? You have That's counterintuitive, right? You have to do intentional engineering work to do intentional engineering work to do intentional engineering work before you're going to see that hockey before you're going to see that hockey before you're going to see that hockey stick curve in productivity improvement. stick curve in productivity improvement. stick curve in productivity improvement. Because we have to do real work in our Because we have to do real work in our Because we have to do real work in our code base first for agents to be code base first for agents to be code base first for agents to be successful there, especially in successful there, especially in successful there, especially in brownfield existing code bases. So they brownfield existing code bases. So they brownfield existing code bases. So they had to build that agent context up. They had to build that agent context up. They had to build that agent context up. They had to improve existing tools error had to improve existing tools error had to improve existing tools error messages so that the model knew what was messages so that the model knew what was messages so that the model knew what was going on when it failed. They built new going on when it failed. They built new going on when it failed. They built new tools, new MCP servers for helping that tools, new MCP servers for helping that tools, new MCP servers for helping that model to actually get done what it model to actually get done what it model to actually get done what it needed to get done. A lot of teams ended needed to get done. A lot of teams ended needed to get done. A lot of teams ended up restructuring their code base so that up restructuring their code base so that up restructuring their code base so that agents could actually navigate it more agents could actually navigate it more agents could actually navigate it more easily. And I've even seen drastic easily. And I've even seen drastic easily. And I've even seen drastic changes like changing the programming changes like changing the programming changes like changing the programming language of the code base.
-
language of the code base. language of the code base. Um often I've seen teams struggle with Um often I've seen teams struggle with Um often I've seen teams struggle with Python, with JavaScript because they're Python, with JavaScript because they're Python, with JavaScript because they're untyped languages. It's hard to test. untyped languages. It's hard to test. untyped languages. It's hard to test. There's no compiler errors. So the model There's no compiler errors. So the model There's no compiler errors. So the model kind of guesses and give it gives it kind of guesses and give it gives it kind of guesses and give it gives it back to you. And so I've seen teams back to you. And so I've seen teams back to you. And so I've seen teams moving to TypeScript. Um Rust has become moving to TypeScript. Um Rust has become moving to TypeScript. Um Rust has become very popular inside of Amazon. The very popular inside of Amazon. The very popular inside of Amazon. The compiler gives great error messages. compiler gives great error messages. compiler gives great error messages. Um you don't have to do that, but I've Um you don't have to do that, but I've Um you don't have to do that, but I've seen a lot of teams making those seen a lot of teams making those seen a lot of teams making those intentional changes for the productivity intentional changes for the productivity intentional changes for the productivity gains that they're able to see. gains that they're able to see. gains that they're able to see. The third one is feeding agents, not The third one is feeding agents, not The third one is feeding agents, not babysitting agents. And for me this was babysitting agents. And for me this was babysitting agents. And for me this was one of those aha moments of why we're one of those aha moments of why we're one of those aha moments of why we're seeing this step function improvement in seeing this step function improvement in seeing this step function improvement in productivity. productivity. productivity. If you are vibe coding, if you are If you are vibe coding, if you are If you are vibe coding, if you are having a back-and-forth conversation having a back-and-forth conversation having a back-and-forth conversation with your agent all day long, of course with your agent all day long, of course with your agent all day long, of course you're not going to see four to five x you're not going to see four to five x you're not going to see four to five x productivity improvements because you productivity improvements because you productivity improvements because you are in the loop the entire time. You're are in the loop the entire time. You're are in the loop the entire time. You're probably sitting there for 30 seconds to probably sitting there for 30 seconds to probably sitting there for 30 seconds to a minute waiting for it to generate code a minute waiting for it to generate code a minute waiting for it to generate code and come back to you with with the code and come back to you with with the code and come back to you with with the code to review.
-
to review. to review. If you're sitting there waiting for it, If you're sitting there waiting for it, If you're sitting there waiting for it, then you can't go off and do other then you can't go off and do other then you can't go off and do other stuff. It's really difficult to run stuff. It's really difficult to run stuff. It's really difficult to run agents in parallel. It's very difficult agents in parallel. It's very difficult agents in parallel. It's very difficult to get to to clone yourself into to get to to clone yourself into to get to to clone yourself into multiple agents. And so if your multiple agents. And so if your multiple agents. And so if your conversations look a bit like this on conversations look a bit like this on conversations look a bit like this on the left, then you're babysitting that the left, then you're babysitting that the left, then you're babysitting that agent. As opposed to the right side agent. As opposed to the right side agent. As opposed to the right side where you're feeding it what it needs to where you're feeding it what it needs to where you're feeding it what it needs to do and how it can self-validate. And do and how it can self-validate. And do and how it can self-validate. And that's really the key so that agents can that's really the key so that agents can that's really the key so that agents can self-correct and only come back to you self-correct and only come back to you self-correct and only come back to you when it meets a certain quality bar, when it meets a certain quality bar, when it meets a certain quality bar, when it when it actually runs and when it when it actually runs and when it when it actually runs and compiles and passes tests, when it's compiles and passes tests, when it's compiles and passes tests, when it's testable, when it it actually has high testable, when it it actually has high testable, when it it actually has high coverage. And of course the next level coverage. And of course the next level coverage. And of course the next level is put all of this content into your is put all of this content into your is put all of this content into your steering file so it does it every time steering file so it does it every time steering file so it does it every time without you having to prompt it. The fourth habit is to make intent The fourth habit is to make intent explicit. At Amazon we practice a lot of explicit. At Amazon we practice a lot of explicit. At Amazon we practice a lot of behavior-driven development. We've built behavior-driven development. We've built behavior-driven development. We've built that into the Q product and so it's very that into the Q product and so it's very that into the Q product and so it's very natural for Amazon engineers to adopt it natural for Amazon engineers to adopt it natural for Amazon engineers to adopt it in Q. Um what what I've typically seen in Q. Um what what I've typically seen in Q. Um what what I've typically seen with live coding as opposed to frontier with live coding as opposed to frontier with live coding as opposed to frontier engineering is giving a very high-level engineering is giving a very high-level engineering is giving a very high-level prompt, letting the agent generate a ton prompt, letting the agent generate a ton prompt, letting the agent generate a ton of code, and then having a of code, and then having a of code, and then having a back-and-forth conversation saying, "Oh, back-and-forth conversation saying, "Oh, back-and-forth conversation saying, "Oh, that's not really what I meant. That you that's not really what I meant. That you that's not really what I meant. That you haven't you haven't exactly gotten the haven't you haven't exactly gotten the haven't you haven't exactly gotten the the requirements right. No, I didn't the requirements right. No, I didn't the requirements right. No, I didn't actually want to build it that way.
-
actually want to build it that way. actually want to build it that way. Here's a technical design." And it is Here's a technical design." And it is Here's a technical design." And it is less I find less productive to iterate less I find less productive to iterate less I find less productive to iterate with the agent on code when the intent with the agent on code when the intent with the agent on code when the intent itself was incorrect. So often will have itself was incorrect. So often will have itself was incorrect. So often will have will see Amazon engineers go through will see Amazon engineers go through will see Amazon engineers go through this process for for ambiguous complex this process for for ambiguous complex this process for for ambiguous complex features of writing the specification. features of writing the specification. features of writing the specification. And in Kiro, of course, you don't have And in Kiro, of course, you don't have And in Kiro, of course, you don't have to write this whole specification. You to write this whole specification. You to write this whole specification. You can have the model generate it, but it's can have the model generate it, but it's can have the model generate it, but it's a lot easier to to iterate with the a lot easier to to iterate with the a lot easier to to iterate with the model in kind of a back and forth model in kind of a back and forth model in kind of a back and forth conversation about a document than it is conversation about a document than it is conversation about a document than it is about code that's code changes that are about code that's code changes that are about code that's code changes that are spread across a code base. spread across a code base. spread across a code base. The fifth one is shift testing left. One The fifth one is shift testing left. One The fifth one is shift testing left. One of the keys here is to give the agent of the keys here is to give the agent of the keys here is to give the agent that fast feedback loop. that fast feedback loop. that fast feedback loop. Because that's what lets it go off for Because that's what lets it go off for Because that's what lets it go off for hours at a time and self-correct. The hours at a time and self-correct. The hours at a time and self-correct. The agent is going to make mistakes and agent is going to make mistakes and agent is going to make mistakes and that's fine. But if you give it the that's fine. But if you give it the that's fine. But if you give it the right signals, it can self-correct and right signals, it can self-correct and right signals, it can self-correct and it can spend a while doing that.
-
it can spend a while doing that. it can spend a while doing that. So, I've seen teams adding linters, So, I've seen teams adding linters, So, I've seen teams adding linters, adding unit tests, integration tests, adding unit tests, integration tests, adding unit tests, integration tests, performance tests, security tests. These performance tests, security tests. These performance tests, security tests. These are all things we all know we should are all things we all know we should are all things we all know we should have been doing all along. This is good have been doing all along. This is good have been doing all along. This is good engineering hygiene and practices. But engineering hygiene and practices. But engineering hygiene and practices. But now the ROI is, I think, finally high now the ROI is, I think, finally high now the ROI is, I think, finally high enough for actually us to actually enough for actually us to actually enough for actually us to actually invest in it. Um one thing that I've invest in it. Um one thing that I've invest in it. Um one thing that I've been seeing a lot of teams do is mock been seeing a lot of teams do is mock been seeing a lot of teams do is mock out services. Often with integration out services. Often with integration out services. Often with integration tests, we would test kind of end-to-end tests, we would test kind of end-to-end tests, we would test kind of end-to-end an entire system including live an entire system including live an entire system including live services. But we've been investing a lot services. But we've been investing a lot services. But we've been investing a lot in in mock services that run entirely in in mock services that run entirely in in mock services that run entirely locally with deterministic responses locally with deterministic responses locally with deterministic responses because it lets the agent do everything because it lets the agent do everything because it lets the agent do everything locally. Um doing everything on your locally. Um doing everything on your locally. Um doing everything on your laptop without having to spin up a bunch laptop without having to spin up a bunch laptop without having to spin up a bunch of other services and and connect to of other services and and connect to of other services and and connect to cloud services makes everything a lot cloud services makes everything a lot cloud services makes everything a lot faster because the the more that your faster because the the more that your faster because the the more that your agent can get fast feedback means the agent can get fast feedback means the agent can get fast feedback means the more loops that it can can do and the more loops that it can can do and the more loops that it can can do and the more productive your own agent can be.
-
more productive your own agent can be. more productive your own agent can be. So, across all of these, these are some So, across all of these, these are some So, across all of these, these are some of the habits we've seen, but of course of the habits we've seen, but of course of the habits we've seen, but of course I would be remiss if I would tell you if I would be remiss if I would tell you if I would be remiss if I would tell you if you adopt all of these habits, you will you adopt all of these habits, you will you adopt all of these habits, you will achieve nirvana. You will be the most achieve nirvana. You will be the most achieve nirvana. You will be the most productive engineering organization the productive engineering organization the productive engineering organization the world has ever seen. Things are still world has ever seen. Things are still world has ever seen. Things are still hard. We are still very much in an early hard. We are still very much in an early hard. We are still very much in an early adopter phase and teams are still adopter phase and teams are still adopter phase and teams are still figuring it out. figuring it out. figuring it out. So, one thing that we've been seeing So, one thing that we've been seeing So, one thing that we've been seeing across our teams just organizationally across our teams just organizationally across our teams just organizationally is the risk of burnout. I did not coin is the risk of burnout. I did not coin is the risk of burnout. I did not coin this term. I forget who did at what this term. I forget who did at what this term. I forget who did at what conference, but flow mat is real. We've conference, but flow mat is real. We've conference, but flow mat is real. We've been seeing engineers staying up late been seeing engineers staying up late been seeing engineers staying up late late at night late at night late at night trying to get that perfect prompt that's trying to get that perfect prompt that's trying to get that perfect prompt that's going to make their agent run for hours going to make their agent run for hours going to make their agent run for hours overnight so that they wake up in the overnight so that they wake up in the overnight so that they wake up in the morning with a code change ready. morning with a code change ready. morning with a code change ready. The cognitive load increases as you run The cognitive load increases as you run The cognitive load increases as you run these multiple agents in parallel. these multiple agents in parallel. these multiple agents in parallel. You're constantly shifting between You're constantly shifting between You're constantly shifting between terminal tabs. terminal tabs. terminal tabs. And then we do see that reviewing AI And then we do see that reviewing AI And then we do see that reviewing AI output is often harder for some than output is often harder for some than output is often harder for some than than actually writing it, especially than actually writing it, especially than actually writing it, especially early in career.
-
early in career. early in career. Senior engineers have have already spent Senior engineers have have already spent Senior engineers have have already spent a large portion of their career a large portion of their career a large portion of their career reviewing others code. reviewing others code. reviewing others code. But early career engineers don't have But early career engineers don't have But early career engineers don't have that muscle yet and so reviewing it can that muscle yet and so reviewing it can that muscle yet and so reviewing it can can feel like a lot more cognitive load can feel like a lot more cognitive load can feel like a lot more cognitive load than they're used to and actually than they're used to and actually than they're used to and actually writing it. writing it. writing it. The other one is organizational change. The other one is organizational change. The other one is organizational change. So, it's already hard to change the way So, it's already hard to change the way So, it's already hard to change the way we work as engineers. The way that we we work as engineers. The way that we we work as engineers. The way that we spend our entire day completely changes spend our entire day completely changes spend our entire day completely changes when we're frontier engineers, but also when we're frontier engineers, but also when we're frontier engineers, but also organizations have to change to enable organizations have to change to enable organizations have to change to enable frontier engineering teams. frontier engineering teams. frontier engineering teams. One that I've seen very commonly is One that I've seen very commonly is One that I've seen very commonly is accepting slowing down to speed up. accepting slowing down to speed up. accepting slowing down to speed up. And I've been guilty of this myself. My And I've been guilty of this myself. My And I've been guilty of this myself. My my fellow leaders have been guilty of of my fellow leaders have been guilty of of my fellow leaders have been guilty of of this of saying, "Well, you have the AI this of saying, "Well, you have the AI this of saying, "Well, you have the AI tools now and the models are so amazing tools now and the models are so amazing tools now and the models are so amazing now. Why are you not going faster? now. Why are you not going faster? now. Why are you not going faster? Um and that's because you have to take Um and that's because you have to take Um and that's because you have to take those two months to invest in your code those two months to invest in your code those two months to invest in your code base, to figure out the best practices base, to figure out the best practices base, to figure out the best practices for your team, to make hard habit for your team, to make hard habit for your team, to make hard habit changes on your team.
-
changes on your team. changes on your team. Um and and if you're constantly Um and and if you're constantly Um and and if you're constantly expecting expecting expecting shipping features every month because shipping features every month because shipping features every month because now we have these amazing models and now we have these amazing models and now we have these amazing models and we're seeing um all of these these we're seeing um all of these these we're seeing um all of these these companies on X saying how they're companies on X saying how they're companies on X saying how they're shipping 20 PRs a day, um we have to shipping 20 PRs a day, um we have to shipping 20 PRs a day, um we have to slow down to speed up. slow down to speed up. slow down to speed up. The second one is actually going too The second one is actually going too The second one is actually going too broad in the organization too fast. I broad in the organization too fast. I broad in the organization too fast. I think that if we had um expected all think that if we had um expected all think that if we had um expected all teams in massive organizations to be teams in massive organizations to be teams in massive organizations to be frontier teams immediately, we would not frontier teams immediately, we would not frontier teams immediately, we would not have had the learnings that we had from have had the learnings that we had from have had the learnings that we had from the Pathfinder, from the from the sprint the Pathfinder, from the from the sprint the Pathfinder, from the from the sprint experiment, from the pilot uh teams experiment, from the pilot uh teams experiment, from the pilot uh teams within Amazon. And now the challenge for within Amazon. And now the challenge for within Amazon. And now the challenge for us is how do we scale it out? And that's us is how do we scale it out? And that's us is how do we scale it out? And that's what 2026 is about for Amazon is how do what 2026 is about for Amazon is how do what 2026 is about for Amazon is how do we scale this out to more and more we scale this out to more and more we scale this out to more and more teams, to the next uh 2,000 teams teams, to the next uh 2,000 teams teams, to the next uh 2,000 teams instead of uh 50 teams. instead of uh 50 teams. instead of uh 50 teams. Um and so I think that when you roll it Um and so I think that when you roll it Um and so I think that when you roll it out too quickly, you have a lot of teams out too quickly, you have a lot of teams out too quickly, you have a lot of teams who don't know what they're doing. You who don't know what they're doing. You who don't know what they're doing. You haven't had time to find the best haven't had time to find the best haven't had time to find the best practices for your own organizations, practices for your own organizations, practices for your own organizations, the the context that your organization the the context that your organization the the context that your organization needs.
-
needs. needs. And the last one is that you're going to And the last one is that you're going to And the last one is that you're going to find new bottlenecks. find new bottlenecks. find new bottlenecks. Previously, code writing code manually Previously, code writing code manually Previously, code writing code manually was the bottleneck. Um I find that was the bottleneck. Um I find that was the bottleneck. Um I find that within Amazon, we've found um the speed within Amazon, we've found um the speed within Amazon, we've found um the speed of decision-making becomes a new of decision-making becomes a new of decision-making becomes a new bottleneck. Um the more that you spend bottleneck. Um the more that you spend bottleneck. Um the more that you spend reviewing the decision to actually build reviewing the decision to actually build reviewing the decision to actually build a new product, the slower it is to build a new product, the slower it is to build a new product, the slower it is to build the product now because the code only the product now because the code only the product now because the code only takes 1 to two months to write. takes 1 to two months to write. takes 1 to two months to write. >> [snorts] >> [snorts] >> [snorts] >> Um all of the review processes >> Um all of the review processes >> Um all of the review processes associated with the launch of a product associated with the launch of a product associated with the launch of a product become the bottleneck. When it used to become the bottleneck. When it used to become the bottleneck. When it used to take 9 to 12 months to build a new take 9 to 12 months to build a new take 9 to 12 months to build a new product, it didn't matter so much in the product, it didn't matter so much in the product, it didn't matter so much in the in the overall wash of things if it took in the overall wash of things if it took in the overall wash of things if it took two months to make the decision to build two months to make the decision to build two months to make the decision to build the product and then two months to the product and then two months to the product and then two months to approve the launch. But now those are approve the launch. But now those are approve the launch. But now those are the bottlenecks. Those are the long the bottlenecks. Those are the long the bottlenecks. Those are the long pole. And so you find all of these all pole. And so you find all of these all pole. And so you find all of these all of these things that slow you down. of these things that slow you down. of these things that slow you down. Often I find that frontier engineering Often I find that frontier engineering Often I find that frontier engineering teams spend more time making decisions teams spend more time making decisions teams spend more time making decisions than they do writing code. And so the than they do writing code. And so the than they do writing code. And so the more that you can make fast decisions, more that you can make fast decisions, more that you can make fast decisions, especially ones that are easy to be especially ones that are easy to be especially ones that are easy to be reversed, the better.
-
reversed, the better. reversed, the better. So my one big takeaway for for everyone So my one big takeaway for for everyone So my one big takeaway for for everyone here is that here is that here is that frontier engineering is about frontier engineering is about frontier engineering is about intentionally changing the way that you intentionally changing the way that you intentionally changing the way that you work. And that is difficult. That takes work. And that is difficult. That takes work. And that is difficult. That takes time. It is forming new habits and a new time. It is forming new habits and a new time. It is forming new habits and a new way of working. way of working. way of working. And that goes across any engineering And that goes across any engineering And that goes across any engineering team as well as your organization. Um so team as well as your organization. Um so team as well as your organization. Um so I encourage you to think about I encourage you to think about I encourage you to think about um how you're interacting with AI tools um how you're interacting with AI tools um how you're interacting with AI tools and how that can change to free yourself and how that can change to free yourself and how that can change to free yourself up from being in the loop. up from being in the loop. up from being in the loop. Um thanks. I'm going to I'll hang out uh Um thanks. I'm going to I'll hang out uh Um thanks. I'm going to I'll hang out uh a little bit if anyone has questions in a little bit if anyone has questions in a little bit if anyone has questions in the back. Um but thanks for the time the back. Um but thanks for the time the back. Um but thanks for the time today.
Summary
The main theme is the significant productivity gains achieved through frontier development practices using AI coding assistants, particularly at Amazon. Key subjects include agentic AI, coding assistance evolution from completion to chat, and the concept of "frontier developers" who write minimal code themselves, interact infrequently with agents, and minimize idle time. The practical takeaway is that adopting these frontier development behaviors can lead to dramatic, step-function improvements in developer productivity.