← Back
AI Engineer July 24, 2026 21m

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

Read full transcript 19 segments
  1. >> My name is Jay and I'm here with Sonya. >> My name is Jay and I'm here with Sonya. We are part of the computer vision team We are part of the computer vision team We are part of the computer vision team at Aruba. We're going to talk to you at Aruba. We're going to talk to you at Aruba. We're going to talk to you about a real world production use. about a real world production use. about a real world production use. Oh, my son done. Okay. Oh, my son done. Okay. Oh, my son done. Okay. Try again. Try again. Try again. Okay. Don't worry. I'll I'll manage. You Okay. Don't worry. I'll I'll manage. You Okay. Don't worry. I'll I'll manage. You hear me now? hear me now? hear me now? Okay, so we're going to talk to you Okay, so we're going to talk to you Okay, so we're going to talk to you today about a real world production use today about a real world production use today about a real world production use case case case and specifically we're going to dive and specifically we're going to dive and specifically we're going to dive into how we design the e-bows and the into how we design the e-bows and the into how we design the e-bows and the e-bow loops. So All right, cool. So just before we get All right, cool. So just before we get into the agent design, into the agent design, into the agent design, we're going to talk about a little bit we're going to talk about a little bit we're going to talk about a little bit about the use case. So our delivery about the use case. So our delivery about the use case. So our delivery marketplace Uber Eats, we do about 90 marketplace Uber Eats, we do about 90 marketplace Uber Eats, we do about 90 billion billion billion run rate per year at the moment. run rate per year at the moment. run rate per year at the moment. We were adding millions of items to the We were adding millions of items to the We were adding millions of items to the marketplace each and every year. marketplace each and every year. marketplace each and every year. Sorry, every every month. We're growing Sorry, every every month. We're growing Sorry, every every month. We're growing at 20% year-on-year and and we operate at 20% year-on-year and and we operate at 20% year-on-year and and we operate in 10,000 cities globally. So not many in 10,000 cities globally. So not many in 10,000 cities globally. So not many people actually know this but our people actually know this but our people actually know this but our delivery marketplace is just as big as delivery marketplace is just as big as delivery marketplace is just as big as the mobility side on Uber today.

  2. the mobility side on Uber today. the mobility side on Uber today. Visual content actually plays a really Visual content actually plays a really Visual content actually plays a really important role for the user experience. important role for the user experience. important role for the user experience. So a photo is quite often the first So a photo is quite often the first So a photo is quite often the first signal that a customer gets that gives signal that a customer gets that gives signal that a customer gets that gives them that initial impression about a them that initial impression about a them that initial impression about a merchant. merchant. merchant. So a good photo can make the difference So a good photo can make the difference So a good photo can make the difference between someone scrolling through the between someone scrolling through the between someone scrolling through the feed and actually clicking on an item feed and actually clicking on an item feed and actually clicking on an item and adding to the cart. And more and and adding to the cart. And more and and adding to the cart. And more and more we're seeing different modalities more we're seeing different modalities more we're seeing different modalities on Uber Eats uh especially video on Uber Eats uh especially video on Uber Eats uh especially video content. content. content. But this is a problem. But this is a problem. But this is a problem. So, our smaller independent merchants So, our smaller independent merchants So, our smaller independent merchants simply just don't have the level of simply just don't have the level of simply just don't have the level of quality for their photos that reflect quality for their photos that reflect quality for their photos that reflect what the eater is actually going to get. what the eater is actually going to get. what the eater is actually going to get. And when we speak to our merchants, And when we speak to our merchants, And when we speak to our merchants, there are three themes that kind of there are three themes that kind of there are three themes that kind of emerge. emerge. emerge. Lack of time, Lack of time, Lack of time, lack of know-how, and costs cuz these lack of know-how, and costs cuz these lack of know-how, and costs cuz these professional um photo shoots actually professional um photo shoots actually professional um photo shoots actually cost a lot of money. cost a lot of money. cost a lot of money. And this can be especially [snorts] And this can be especially [snorts] And this can be especially [snorts] problematic if the merchant is updating problematic if the merchant is updating problematic if the merchant is updating their menu over time.

  3. So, this problem is actually pretty So, this problem is actually pretty challenging to solve for at scale, challenging to solve for at scale, challenging to solve for at scale, right? Because our consumers, they want right? Because our consumers, they want right? Because our consumers, they want authentic, real-looking photos, authentic, real-looking photos, authentic, real-looking photos, um but a meaningful fraction of uh um but a meaningful fraction of uh um but a meaningful fraction of uh consumers actually distrust anything consumers actually distrust anything consumers actually distrust anything that is AI-generated. So, if you open up that is AI-generated. So, if you open up that is AI-generated. So, if you open up the Uber Eats app, the last thing that the Uber Eats app, the last thing that the Uber Eats app, the last thing that you want is to be scrolling through uh you want is to be scrolling through uh you want is to be scrolling through uh you know, food photography that looks you know, food photography that looks you know, food photography that looks like AI's lock. like AI's lock. like AI's lock. So, we're threading the needle here. We So, we're threading the needle here. We So, we're threading the needle here. We need to be able to stay faithful to the need to be able to stay faithful to the need to be able to stay faithful to the original image, preserve the brand of original image, preserve the brand of original image, preserve the brand of the merchant, and avoid everything the merchant, and avoid everything the merchant, and avoid everything looking the same. If we have the same looking the same. If we have the same looking the same. If we have the same prompt for every photo that we're prompt for every photo that we're prompt for every photo that we're editing, the diversity of the editing, the diversity of the editing, the diversity of the marketplace is going to collapse. We also, because we operate globally, we We also, because we operate globally, we also have this long-tail distribution of also have this long-tail distribution of also have this long-tail distribution of different quality that we see it across different quality that we see it across different quality that we see it across the marketplace. the marketplace. the marketplace. Um so, we've got some examples here. You Um so, we've got some examples here. You Um so, we've got some examples here. You might see food photography that, you might see food photography that, you might see food photography that, you know, has poor sharpness, poor know, has poor sharpness, poor know, has poor sharpness, poor composition, not centered, uh or or poor composition, not centered, uh or or poor composition, not centered, uh or or poor colors as well. We also have a wide colors as well. We also have a wide colors as well. We also have a wide range of spectrum of user-generated range of spectrum of user-generated range of spectrum of user-generated content on the platform as well.

  4. So, what are our goals when we're So, what are our goals when we're designing these agents? When you think designing these agents? When you think designing these agents? When you think through these goals, you might actually through these goals, you might actually through these goals, you might actually be thinking through, you know, your own be thinking through, you know, your own be thinking through, you know, your own agents that you're building yourself. agents that you're building yourself. agents that you're building yourself. But for us, it's about one, preserving But for us, it's about one, preserving But for us, it's about one, preserving authenticity and trust. authenticity and trust. authenticity and trust. Two, improving the quality when we need Two, improving the quality when we need Two, improving the quality when we need to. So, we want to be able to improve to. So, we want to be able to improve to. So, we want to be able to improve quality selectively. quality selectively. quality selectively. We want to optimize globally for for the We want to optimize globally for for the We want to optimize globally for for the entire marketplace. We don't entire marketplace. We don't entire marketplace. We don't cannibalize certain merchants. We want cannibalize certain merchants. We want cannibalize certain merchants. We want to ship safely, and this is going to be to ship safely, and this is going to be to ship safely, and this is going to be an important theme throughout the talk. an important theme throughout the talk. an important theme throughout the talk. We want to learn continuously, We want to learn continuously, We want to learn continuously, and we want to operate at scale in a and we want to operate at scale in a and we want to operate at scale in a cost-efficient manner. cost-efficient manner. cost-efficient manner. So, agents are actually really So, agents are actually really So, agents are actually really well-suited to solve this problem. well-suited to solve this problem. well-suited to solve this problem. So, if you imagine a spectrum, on the So, if you imagine a spectrum, on the So, if you imagine a spectrum, on the one side, you've got something that's one side, you've got something that's one side, you've got something that's more deterministic. It's more more deterministic. It's more more deterministic. It's more rules-based. Um and uh you you have more rules-based. Um and uh you you have more rules-based. Um and uh you you have more control over it, but it's fairly it's a control over it, but it's fairly it's a control over it, but it's fairly it's a brittle system. It's not actually going brittle system. It's not actually going brittle system. It's not actually going to be able to scale for the entire to be able to scale for the entire to be able to scale for the entire marketplace. marketplace. marketplace. Imagine the other side, you provide an Imagine the other side, you provide an Imagine the other side, you provide an agent with obviously a lot of agent with obviously a lot of agent with obviously a lot of creativity, it has a lot of agency. Um creativity, it has a lot of agency. Um creativity, it has a lot of agency. Um and that's actually what we want to lean and that's actually what we want to lean and that's actually what we want to lean into, into, into, but we can't leave that unconstrained, but we can't leave that unconstrained, but we can't leave that unconstrained, right? Cuz we have certain safety and right? Cuz we have certain safety and right? Cuz we have certain safety and certain guardrails in place that we need certain guardrails in place that we need certain guardrails in place that we need to adhere to.

  5. to adhere to. to adhere to. So, we want to find a balancing act. So, we want to find a balancing act. So, we want to find a balancing act. Uh and that's kind of set the principle Uh and that's kind of set the principle Uh and that's kind of set the principle for the way that we think and design for the way that we think and design for the way that we think and design around agents and evals. around agents and evals. around agents and evals. So, now we're going to actually like So, now we're going to actually like So, now we're going to actually like dive a little bit deeper into a dive a little bit deeper into a dive a little bit deeper into a simplified but representative example of simplified but representative example of simplified but representative example of what we have in production. what we have in production. what we have in production. And we're going to go through each stage And we're going to go through each stage And we're going to go through each stage and how we eval it, and then talk and how we eval it, and then talk and how we eval it, and then talk through some continuous learning loops through some continuous learning loops through some continuous learning loops as well. as well. as well. So, first up, we have what we call an So, first up, we have what we call an So, first up, we have what we call an image understanding and routing agents. image understanding and routing agents. image understanding and routing agents. So, this is where multimodality is is So, this is where multimodality is is So, this is where multimodality is is pretty important. We actually ask the pretty important. We actually ask the pretty important. We actually ask the LLM to describe what it sees in the LLM to describe what it sees in the LLM to describe what it sees in the photo. photo. photo. Um and then we we create a structured Um and then we we create a structured Um and then we we create a structured output from that, and we send it to a output from that, and we send it to a output from that, and we send it to a router. router. router. The router will then determine, do we The router will then determine, do we The router will then determine, do we enhance it, or do we skip it? We skip enhance it, or do we skip it? We skip enhance it, or do we skip it? We skip it, we will keep the original. it, we will keep the original. it, we will keep the original. If we enhance it, we send it to our next If we enhance it, we send it to our next If we enhance it, we send it to our next agent, agent, agent, which is an image editing agent. And which is an image editing agent. And which is an image editing agent. And this can actually run in a loop. So, it this can actually run in a loop. So, it this can actually run in a loop. So, it gets feedback from a QA agent. Um it can gets feedback from a QA agent. Um it can gets feedback from a QA agent. Um it can edit uh edit uh edit uh in the in this loop and self-correct and in the in this loop and self-correct and in the in this loop and self-correct and fix things fix things fix things um as it goes.

  6. um as it goes. um as it goes. If it goes through a number of loops and If it goes through a number of loops and If it goes through a number of loops and it still fails, we we don't publish it. it still fails, we we don't publish it. it still fails, we we don't publish it. Then we actually send it to a final Then we actually send it to a final Then we actually send it to a final post-processing and QA step. post-processing and QA step. post-processing and QA step. If that's all good, we'll publish it to If that's all good, we'll publish it to If that's all good, we'll publish it to the menu. the menu. the menu. And the last thing that's really And the last thing that's really And the last thing that's really critical is we log everything. critical is we log everything. critical is we log everything. Just a quick note about logging. Just a quick note about logging. Just a quick note about logging. Don't know if you can actually read the Don't know if you can actually read the Don't know if you can actually read the JSON here, but you might notice that all JSON here, but you might notice that all JSON here, but you might notice that all of the agents in this end-to-end of the agents in this end-to-end of the agents in this end-to-end orchestration is within one It's It's orchestration is within one It's It's orchestration is within one It's It's basically a flat structure in this JSON. basically a flat structure in this JSON. basically a flat structure in this JSON. Um and so, this is actually incredibly Um and so, this is actually incredibly Um and so, this is actually incredibly useful for the entire team useful for the entire team useful for the entire team because anyone, be it non-technical uh because anyone, be it non-technical uh because anyone, be it non-technical uh technical um folks on engineering technical um folks on engineering technical um folks on engineering product, can actually dive in um and product, can actually dive in um and product, can actually dive in um and look at specific cases to diagnose and look at specific cases to diagnose and look at specific cases to diagnose and also roll up things to look in also roll up things to look in also roll up things to look in aggregates. aggregates. aggregates. Um and it's important to note here that, Um and it's important to note here that, Um and it's important to note here that, you know, we think this is important to you know, we think this is important to you know, we think this is important to start with. You want to start with your start with. You want to start with your start with. You want to start with your logging cuz if you don't start with it, logging cuz if you don't start with it, logging cuz if you don't start with it, you have nothing to optimize for, let you have nothing to optimize for, let you have nothing to optimize for, let alone set up a self-learning loop. And alone set up a self-learning loop. And alone set up a self-learning loop. And at Uber, we um we use our eyes.

  7. at Uber, we um we use our eyes. at Uber, we um we use our eyes. Cool. We're going to dive um a bit Cool. We're going to dive um a bit Cool. We're going to dive um a bit deeper into the router. So, the router's actually pretty So, the router's actually pretty straightforward. If you remember, we you straightforward. If you remember, we you straightforward. If you remember, we you know, we have this multimodality input. know, we have this multimodality input. know, we have this multimodality input. We look at certain text description We look at certain text description We look at certain text description metadata, the image itself. We ask it to metadata, the image itself. We ask it to metadata, the image itself. We ask it to to to to under um describe what it's to to to under um describe what it's to to to under um describe what it's seeing. We create structured output from seeing. We create structured output from seeing. We create structured output from that. With that structured output, we that. With that structured output, we that. With that structured output, we can then grade against a rubric. So, we can then grade against a rubric. So, we can then grade against a rubric. So, we have these pass and fail criteria. The have these pass and fail criteria. The have these pass and fail criteria. The last step is we want to decide whether last step is we want to decide whether last step is we want to decide whether or not we should enhance or skip. or not we should enhance or skip. or not we should enhance or skip. How do we actually eval this? How do we actually eval this? How do we actually eval this? This is you could think of this as a This is you could think of this as a This is you could think of this as a more sort of traditional classifier. So, more sort of traditional classifier. So, more sort of traditional classifier. So, here we we have a confusion matrix. You here we we have a confusion matrix. You here we we have a confusion matrix. You know, many of you are probably pretty know, many of you are probably pretty know, many of you are probably pretty familiar with this. familiar with this. familiar with this. Um but we can look at things like the Um but we can look at things like the Um but we can look at things like the true positive cases, the false negative true positive cases, the false negative true positive cases, the false negative negative cases, and so on and so forth. negative cases, and so on and so forth. negative cases, and so on and so forth. Essentially, what we're doing is we're Essentially, what we're doing is we're Essentially, what we're doing is we're measuring the precision recall. measuring the precision recall. measuring the precision recall. In practice, your routers might actually In practice, your routers might actually In practice, your routers might actually be much more sophisticated. So, for be much more sophisticated. So, for be much more sophisticated. So, for example, we might want to route an image example, we might want to route an image example, we might want to route an image to a lower latency smaller model to be to a lower latency smaller model to be to a lower latency smaller model to be able to save on cost and improve the able to save on cost and improve the able to save on cost and improve the user experience at the trade-off of user experience at the trade-off of user experience at the trade-off of quality.

  8. quality. quality. And if that's the case, instead of And if that's the case, instead of And if that's the case, instead of having a 2 by 2 matrix for your having a 2 by 2 matrix for your having a 2 by 2 matrix for your confusion matrix, you might actually confusion matrix, you might actually confusion matrix, you might actually have an n by n matrix. have an n by n matrix. have an n by n matrix. Where each grid is actually telling you Where each grid is actually telling you Where each grid is actually telling you whether or not you're correctly routing whether or not you're correctly routing whether or not you're correctly routing to that specific branch. to that specific branch. to that specific branch. So, I'm going to now hand over to Somya So, I'm going to now hand over to Somya So, I'm going to now hand over to Somya who's going to dive a little bit deeper who's going to dive a little bit deeper who's going to dive a little bit deeper into how we handle drift and human into how we handle drift and human into how we handle drift and human alignment. >> So, now that we spoke about how we eval >> So, now that we spoke about how we eval the routing, I want to talk about how do the routing, I want to talk about how do the routing, I want to talk about how do you get the first version of the model you get the first version of the model you get the first version of the model out. out. out. For our use case, we consider human For our use case, we consider human For our use case, we consider human labels as the golden source of truth. labels as the golden source of truth. labels as the golden source of truth. And this is what we want to align our And this is what we want to align our And this is what we want to align our models to. models to. models to. The way we do about this is we go The way we do about this is we go The way we do about this is we go collect a dataset which is collect a dataset which is collect a dataset which is representative. So, you know, different representative. So, you know, different representative. So, you know, different cuts, geographies, dish type, image cuts, geographies, dish type, image cuts, geographies, dish type, image quality type. Send it to our human quality type. Send it to our human quality type. Send it to our human labelers and give them a very objective labelers and give them a very objective labelers and give them a very objective guideline to label on. guideline to label on. guideline to label on. This is to remove any subjective biases This is to remove any subjective biases This is to remove any subjective biases or any noise coming in from human or any noise coming in from human or any noise coming in from human labelers.

  9. labelers. labelers. Once we've got that system set up is Once we've got that system set up is Once we've got that system set up is when we start tuning our model. We take when we start tuning our model. We take when we start tuning our model. We take our agent, we go ahead get output from our agent, we go ahead get output from our agent, we go ahead get output from the agent, compare it to your golden the agent, compare it to your golden the agent, compare it to your golden dataset, evaluate if it's good enough to dataset, evaluate if it's good enough to dataset, evaluate if it's good enough to ship, if it meets your guardrail ship, if it meets your guardrail ship, if it meets your guardrail metrics, you go ahead and ship it. If metrics, you go ahead and ship it. If metrics, you go ahead and ship it. If not, then you go tune and you keep doing not, then you go tune and you keep doing not, then you go tune and you keep doing this until you meet your guardrail this until you meet your guardrail this until you meet your guardrail metrics. metrics. metrics. For routing, our guardrail metric is For routing, our guardrail metric is For routing, our guardrail metric is recall. We don't want any bad image to recall. We don't want any bad image to recall. We don't want any bad image to slip through our system. Here are some examples of the failures Here are some examples of the failures we've seen. we've seen. we've seen. Uh on your left you see a very good Uh on your left you see a very good Uh on your left you see a very good image of cheeseburger. Uh on the right image of cheeseburger. Uh on the right image of cheeseburger. Uh on the right you notice that the routing agent you notice that the routing agent you notice that the routing agent actually failed this. It said the actually failed this. It said the actually failed this. It said the technical is low ball and it will go technical is low ball and it will go technical is low ball and it will go send this image for enhancement. Now send this image for enhancement. Now send this image for enhancement. Now there's two challenges when you send there's two challenges when you send there's two challenges when you send this image for enhancement. Firstly, you this image for enhancement. Firstly, you this image for enhancement. Firstly, you pay the compute cost for a zero quality pay the compute cost for a zero quality pay the compute cost for a zero quality lift from this image. And secondly, uh lift from this image. And secondly, uh lift from this image. And secondly, uh there is a risk of degrading this image there is a risk of degrading this image there is a risk of degrading this image given it's already such a high quality given it's already such a high quality given it's already such a high quality image. And on the other end of the spectrum, And on the other end of the spectrum, you have a recall miss. So on your left you have a recall miss. So on your left you have a recall miss. So on your left you have an image with six chicken wings you have an image with six chicken wings you have an image with six chicken wings and on your right if you notice the dish and on your right if you notice the dish and on your right if you notice the dish name, it says eight pieces chicken name, it says eight pieces chicken name, it says eight pieces chicken wings.

  10. wings. wings. And your routing agent approved this And your routing agent approved this And your routing agent approved this image. That means So now there's a risk image. That means So now there's a risk image. That means So now there's a risk here if you send up send this image for here if you send up send this image for here if you send up send this image for enhancement and you only see six chicken enhancement and you only see six chicken enhancement and you only see six chicken wings, there's a chance your model's wings, there's a chance your model's wings, there's a chance your model's going to hallucinate these two extra going to hallucinate these two extra going to hallucinate these two extra wings to to match the description. wings to to match the description. wings to to match the description. And that's also an that's a the And that's also an that's a the And that's also an that's a the cut we take at our faithfulness metric cut we take at our faithfulness metric cut we take at our faithfulness metric that Jay earlier showed us. that Jay earlier showed us. that Jay earlier showed us. So the meta point I'm trying to get here So the meta point I'm trying to get here So the meta point I'm trying to get here is you've trained your offline model, is you've trained your offline model, is you've trained your offline model, but there will be long cases where your but there will be long cases where your but there will be long cases where your model is going to continue to fail and model is going to continue to fail and model is going to continue to fail and the static model will not work in the the static model will not work in the the static model will not work in the real system. You need a way such that real system. You need a way such that real system. You need a way such that your prompts, agents, system itself is your prompts, agents, system itself is your prompts, agents, system itself is evolving over time. evolving over time. evolving over time. And that's what we've done And that's what we've done And that's what we've done uh for our system as well. And I'm uh for our system as well. And I'm uh for our system as well. And I'm talking more from the routing talking more from the routing talking more from the routing perspective, but every component in our perspective, but every component in our perspective, but every component in our system is able to tune itself uh for any system is able to tune itself uh for any system is able to tune itself uh for any drift online. drift online. drift online. So what we do is we sample production So what we do is we sample production So what we do is we sample production data at regular cadence, data at regular cadence, data at regular cadence, uh send this to the human labelers with uh send this to the human labelers with uh send this to the human labelers with the same guidelines that we have seen the same guidelines that we have seen the same guidelines that we have seen before. Once you've got that data, we before. Once you've got that data, we before. Once you've got that data, we compare our agents' output with the compare our agents' output with the compare our agents' output with the output we got from the labelers and see output we got from the labelers and see output we got from the labelers and see if there's a mismatch. If there's a if there's a mismatch. If there's a if there's a mismatch. If there's a mismatch, we have an umbrella diagnosis mismatch, we have an umbrella diagnosis mismatch, we have an umbrella diagnosis agent which takes in the feedback, agent which takes in the feedback, agent which takes in the feedback, localizes where this issue is happening, localizes where this issue is happening, localizes where this issue is happening, and and and triggers our auto-tuning and and and triggers our auto-tuning and and and triggers our auto-tuning pipeline.

  11. pipeline. pipeline. Once we tune this agent, we go and Once we tune this agent, we go and Once we tune this agent, we go and benchmark it against our golden data set benchmark it against our golden data set benchmark it against our golden data set that we saw earlier, and if we pass our that we saw earlier, and if we pass our that we saw earlier, and if we pass our golden data set on the metrics that we golden data set on the metrics that we golden data set on the metrics that we had designed, we go ahead and ship this had designed, we go ahead and ship this had designed, we go ahead and ship this model. model. model. Uh if not, then you kind of keep Uh if not, then you kind of keep Uh if not, then you kind of keep iterating. And this happens on a regular iterating. And this happens on a regular iterating. And this happens on a regular basis on production data set. basis on production data set. basis on production data set. Um the beauty of this is this is Um the beauty of this is this is Um the beauty of this is this is completely config driven and doesn't completely config driven and doesn't completely config driven and doesn't require human in the loop. Your require human in the loop. Your require human in the loop. Your diagnoser agent can write your config diagnoser agent can write your config diagnoser agent can write your config and trigger the auto-tuning pipeline and trigger the auto-tuning pipeline and trigger the auto-tuning pipeline here. And this is what will keep your here. And this is what will keep your here. And this is what will keep your model sharp over time. You will have one model sharp over time. You will have one model sharp over time. You will have one static model with the offline, but this static model with the offline, but this static model with the offline, but this is what is going to keep your system is what is going to keep your system is what is going to keep your system alive. Um so Jay is going to spend more time on Um so Jay is going to spend more time on the diagnosis side of it. What I want to the diagnosis side of it. What I want to the diagnosis side of it. What I want to do is zoom into the auto-tuning bit. And do is zoom into the auto-tuning bit. And do is zoom into the auto-tuning bit. And again, we're looking at routing, but again, we're looking at routing, but again, we're looking at routing, but this is how we tune every agent in our this is how we tune every agent in our this is how we tune every agent in our system. system. system. Uh so we start with a target agent, and Uh so we start with a target agent, and Uh so we start with a target agent, and we've already got these uh unseen eval we've already got these uh unseen eval we've already got these uh unseen eval samples from our humans. samples from our humans. samples from our humans. We go find out the mismatch and matches We go find out the mismatch and matches We go find out the mismatch and matches and call a prompt optimizer agent. Now, and call a prompt optimizer agent. Now, and call a prompt optimizer agent. Now, this itself is two sub-agents. There's this itself is two sub-agents. There's this itself is two sub-agents. There's the reflect agent and the up synthesize the reflect agent and the up synthesize the reflect agent and the up synthesize agent. What reflect does is it it just agent. What reflect does is it it just agent. What reflect does is it it just looks at the mismatches, tries to find looks at the mismatches, tries to find looks at the mismatches, tries to find remove any noise, find any systemic remove any noise, find any systemic remove any noise, find any systemic issues that might be in your data set, issues that might be in your data set, issues that might be in your data set, and and and reflect on it and send that feedback to reflect on it and send that feedback to reflect on it and send that feedback to the synthesize agent. Now, the the synthesize agent. Now, the the synthesize agent. Now, the synthesize agent takes this feedback. It synthesize agent takes this feedback. It synthesize agent takes this feedback. It has your agent config. It goes and has your agent config. It goes and has your agent config. It goes and updates your agent with the new config updates your agent with the new config updates your agent with the new config based on the feedback it's getting. And based on the feedback it's getting. And based on the feedback it's getting. And goes and benchmarks again. If this goes and benchmarks again. If this goes and benchmarks again. If this benchmark is passed, you actually benchmark is passed, you actually benchmark is passed, you actually register this new agent in the new agent register this new agent in the new agent register this new agent in the new agent store. And next time your production

  12. store. And next time your production store. And next time your production runs, you pick up the new version of the runs, you pick up the new version of the runs, you pick up the new version of the agent. agent. agent. And this is a closed-loop system as I And this is a closed-loop system as I And this is a closed-loop system as I mentioned, no human in the loop. We mentioned, no human in the loop. We mentioned, no human in the loop. We definitely have observability on the definitely have observability on the definitely have observability on the guardrails, quick rollback built in in guardrails, quick rollback built in in guardrails, quick rollback built in in case of any issues with the system case of any issues with the system case of any issues with the system itself. Moving on to the next step of our Moving on to the next step of our orchestration flow. So we spoke about orchestration flow. So we spoke about orchestration flow. So we spoke about routing, moving on to the enhancement routing, moving on to the enhancement routing, moving on to the enhancement bit of it. It's a three-step process. bit of it. It's a three-step process. bit of it. It's a three-step process. What we do is the first step, we What we do is the first step, we What we do is the first step, we generate a prompt specific to this generate a prompt specific to this generate a prompt specific to this image. We take in the description, we image. We take in the description, we image. We take in the description, we take in the directives we were getting take in the directives we were getting take in the directives we were getting from our routing agent, and we go ahead from our routing agent, and we go ahead from our routing agent, and we go ahead and generate a prompt for this image. and generate a prompt for this image. and generate a prompt for this image. What needs improvement in this image What needs improvement in this image What needs improvement in this image specifically? And we go ahead and specifically? And we go ahead and specifically? And we go ahead and enhance this image. Then you've got the enhance this image. Then you've got the enhance this image. Then you've got the QA gate, which is a multi-dimensional QA gate, which is a multi-dimensional QA gate, which is a multi-dimensional gate, looks at multiple things like gate, looks at multiple things like gate, looks at multiple things like plating, faithfulness, colors. And if it plating, faithfulness, colors. And if it plating, faithfulness, colors. And if it passes is when you actually go ahead and passes is when you actually go ahead and passes is when you actually go ahead and publish this. If it doesn't pass, you publish this. If it doesn't pass, you publish this. If it doesn't pass, you take the feedback back from the QA gate, take the feedback back from the QA gate, take the feedback back from the QA gate, push it back to your generate prompt push it back to your generate prompt push it back to your generate prompt along with the initial inputs you sent along with the initial inputs you sent along with the initial inputs you sent it, and go ahead and enhance it again. it, and go ahead and enhance it again. it, and go ahead and enhance it again. So, there's two end results here. You So, there's two end results here. You So, there's two end results here. You either keep enhancing for K iterations either keep enhancing for K iterations either keep enhancing for K iterations and you pass your QA gate and you and you pass your QA gate and you and you pass your QA gate and you publish, or you take a coverage hit and publish, or you take a coverage hit and publish, or you take a coverage hit and you never enhance this image.

  13. Here's an example. On your left, you see Here's an example. On your left, you see a bowl of sweet potato fries. We send it a bowl of sweet potato fries. We send it a bowl of sweet potato fries. We send it up for the first iteration and our QA up for the first iteration and our QA up for the first iteration and our QA agent rejects it because the portion agent rejects it because the portion agent rejects it because the portion size is incorrect, the plating is very size is incorrect, the plating is very size is incorrect, the plating is very unrealistic. We take that feedback in, unrealistic. We take that feedback in, unrealistic. We take that feedback in, go for the second iteration, and we're go for the second iteration, and we're go for the second iteration, and we're actually able to pass it the second actually able to pass it the second actually able to pass it the second iteration. So, the metric we are iteration. So, the metric we are iteration. So, the metric we are measuring here is pass at K. Pass at K measuring here is pass at K. Pass at K measuring here is pass at K. Pass at K is essentially the pass rate at Kth is essentially the pass rate at Kth is essentially the pass rate at Kth iteration. And ideally with the more the iteration. And ideally with the more the iteration. And ideally with the more the iterations, your pass rate will increase iterations, your pass rate will increase iterations, your pass rate will increase because you're getting more feedback in. because you're getting more feedback in. because you're getting more feedback in. Now, I'll pass it on back to Jay to Now, I'll pass it on back to Jay to Now, I'll pass it on back to Jay to cover the rest of this. >> Thanks Thanks, Somya. >> Thanks Thanks, Somya. Um Um Um So, yeah, just before we end here on the So, yeah, just before we end here on the So, yeah, just before we end here on the um on on the generation of the vowels, um on on the generation of the vowels, um on on the generation of the vowels, we use what's called pairwise we use what's called pairwise we use what's called pairwise comparison, comparison, comparison, right, for our pass at K. So, it's right, for our pass at K. So, it's right, for our pass at K. So, it's looking at the input image and the the looking at the input image and the the looking at the input image and the the output image, and it's assessing whether output image, and it's assessing whether output image, and it's assessing whether or not it's better. or not it's better. or not it's better. But how do we actually find what's But how do we actually find what's But how do we actually find what's better? So, better? So, better? So, um we're not going to dive into too much um we're not going to dive into too much um we're not going to dive into too much of the details here cuz this is kind of of the details here cuz this is kind of of the details here cuz this is kind of like proprietary stuff, and so we'll like proprietary stuff, and so we'll like proprietary stuff, and so we'll just mention it at a higher level that just mention it at a higher level that just mention it at a higher level that this is where you sort of For least for this is where you sort of For least for this is where you sort of For least for us at Rue Ba, we have to make sure that us at Rue Ba, we have to make sure that us at Rue Ba, we have to make sure that we're aligning with product, design, we're aligning with product, design, we're aligning with product, design, policy, legal. And this is where we're policy, legal. And this is where we're policy, legal. And this is where we're baking in what we define as a better baking in what we define as a better baking in what we define as a better image on the platform into our Evals.

  14. image on the platform into our Evals. image on the platform into our Evals. Um so, examples here, is it faithful? Is Um so, examples here, is it faithful? Is Um so, examples here, is it faithful? Is it complete? Is it natural? Is it it complete? Is it natural? Is it it complete? Is it natural? Is it realistic? And there's a bunch of other realistic? And there's a bunch of other realistic? And there's a bunch of other things as well. The output of this is things as well. The output of this is things as well. The output of this is then uh a yes, no, or unsure. then uh a yes, no, or unsure. then uh a yes, no, or unsure. So, here are some examples of failure So, here are some examples of failure So, here are some examples of failure modes. modes. modes. So, input and output on the right. The So, input and output on the right. The So, input and output on the right. The inputs on the left, outputs on the inputs on the left, outputs on the inputs on the left, outputs on the right-hand side. This might be a little right-hand side. This might be a little right-hand side. This might be a little bit uh difficult to to see at the first bit uh difficult to to see at the first bit uh difficult to to see at the first pass. We actually added shrimp here and pass. We actually added shrimp here and pass. We actually added shrimp here and we shouldn't be. So, we failed we shouldn't be. So, we failed we shouldn't be. So, we failed faithfulness. faithfulness. faithfulness. This is where we go the other way. This is where we go the other way. This is where we go the other way. So, the input um So, the input um So, the input um has some source at the bottom of the has some source at the bottom of the has some source at the bottom of the sushi. We actually remove it. sushi. We actually remove it. sushi. We actually remove it. So, we failed completeness. So, we failed completeness. So, we failed completeness. Here's actually a pretty interesting Here's actually a pretty interesting Here's actually a pretty interesting example where the agent actually example where the agent actually example where the agent actually attempted a more creative edit in the attempted a more creative edit in the attempted a more creative edit in the first iteration. first iteration. first iteration. Um and then the QA said, "Nope, it's not Um and then the QA said, "Nope, it's not Um and then the QA said, "Nope, it's not good enough." good enough." good enough." Uh and then it actually oversteers the Uh and then it actually oversteers the Uh and then it actually oversteers the other other way. other other way. other other way. And it becomes overly conservative. And it becomes overly conservative. And it becomes overly conservative. Sort of falls back to this generic Sort of falls back to this generic Sort of falls back to this generic ceramic plate uh ceramic bowl, sorry.

  15. ceramic plate uh ceramic bowl, sorry. ceramic plate uh ceramic bowl, sorry. So, this is an example of a reward So, this is an example of a reward So, this is an example of a reward hacking actually. And and this is a hacking actually. And and this is a hacking actually. And and this is a nugatory change, but something that we nugatory change, but something that we nugatory change, but something that we don't think is a meaningful or don't think is a meaningful or don't think is a meaningful or influential change despite the actual influential change despite the actual influential change despite the actual raw pixels of the input and output being raw pixels of the input and output being raw pixels of the input and output being pretty different. pretty different. pretty different. Here's another example where in the Here's another example where in the Here's another example where in the output the plate is covering the sauce. output the plate is covering the sauce. output the plate is covering the sauce. This is an example where for some some This is an example where for some some This is an example where for some some of the frontier models that we're using of the frontier models that we're using of the frontier models that we're using for the actual image editing, some of for the actual image editing, some of for the actual image editing, some of their um their um their um some of their problems will actually some of their problems will actually some of their problems will actually sort of leak up into our applied use sort of leak up into our applied use sort of leak up into our applied use case. case. case. Um and so, so object coherence and Um and so, so object coherence and Um and so, so object coherence and physics plausibility of the Evals that physics plausibility of the Evals that physics plausibility of the Evals that sometimes will coordinate with the sometimes will coordinate with the sometimes will coordinate with the frontier teams and and let them know frontier teams and and let them know frontier teams and and let them know about these problems and work together about these problems and work together about these problems and work together with them. with them. with them. Here's uh an example of why Here's uh an example of why Here's uh an example of why multimodality is is pretty important. In multimodality is is pretty important. In multimodality is is pretty important. In the input and the output, we we can't the input and the output, we we can't the input and the output, we we can't actually see that there are eight pieces actually see that there are eight pieces actually see that there are eight pieces here of of the wontons. here of of the wontons. here of of the wontons. So, we're not confident, actually. We're So, we're not confident, actually. We're So, we're not confident, actually. We're not sure. And so, this is an example not sure. And so, this is an example not sure. And so, this is an example where we would actually reject it in where we would actually reject it in where we would actually reject it in production and and it wouldn't it production and and it wouldn't it production and and it wouldn't it wouldn't go through.

  16. So, the last step after all of that is a So, the last step after all of that is a post-processing and what we refer to as post-processing and what we refer to as post-processing and what we refer to as the publish-ready QA. This is the final the publish-ready QA. This is the final the publish-ready QA. This is the final gate before we decide we want to publish gate before we decide we want to publish gate before we decide we want to publish something to production. Here, we do some policy checks. Here, we do some policy checks. We also do some more quality checks. We also do some more quality checks. We also do some more quality checks. Um and you might be wondering, like, Um and you might be wondering, like, Um and you might be wondering, like, we've already done some QA. Like, why we've already done some QA. Like, why we've already done some QA. Like, why are we going to do another step of QA? are we going to do another step of QA? are we going to do another step of QA? The reason is because we think of this The reason is because we think of this The reason is because we think of this like a Swiss cheese model. like a Swiss cheese model. like a Swiss cheese model. So, we want to try and optimize for So, we want to try and optimize for So, we want to try and optimize for reducing the chance of a failure getting reducing the chance of a failure getting reducing the chance of a failure getting into production. And so, there is some into production. And so, there is some into production. And so, there is some redundancy here or there. redundancy here or there. redundancy here or there. And that's okay. And that's okay. And that's okay. Um and so, this QA gate is is a little Um and so, this QA gate is is a little Um and so, this QA gate is is a little bit more holistic. It captures more bit more holistic. It captures more bit more holistic. It captures more things. But, it also will will try and things. But, it also will will try and things. But, it also will will try and flag things that we should have caught flag things that we should have caught flag things that we should have caught upstream, as well. All right. So, we've talked about All right. So, we've talked about a couple of uh a couple of uh a couple of uh feedback loops here. So, to summarize, feedback loops here. So, to summarize, feedback loops here. So, to summarize, we talked about predominantly this first we talked about predominantly this first we talked about predominantly this first one here, which is the model loop. And one here, which is the model loop. And one here, which is the model loop. And this is accounting for drifts and this is accounting for drifts and this is accounting for drifts and aligning with human labeled data set aligning with human labeled data set aligning with human labeled data set that we have and we've established that we have and we've established that we have and we've established offline.

  17. offline. offline. But, we actually have more feedback But, we actually have more feedback But, we actually have more feedback loops. loops. loops. So, we we have that Uber what we we have So, we we have that Uber what we we have So, we we have that Uber what we we have is a is a great sort of dog dog feeding is a is a great sort of dog dog feeding is a is a great sort of dog dog feeding culture. culture. culture. Um we will test apps before they go Um we will test apps before they go Um we will test apps before they go live. Um but, we also have when it goes live. Um but, we also have when it goes live. Um but, we also have when it goes live in production, how do we get that live in production, how do we get that live in production, how do we get that feedback back into our agent to be able feedback back into our agent to be able feedback back into our agent to be able to steer it appropriately? to steer it appropriately? to steer it appropriately? So, as we're adding more of these So, as we're adding more of these So, as we're adding more of these feedback loops, we want to be able to feedback loops, we want to be able to feedback loops, we want to be able to generalize the system. generalize the system. generalize the system. So, this is where we've actually created So, this is where we've actually created So, this is where we've actually created um a higher level of abstraction on top, um a higher level of abstraction on top, um a higher level of abstraction on top, which we call the diagnoser. which we call the diagnoser. which we call the diagnoser. So, the diagnoser can take in any input So, the diagnoser can take in any input So, the diagnoser can take in any input from these different feedback loops that from these different feedback loops that from these different feedback loops that we're we're capturing. It can reflect on we're we're capturing. It can reflect on we're we're capturing. It can reflect on what actual agent within the overall what actual agent within the overall what actual agent within the overall system needs to be optimized, and it can system needs to be optimized, and it can system needs to be optimized, and it can route that agent to be able to fix that route that agent to be able to fix that route that agent to be able to fix that configuration specifically. It could be configuration specifically. It could be configuration specifically. It could be one agent, it could be multiple agents. So, here's an example of internal dog So, here's an example of internal dog fooding. You might see these in sort of fooding. You might see these in sort of fooding. You might see these in sort of different apps that you've got where you different apps that you've got where you different apps that you've got where you got the thumbs down and the thumbs up.

  18. got the thumbs down and the thumbs up. got the thumbs down and the thumbs up. We also take some free form feedback as We also take some free form feedback as We also take some free form feedback as well. well. well. Uh and this is actually great cuz we'll Uh and this is actually great cuz we'll Uh and this is actually great cuz we'll get feedback from merchants directly. get feedback from merchants directly. get feedback from merchants directly. We'll get feedback from, you know, We'll get feedback from, you know, We'll get feedback from, you know, design teams, other product teams uh at design teams, other product teams uh at design teams, other product teams uh at Uber. And we'll incorporate that Uber. And we'll incorporate that Uber. And we'll incorporate that feedback back into our diagnosis step feedback back into our diagnosis step feedback back into our diagnosis step and tune the system over time. and tune the system over time. and tune the system over time. Again, similar sort of workflow pattern Again, similar sort of workflow pattern Again, similar sort of workflow pattern here. We'll replay the examples that we here. We'll replay the examples that we here. We'll replay the examples that we know are those ones that have been know are those ones that have been know are those ones that have been flagged, be it good examples, be it bad flagged, be it good examples, be it bad flagged, be it good examples, be it bad examples, uh and then we'll benchmark examples, uh and then we'll benchmark examples, uh and then we'll benchmark the metrics before we push the latest the metrics before we push the latest the metrics before we push the latest config version. config version. config version. The last step is is actually getting The last step is is actually getting The last step is is actually getting this into production. this into production. this into production. And and this is where we're looking for And and this is where we're looking for And and this is where we're looking for a whole heap of different metrics we a whole heap of different metrics we a whole heap of different metrics we track for for the marketplace quality track for for the marketplace quality track for for the marketplace quality and health. Uh I've just called out one and health. Uh I've just called out one and health. Uh I've just called out one here, which is conversion. So, we're here, which is conversion. So, we're here, which is conversion. So, we're looking for improvements in people looking for improvements in people looking for improvements in people adding to cart, converting, completing adding to cart, converting, completing adding to cart, converting, completing their orders. their orders. their orders. Um I think this one's actually an Um I think this one's actually an Um I think this one's actually an interesting one to call out because now interesting one to call out because now interesting one to call out because now at I mean, at least at Uber, but at I mean, at least at Uber, but at I mean, at least at Uber, but especially in production um settings at especially in production um settings at especially in production um settings at scale, you have a wide um scale, you have a wide um scale, you have a wide um uh uh uh you have a lot of data that you can you have a lot of data that you can you have a lot of data that you can actually slice and dice.

  19. actually slice and dice. actually slice and dice. So, in this area as opposed to the So, in this area as opposed to the So, in this area as opposed to the others, what we can do is sort of slice others, what we can do is sort of slice others, what we can do is sort of slice by geos, by device type, by dish type, by geos, by device type, by dish type, by geos, by device type, by dish type, etc. And we can look at where things are etc. And we can look at where things are etc. And we can look at where things are improving in different segments and improving in different segments and improving in different segments and actually tune on certain segments as actually tune on certain segments as actually tune on certain segments as well. Cool, and that's it for our Cool, and that's it for our presentation. Appreciate it. presentation. Appreciate it. presentation. Appreciate it. >> [applause]

Summary

The main theme is using computer vision for real-world production at Uber Eats, focusing on improving merchant photos. Key subjects include the delivery marketplace's scale, the importance of visual content, and the challenges faced by smaller merchants with time, know-how, and cost for professional photography. The practical takeaway is the need to ethically leverage AI to enhance visuals while maintaining authenticity and customer trust, as many consumers distrust AI-generated content.

View original episode ↗