← Back
AI Engineer September 1, 2026 21m

Multimodal Collaborative Agents for Next-Gen Commerce — Nidhi Kaushik Vyas, Google DeepMind

Read full transcript 16 segments
  1. >> Okay, good morning folks. Thank you for >> Okay, good morning folks. Thank you for showing up. I'm Nidhi. I'm a product showing up. I'm Nidhi. I'm a product showing up. I'm Nidhi. I'm a product person at Google DeepMind. And today person at Google DeepMind. And today person at Google DeepMind. And today I'll be talking about multimodal I'll be talking about multimodal I'll be talking about multimodal collaborative agents. Typically, these collaborative agents. Typically, these collaborative agents. Typically, these are agents that work with fuzzy intent are agents that work with fuzzy intent are agents that work with fuzzy intent intent when the intent when it's not intent when the intent when it's not intent when the intent when it's not clear the user hasn't gotten the right clear the user hasn't gotten the right clear the user hasn't gotten the right keywords to specify and when they come keywords to specify and when they come keywords to specify and when they come in for a shopping intent, how can you in for a shopping intent, how can you in for a shopping intent, how can you have the agents still guide the user have the agents still guide the user have the agents still guide the user towards their goal with high agency towards their goal with high agency towards their goal with high agency execution, proactive elicitation, and a execution, proactive elicitation, and a execution, proactive elicitation, and a lot of hand-holding. lot of hand-holding. lot of hand-holding. The framework that we'll be discussing The framework that we'll be discussing The framework that we'll be discussing today are grounded in shopping or today are grounded in shopping or today are grounded in shopping or commerce because I wanted to show you a commerce because I wanted to show you a commerce because I wanted to show you a few patterns that are much easier to see few patterns that are much easier to see few patterns that are much easier to see in commerce, but they're quite in commerce, but they're quite in commerce, but they're quite applicable in other consumer verticals applicable in other consumer verticals applicable in other consumer verticals as well, finance, education, or whatever as well, finance, education, or whatever as well, finance, education, or whatever you guys work on. you guys work on. you guys work on. Um Um Um Yeah, so before we even start into the Yeah, so before we even start into the Yeah, so before we even start into the frameworks, I wanted to discuss about frameworks, I wanted to discuss about frameworks, I wanted to discuss about why this is an existing problem. why this is an existing problem. why this is an existing problem. Currently, a lot of the agents that we Currently, a lot of the agents that we Currently, a lot of the agents that we have act more like a wrapper to the have act more like a wrapper to the have act more like a wrapper to the search bar. search bar. search bar. They assume that the user has a They assume that the user has a They assume that the user has a well-defined intent, has the right well-defined intent, has the right well-defined intent, has the right keywords, already knows what they're keywords, already knows what they're keywords, already knows what they're looking for. And so when they come in, looking for. And so when they come in, looking for. And so when they come in, they just have the right vocabulary to they just have the right vocabulary to they just have the right vocabulary to interact with the agent. However, there interact with the agent. However, there interact with the agent. However, there is quite a huge articulation gap. When is quite a huge articulation gap. When is quite a huge articulation gap. When the users come in, rarely they have the users come in, rarely they have the users come in, rarely they have their intent well formed. Rather, they their intent well formed. Rather, they their intent well formed. Rather, they kind of have a fuzzy feeling or a vibe kind of have a fuzzy feeling or a vibe kind of have a fuzzy feeling or a vibe when they are kind of looking to shop.

  2. when they are kind of looking to shop. when they are kind of looking to shop. So, So, So, the agent needs to play quite a huge the agent needs to play quite a huge the agent needs to play quite a huge role in hand-holding them. They need to role in hand-holding them. They need to role in hand-holding them. They need to work with the user in first work with the user in first work with the user in first understanding their preferences or even understanding their preferences or even understanding their preferences or even eliciting these preferences more eliciting these preferences more eliciting these preferences more proactively. proactively. proactively. And then helping the user And then helping the user And then helping the user kind of show them different kind of show them different kind of show them different possibilities of what could be possible possibilities of what could be possible possibilities of what could be possible that they can shop that they can shop that they can shop and then eventually move towards and then eventually move towards and then eventually move towards kind of recommendations that work with kind of recommendations that work with kind of recommendations that work with these constraints that they have in these constraints that they have in these constraints that they have in mind. So, what we're going to be mind. So, what we're going to be mind. So, what we're going to be discussing is this kind of a flywheel or discussing is this kind of a flywheel or discussing is this kind of a flywheel or a loop that goes from very, very fuzzy a loop that goes from very, very fuzzy a loop that goes from very, very fuzzy intent all the way towards kind of intent all the way towards kind of intent all the way towards kind of achieving the user's goal. achieving the user's goal. achieving the user's goal. So, this is how the shopping loop looks So, this is how the shopping loop looks So, this is how the shopping loop looks like right now. The very first thing like right now. The very first thing like right now. The very first thing that the agent prepares for when the that the agent prepares for when the that the agent prepares for when the user comes in is what we call the user comes in is what we call the user comes in is what we call the discovery phase. discovery phase. discovery phase. This is where the where the agent takes This is where the where the agent takes This is where the where the agent takes all the different contextual information all the different contextual information all the different contextual information that they have about the user. This that they have about the user. This that they have about the user. This could be past conversations, what the could be past conversations, what the could be past conversations, what the user has specified in their query. It user has specified in their query. It user has specified in their query. It could be information present in their could be information present in their could be information present in their personal context, or even the references personal context, or even the references personal context, or even the references that the user have provided. And then that the user have provided. And then that the user have provided. And then comes up with a collaborative strategy, comes up with a collaborative strategy, comes up with a collaborative strategy, and we'll be going to the details of and we'll be going to the details of and we'll be going to the details of this, but it comes up with a this, but it comes up with a this, but it comes up with a collaborative strategy on what more does collaborative strategy on what more does collaborative strategy on what more does the agent need to elicitate back from the agent need to elicitate back from the agent need to elicitate back from the user in order to help get the intent the user in order to help get the intent the user in order to help get the intent in a better shape, to get clarity on in a better shape, to get clarity on in a better shape, to get clarity on what exactly the user is looking for.

  3. what exactly the user is looking for. what exactly the user is looking for. And then it moves to the second phase, And then it moves to the second phase, And then it moves to the second phase, which is what we call the research which is what we call the research which is what we call the research phase. phase. phase. Which is again a two-step process. First Which is again a two-step process. First Which is again a two-step process. First is where is where is where it learns what is the best way to it learns what is the best way to it learns what is the best way to elicitate this preference. So, for elicitate this preference. So, for elicitate this preference. So, for example, a lot of times the text the example, a lot of times the text the example, a lot of times the text the text base elicitation may not be the text base elicitation may not be the text base elicitation may not be the best way to get the preferences from the best way to get the preferences from the best way to get the preferences from the user because sometimes a user don't know user because sometimes a user don't know user because sometimes a user don't know what they're looking for. So, then how what they're looking for. So, then how what they're looking for. So, then how can the agent start being more creative can the agent start being more creative can the agent start being more creative in terms of elicitating these in terms of elicitating these in terms of elicitating these preferences? Can they start using some preferences? Can they start using some preferences? Can they start using some kind of visual references or visual kind of visual references or visual kind of visual references or visual visual inspiration boards for grounding visual inspiration boards for grounding visual inspiration boards for grounding and speaking a common language with the and speaking a common language with the and speaking a common language with the user? And then the second phase to this user? And then the second phase to this user? And then the second phase to this is also coming up going into the is also coming up going into the is also coming up going into the background and doing the heavy lifting background and doing the heavy lifting background and doing the heavy lifting for the user, so taking the burden away for the user, so taking the burden away for the user, so taking the burden away from the user to describing what they from the user to describing what they from the user to describing what they want, and rather going into the want, and rather going into the want, and rather going into the background and doing all the background and doing all the background and doing all the comparisons, trade-offs, comparisons, trade-offs, comparisons, trade-offs, summarization of all the information summarization of all the information summarization of all the information that they're looking for, and coming that they're looking for, and coming that they're looking for, and coming back with the right set of potential back with the right set of potential back with the right set of potential options for the user. And then the third options for the user. And then the third options for the user. And then the third phase, once um phase, once um phase, once um the user and the agent is ready to go the user and the agent is ready to go the user and the agent is ready to go into some kind of a into some kind of a into some kind of a is ready to go into the last phase, is ready to go into the last phase, is ready to go into the last phase, which is where um which is where um which is where um the agent is ready to give out a the agent is ready to give out a the agent is ready to give out a response. This is where the agent needs response. This is where the agent needs response. This is where the agent needs to start adapting the response in a way to start adapting the response in a way to start adapting the response in a way that is most useful for the user. So, that is most useful for the user. So, that is most useful for the user. So, typically a lot of systems fail here typically a lot of systems fail here typically a lot of systems fail here where they just give out a text-heavy where they just give out a text-heavy where they just give out a text-heavy response. What the agent should rather response. What the agent should rather response. What the agent should rather be doing is adapting to the query that be doing is adapting to the query that be doing is adapting to the query that the user had in mind. So, should it be the user had in mind. So, should it be the user had in mind. So, should it be using um some kind of a bulleted list?

  4. using um some kind of a bulleted list? using um some kind of a bulleted list? Should it be using comparison tables? Should it be using comparison tables? Should it be using comparison tables? Should it be using um visual boards? So, Should it be using um visual boards? So, Should it be using um visual boards? So, this is where the there's adaptive this is where the there's adaptive this is where the there's adaptive response happening as well, where the uh response happening as well, where the uh response happening as well, where the uh agent starts to develop uh more more of agent starts to develop uh more more of agent starts to develop uh more more of a uh smarter uh presentation skills uh a uh smarter uh presentation skills uh a uh smarter uh presentation skills uh for the user to really find the answer for the user to really find the answer for the user to really find the answer that they're looking for. So, we'll be that they're looking for. So, we'll be that they're looking for. So, we'll be diving into all these topics as I uh diving into all these topics as I uh diving into all these topics as I uh kind of work through the presentation. kind of work through the presentation. kind of work through the presentation. So, first and foremost, we're looking at So, first and foremost, we're looking at So, first and foremost, we're looking at discovery. Like I was mentioning, this discovery. Like I was mentioning, this discovery. Like I was mentioning, this is where uh the the the agent needs to is where uh the the the agent needs to is where uh the the the agent needs to remember what exactly matters. remember what exactly matters. remember what exactly matters. Um you the agent starts to look at the Um you the agent starts to look at the Um you the agent starts to look at the context across a bunch of different context across a bunch of different context across a bunch of different signals. It starts to look at past signals. It starts to look at past signals. It starts to look at past conversations. It looks at some of the conversations. It looks at some of the conversations. It looks at some of the reference images or reference links that reference images or reference links that reference images or reference links that the user might have provided. It also the user might have provided. It also the user might have provided. It also starts to look at personal context and starts to look at personal context and starts to look at personal context and starts to build out a working state. So, starts to build out a working state. So, starts to build out a working state. So, if you look at the sample code that we if you look at the sample code that we if you look at the sample code that we have here, like there is there is a goal have here, like there is there is a goal have here, like there is there is a goal uh one of the let's let's work with this uh one of the let's let's work with this uh one of the let's let's work with this query where the user is trying to redo query where the user is trying to redo query where the user is trying to redo their living room with a certain budget their living room with a certain budget their living room with a certain budget in mind. And some of the um some of the in mind. And some of the um some of the in mind. And some of the um some of the things that the agent develops as part things that the agent develops as part things that the agent develops as part of the working state is the session of the working state is the session of the working state is the session history. It has a user context. It also history. It has a user context. It also history. It has a user context. It also kind of extracts out the hard constraint kind of extracts out the hard constraint kind of extracts out the hard constraint that the user might have provided in the that the user might have provided in the that the user might have provided in the query.

  5. query. query. But, things start getting getting But, things start getting getting But, things start getting getting interesting when we come to the softer interesting when we come to the softer interesting when we come to the softer constraints. constraints. constraints. So, this is where the user may not be So, this is where the user may not be So, this is where the user may not be able to describe what they're looking able to describe what they're looking able to describe what they're looking for and may might have provided like a for and may might have provided like a for and may might have provided like a reference image on, you know, an reference image on, you know, an reference image on, you know, an inspiration that they had in mind or inspiration that they had in mind or inspiration that they had in mind or some kind of a layout design that they some kind of a layout design that they some kind of a layout design that they really liked and is trying to get to the really liked and is trying to get to the really liked and is trying to get to the agent in terms of uh this is what speaks agent in terms of uh this is what speaks agent in terms of uh this is what speaks more to me. So, this is where the agent more to me. So, this is where the agent more to me. So, this is where the agent starts to be more proactive and pulls starts to be more proactive and pulls starts to be more proactive and pulls out some of the salient signals from the out some of the salient signals from the out some of the salient signals from the from the reference images and starts to from the reference images and starts to from the reference images and starts to develop a mental model of what the user develop a mental model of what the user develop a mental model of what the user might really be looking to get at. So, might really be looking to get at. So, might really be looking to get at. So, here the user is uh so the agent has here the user is uh so the agent has here the user is uh so the agent has specified uh has identified that maybe specified uh has identified that maybe specified uh has identified that maybe the style that is working out well for the style that is working out well for the style that is working out well for them is of a certain kind. It also uh is them is of a certain kind. It also uh is them is of a certain kind. It also uh is also is also starting to work towards a also is also starting to work towards a also is also starting to work towards a confidence score on like how confident confidence score on like how confident confidence score on like how confident it is in terms of pulling out some of it is in terms of pulling out some of it is in terms of pulling out some of this information present in the this information present in the this information present in the multimodal inputs provided. And then the multimodal inputs provided. And then the multimodal inputs provided. And then the last thing uh that happens as part of last thing uh that happens as part of last thing uh that happens as part of developing this working state is also developing this working state is also developing this working state is also figuring out what are the variables that figuring out what are the variables that figuring out what are the variables that the agent needs to pull out almost in the agent needs to pull out almost in the agent needs to pull out almost in real time because these variables real time because these variables real time because these variables uh this variables do affect how the uh this variables do affect how the uh this variables do affect how the results will be displayed back to the results will be displayed back to the results will be displayed back to the user. This could be variables that need user. This could be variables that need user. This could be variables that need to be refreshed in real time like uh to be refreshed in real time like uh to be refreshed in real time like uh inventory because if if what you're inventory because if if what you're inventory because if if what you're providing back to the user is stale providing back to the user is stale providing back to the user is stale information, then it's kind of a moot information, then it's kind of a moot information, then it's kind of a moot point. So, these are the uh these are point. So, these are the uh these are point. So, these are the uh these are the variables that you want to refresh the variables that you want to refresh the variables that you want to refresh in real time and is and becomes a part in real time and is and becomes a part in real time and is and becomes a part of the agent's working state.

  6. of the agent's working state. of the agent's working state. Um Um Um few ways that you can evaluate this few ways that you can evaluate this few ways that you can evaluate this state uh the way we have developed our state uh the way we have developed our state uh the way we have developed our auto raters, we we do make sure that all auto raters, we we do make sure that all auto raters, we we do make sure that all facts are retained meaning that whatever facts are retained meaning that whatever facts are retained meaning that whatever was mentioned in the part of uh whatever was mentioned in the part of uh whatever was mentioned in the part of uh whatever was mentioned in part of the context is was mentioned in part of the context is was mentioned in part of the context is properly represented in the working properly represented in the working properly represented in the working state. We also capture if the confidence state. We also capture if the confidence state. We also capture if the confidence collaboration was within a certain error collaboration was within a certain error collaboration was within a certain error bound because uh the agent needs to be bound because uh the agent needs to be bound because uh the agent needs to be able to confidently pull out these able to confidently pull out these able to confidently pull out these signals from the input. We also focus on signals from the input. We also focus on signals from the input. We also focus on getting out um the counterfactual getting out um the counterfactual getting out um the counterfactual sensitivity uh counterfactual sensitivity uh counterfactual sensitivity uh counterfactual sensitivity. So, we do this by flipping sensitivity. So, we do this by flipping sensitivity. So, we do this by flipping some parts of the queries and making some parts of the queries and making some parts of the queries and making sure that when the query changes, the sure that when the query changes, the sure that when the query changes, the underlying constraints pulled out by the underlying constraints pulled out by the underlying constraints pulled out by the agent the those also change and the ones agent the those also change and the ones agent the those also change and the ones that are not relevant stay the same. So, that are not relevant stay the same. So, that are not relevant stay the same. So, we kind of want to measure the we kind of want to measure the we kind of want to measure the sensitivity both ways. sensitivity both ways. sensitivity both ways. And then the second part that happens in And then the second part that happens in And then the second part that happens in the discovery phase as well is once the the discovery phase as well is once the the discovery phase as well is once the agent knows what they have what they agent knows what they have what they agent knows what they have what they already know about the user what more do already know about the user what more do already know about the user what more do they need to know about the user. So they need to know about the user. So they need to know about the user. So what we call this what we call this is what we call this what we call this is what we call this what we call this is kind of discovering the intent gap.

  7. kind of discovering the intent gap. kind of discovering the intent gap. There's a lot of times that there are a There's a lot of times that there are a There's a lot of times that there are a lot of unknown variables before the lot of unknown variables before the lot of unknown variables before the agent can provide the best answer and agent can provide the best answer and agent can provide the best answer and amongst all these unknown variables the amongst all these unknown variables the amongst all these unknown variables the agent doesn't need to go and find all agent doesn't need to go and find all agent doesn't need to go and find all the unknown variables up front. So what the unknown variables up front. So what the unknown variables up front. So what I mean by that is in this unknown in in I mean by that is in this unknown in in I mean by that is in this unknown in in our working example some of the unknown our working example some of the unknown our working example some of the unknown variables that the agent said they need variables that the agent said they need variables that the agent said they need to find out more about is maybe knowing to find out more about is maybe knowing to find out more about is maybe knowing what the room width would be for the what the room width would be for the what the room width would be for the user or even kind of working on user or even kind of working on user or even kind of working on improving the confidence for the style improving the confidence for the style improving the confidence for the style before they can recommend back the before they can recommend back the before they can recommend back the results. Now results. Now results. Now once these unknown variables are figured once these unknown variables are figured once these unknown variables are figured out the agent also needs to work on a out the agent also needs to work on a out the agent also needs to work on a collaborative strategy compare all the collaborative strategy compare all the collaborative strategy compare all the different moves possible and then figure different moves possible and then figure different moves possible and then figure out what is that one unknown variable out what is that one unknown variable out what is that one unknown variable that it should prioritize such that it that it should prioritize such that it that it should prioritize such that it has the maximal information gain at that has the maximal information gain at that has the maximal information gain at that point. So in this case point. So in this case point. So in this case I mean one could argue that maybe I mean one could argue that maybe I mean one could argue that maybe finding out the room width is the best finding out the room width is the best finding out the room width is the best next move for the agent because if the next move for the agent because if the next move for the agent because if the products that the the agent is products that the the agent is products that the the agent is recommending doesn't fit into the room recommending doesn't fit into the room recommending doesn't fit into the room then again it's a moot point and that is then again it's a moot point and that is then again it's a moot point and that is one variable that is going to one variable that is going to one variable that is going to meaningfully change the meaningfully change the meaningfully change the the direction of the conversation and the direction of the conversation and the direction of the conversation and that's what the agent works on in this that's what the agent works on in this that's what the agent works on in this step kind of figuring out what's the step kind of figuring out what's the step kind of figuring out what's the best next thing to ask to the user and best next thing to ask to the user and best next thing to ask to the user and what's why is that the best next thing what's why is that the best next thing what's why is that the best next thing to ask as well. Once it has okay so I'll to ask as well. Once it has okay so I'll to ask as well. Once it has okay so I'll go ahead into the auto data section on go ahead into the auto data section on go ahead into the auto data section on like why how do we evaluate this like why how do we evaluate this like why how do we evaluate this collaborative strategy we first focus on collaborative strategy we first focus on collaborative strategy we first focus on making sure that the agent is able to making sure that the agent is able to making sure that the agent is able to identify all the different blockers that identify all the different blockers that identify all the different blockers that are needed to be answered before the are needed to be answered before the are needed to be answered before the agent can come back with meaningful agent can come back with meaningful agent can come back with meaningful responses. We also work towards making responses. We also work towards making responses. We also work towards making sure that the agent is um

  8. sure that the agent is um sure that the agent is um optimal in trying to get some of these optimal in trying to get some of these optimal in trying to get some of these responses so we also don't want to have responses so we also don't want to have responses so we also don't want to have the agent constantly going into the loop the agent constantly going into the loop the agent constantly going into the loop and continually continuously asking and continually continuously asking and continually continuously asking these questions so over asking is these questions so over asking is these questions so over asking is definitely something we flag. We also definitely something we flag. We also definitely something we flag. We also work on question utility. So, again, how work on question utility. So, again, how work on question utility. So, again, how optimally is the question being asked? optimally is the question being asked? optimally is the question being asked? Is the question indeed useful to get the Is the question indeed useful to get the Is the question indeed useful to get the right or elicit the right preference right or elicit the right preference right or elicit the right preference from the user? And many more. Okay. Okay. Moving on. The The second part that I Moving on. The The second part that I Moving on. The The second part that I was mentioning is the multimodal was mentioning is the multimodal was mentioning is the multimodal elicitation. This is where the research elicitation. This is where the research elicitation. This is where the research phase happens. phase happens. phase happens. So, play along, but like let's say the So, play along, but like let's say the So, play along, but like let's say the the agent has asked about the room the agent has asked about the room the agent has asked about the room width, the user has provided the width, the user has provided the width, the user has provided the response, and then the agent goes to the response, and then the agent goes to the response, and then the agent goes to the next step, which is where the agent is next step, which is where the agent is next step, which is where the agent is asking about the next constraint that it asking about the next constraint that it asking about the next constraint that it needs to know about is what was the needs to know about is what was the needs to know about is what was the style preference of the user. style preference of the user. style preference of the user. Now, the first first thing that the Now, the first first thing that the Now, the first first thing that the agent needs to do is kind of form this agent needs to do is kind of form this agent needs to do is kind of form this temporary bridge between the constraint temporary bridge between the constraint temporary bridge between the constraint itself and how that maps back to the itself and how that maps back to the itself and how that maps back to the the product catalog and the ontology in the product catalog and the ontology in the product catalog and the ontology in your knowledge knowledge database. And your knowledge knowledge database. And your knowledge knowledge database. And this is going to be important because this is going to be important because this is going to be important because when you start retrieving these when you start retrieving these when you start retrieving these products, you want to have a way to map products, you want to have a way to map products, you want to have a way to map these constraints back to your knowledge these constraints back to your knowledge these constraints back to your knowledge base. So, this almost happens in real base. So, this almost happens in real base. So, this almost happens in real time where we map the known constraint time where we map the known constraint time where we map the known constraint or the constraint that the agent is or the constraint that the agent is or the constraint that the agent is exploring back to the knowledge exploring back to the knowledge exploring back to the knowledge database. We also work towards um database. We also work towards um database. We also work towards um figuring out We also work on like the figuring out We also work on like the figuring out We also work on like the agent knowing what is the best way to agent knowing what is the best way to agent knowing what is the best way to get the response for this constraint as

  9. get the response for this constraint as get the response for this constraint as well. So, in this case, the agent has well. So, in this case, the agent has well. So, in this case, the agent has decided that maybe since this is this decided that maybe since this is this decided that maybe since this is this kind of a subjective constraint, the kind of a subjective constraint, the kind of a subjective constraint, the textual the actual elicitation is not textual the actual elicitation is not textual the actual elicitation is not the best way to do this. So, one of the the best way to do this. So, one of the the best way to do this. So, one of the ways that the agent thinks this could ways that the agent thinks this could ways that the agent thinks this could this could be elicited back from the this could be elicited back from the this could be elicited back from the user is using some kind of a visual user is using some kind of a visual user is using some kind of a visual preference preference preference board. board. board. So, the agent then goes back to So, the agent then goes back to So, the agent then goes back to determining what is the best form of determining what is the best form of determining what is the best form of options to show to the user. This could options to show to the user. This could options to show to the user. This could be a combination of figuring out from be a combination of figuring out from be a combination of figuring out from the existing constraints, the past the existing constraints, the past the existing constraints, the past conversation, and then the temporary conversation, and then the temporary conversation, and then the temporary mapping that you've created from the mapping that you've created from the mapping that you've created from the constraints back to your product constraints back to your product constraints back to your product ontology. So, in this case, the the ontology. So, in this case, the the ontology. So, in this case, the the agent thinks that maybe coming up with a agent thinks that maybe coming up with a agent thinks that maybe coming up with a few styles that are most similar to what few styles that are most similar to what few styles that are most similar to what the user had provided as reference image the user had provided as reference image the user had provided as reference image could be a good way to start thinking or could be a good way to start thinking or could be a good way to start thinking or guiding the user towards a common guiding the user towards a common guiding the user towards a common language on what could be uh something language on what could be uh something language on what could be uh something that the user is interested in. that the user is interested in. that the user is interested in. And then the the then the agent also And then the the then the agent also And then the the then the agent also goes into kind of um observing the space goes into kind of um observing the space goes into kind of um observing the space of what kind of reactions the user is of what kind of reactions the user is of what kind of reactions the user is giving. So, it could start looking at giving. So, it could start looking at giving. So, it could start looking at these micro signals of if there was a these micro signals of if there was a these micro signals of if there was a hover or a click in a certain direction hover or a click in a certain direction hover or a click in a certain direction and starts updating its its confidence and starts updating its its confidence and starts updating its its confidence model on what kind of signals uh what model on what kind of signals uh what model on what kind of signals uh what kind of signals can be used to im- kind of signals can be used to im- kind of signals can be used to im- improve the confidence in like what improve the confidence in like what improve the confidence in like what could be the style preference for the could be the style preference for the could be the style preference for the user.

  10. user. user. Um some of the auto raters we use here, Um some of the auto raters we use here, Um some of the auto raters we use here, we do look at how efficient the how we do look at how efficient the how we do look at how efficient the how efficient the agent is in discovering efficient the agent is in discovering efficient the agent is in discovering hidden preferences. So, typically, we hidden preferences. So, typically, we hidden preferences. So, typically, we would use a user simulator, feed it with would use a user simulator, feed it with would use a user simulator, feed it with some constraints, and then we'll see how some constraints, and then we'll see how some constraints, and then we'll see how how efficient the agent was in kind of how efficient the agent was in kind of how efficient the agent was in kind of eliciting some of these constraints. We eliciting some of these constraints. We eliciting some of these constraints. We also focus on turn efficiency. So, also focus on turn efficiency. So, also focus on turn efficiency. So, how efficient was the user in how many how efficient was the user in how many how efficient was the user in how many turns did it take for the user to be turns did it take for the user to be turns did it take for the user to be able to to be able to uncover all these able to to be able to uncover all these able to to be able to uncover all these hidden preferences. Ideally, we don't hidden preferences. Ideally, we don't hidden preferences. Ideally, we don't want the user to go into this loop and want the user to go into this loop and want the user to go into this loop and keep asking the same questions again and keep asking the same questions again and keep asking the same questions again and again and again. Or also, we don't want again and again. Or also, we don't want again and again. Or also, we don't want to go into the loop of asking some to go into the loop of asking some to go into the loop of asking some um some questions which may not give you um some questions which may not give you um some questions which may not give you the best uh which may not give you the the best uh which may not give you the the best uh which may not give you the best information required to proceed the best information required to proceed the best information required to proceed the conversation. And then we also look at conversation. And then we also look at conversation. And then we also look at forma- format selection accuracy. So, we forma- format selection accuracy. So, we forma- format selection accuracy. So, we typically also look for if the agent is typically also look for if the agent is typically also look for if the agent is asking the right right question in the asking the right right question in the asking the right right question in the right format. So, for example, if the if right format. So, for example, if the if right format. So, for example, if the if the question was something that was the question was something that was the question was something that was easily statable, the right format could easily statable, the right format could easily statable, the right format could be a textual elicitation. But if if this be a textual elicitation. But if if this be a textual elicitation. But if if this was more of a was more of a was more of a fuzzy question where the user is clearly fuzzy question where the user is clearly fuzzy question where the user is clearly having an articulation gap and is not having an articulation gap and is not having an articulation gap and is not able to describe the preference, maybe able to describe the preference, maybe able to describe the preference, maybe the best way to do this is speak a the best way to do this is speak a the best way to do this is speak a common language and come up with some common language and come up with some common language and come up with some visual anchor points for the user to say visual anchor points for the user to say visual anchor points for the user to say what speaks more to them.

  11. Um Um Yeah, so then the last step in this Yeah, so then the last step in this Yeah, so then the last step in this process, once the So, to recap like process, once the So, to recap like process, once the So, to recap like basically the Now the agent knows basically the Now the agent knows basically the Now the agent knows exactly what they what the user is exactly what they what the user is exactly what they what the user is looking for. The agent has identified a looking for. The agent has identified a looking for. The agent has identified a collaboration strategy, has figured out collaboration strategy, has figured out collaboration strategy, has figured out all the preferences for the user, and all the preferences for the user, and all the preferences for the user, and knows what is going to be most optimal knows what is going to be most optimal knows what is going to be most optimal in terms of um the different priorities in terms of um the different priorities in terms of um the different priorities for the user. The last step in this for the user. The last step in this for the user. The last step in this process is to also use the model process is to also use the model process is to also use the model intelligence to figure out what is the intelligence to figure out what is the intelligence to figure out what is the best way to provide this response back best way to provide this response back best way to provide this response back to the user. So, in this working to the user. So, in this working to the user. So, in this working example, we found out that the agent example, we found out that the agent example, we found out that the agent knows what the style preference is, has knows what the style preference is, has knows what the style preference is, has figured out that the user is looking to figured out that the user is looking to figured out that the user is looking to buy products under a certain budget, and buy products under a certain budget, and buy products under a certain budget, and knows what are the different dimensions knows what are the different dimensions knows what are the different dimensions along which it needs to find uh these along which it needs to find uh these along which it needs to find uh these products for because through the products for because through the products for because through the conversation they figured out some of conversation they figured out some of conversation they figured out some of the different um the different um the different um different metadata information that is different metadata information that is different metadata information that is going to be relevant to surface when going to be relevant to surface when going to be relevant to surface when they're giving out this product they're giving out this product they're giving out this product information back to the user. information back to the user. information back to the user. So, the one one important step that So, the one one important step that So, the one one important step that happens at this this stage is figuring happens at this this stage is figuring happens at this this stage is figuring out what's the best way to surface back out what's the best way to surface back out what's the best way to surface back this response. So, for example, like if this response. So, for example, like if this response. So, for example, like if the user was looking for a particular the user was looking for a particular the user was looking for a particular policy or review information about a policy or review information about a policy or review information about a particular product, maybe the best way particular product, maybe the best way particular product, maybe the best way to do this is to go give out a summary to do this is to go give out a summary to do this is to go give out a summary or a bulleted list. But if the if the or a bulleted list. But if the if the or a bulleted list. But if the if the user was looking more towards comparing user was looking more towards comparing user was looking more towards comparing two different products, maybe the best two different products, maybe the best two different products, maybe the best response is to give out a trade-off response is to give out a trade-off response is to give out a trade-off table or a comparison table across the table or a comparison table across the table or a comparison table across the different axis that the user cares different axis that the user cares different axis that the user cares about.

  12. about. about. Um and in case that in our example where Um and in case that in our example where Um and in case that in our example where the user was looking more towards kind the user was looking more towards kind the user was looking more towards kind of style inspiration or ideas on how of style inspiration or ideas on how of style inspiration or ideas on how they can kind of redo their room, then they can kind of redo their room, then they can kind of redo their room, then maybe the best way is to give out some maybe the best way is to give out some maybe the best way is to give out some visual references and product visual references and product visual references and product inspiration photos on like how What are inspiration photos on like how What are inspiration photos on like how What are the different options and possibilities the different options and possibilities the different options and possibilities uh for the user to uncover. So, uh for the user to uncover. So, uh for the user to uncover. So, some other ways that we focus on some other ways that we focus on some other ways that we focus on evaluating uh this stage is focusing on evaluating uh this stage is focusing on evaluating uh this stage is focusing on the format accuracy. So, we want to make the format accuracy. So, we want to make the format accuracy. So, we want to make sure that the response format is kind of sure that the response format is kind of sure that the response format is kind of optimal for the user query so that they optimal for the user query so that they optimal for the user query so that they they find out exactly what they're they find out exactly what they're they find out exactly what they're looking for. looking for. looking for. Uh ideally the information that they're Uh ideally the information that they're Uh ideally the information that they're looking for should not be buried in the looking for should not be buried in the looking for should not be buried in the response, but should be easy for the response, but should be easy for the response, but should be easy for the user to spot so they can come into the user to spot so they can come into the user to spot so they can come into the next stage in the intent journey, which next stage in the intent journey, which next stage in the intent journey, which is to basically buy the product. We also is to basically buy the product. We also is to basically buy the product. We also look at data fidelity, which is to make look at data fidelity, which is to make look at data fidelity, which is to make sure that the model is not hallucinating sure that the model is not hallucinating sure that the model is not hallucinating and it's really capturing the and it's really capturing the and it's really capturing the information in the correct format in the information in the correct format in the information in the correct format in the you know, the the information is just you know, the the information is just you know, the the information is just accurate and is captured across accurate and is captured across accurate and is captured across the response. We also look at user the response. We also look at user the response. We also look at user actionability. So, this is ensuring that actionability. So, this is ensuring that actionability. So, this is ensuring that the response format is such that the the response format is such that the the response format is such that the user is very confident and and commits user is very confident and and commits user is very confident and and commits to the next action, which is like I to the next action, which is like I to the next action, which is like I said, the action to purchase the said, the action to purchase the said, the action to purchase the product.

  13. product. product. So, So, So, just to recap so far, what we have is um just to recap so far, what we have is um just to recap so far, what we have is um you want to design the product. Um you want to design the product. Um you want to design the product. Um you want to you want to design the you want to you want to design the you want to you want to design the product such that you're prepared to product such that you're prepared to product such that you're prepared to accept vibes, like I say. So, that is to accept vibes, like I say. So, that is to accept vibes, like I say. So, that is to say that users will come with fuzzy say that users will come with fuzzy say that users will come with fuzzy intent. Users will not have a intent. Users will not have a intent. Users will not have a well-defined goal. So, you want to make well-defined goal. So, you want to make well-defined goal. So, you want to make sure that your system is able to work sure that your system is able to work sure that your system is able to work through queries that are not clean. through queries that are not clean. through queries that are not clean. Um the second takeaway is you want to Um the second takeaway is you want to Um the second takeaway is you want to focus on showing and asking rather than focus on showing and asking rather than focus on showing and asking rather than always asking with textual textual with always asking with textual textual with always asking with textual textual with textual elicitations. Again, like textual elicitations. Again, like textual elicitations. Again, like visuals and comparisons do reveal visuals and comparisons do reveal visuals and comparisons do reveal preferences much much faster. It allows preferences much much faster. It allows preferences much much faster. It allows you the agent and the user to speak a you the agent and the user to speak a you the agent and the user to speak a common language. common language. common language. Third one I would say is shape the Third one I would say is shape the Third one I would say is shape the answer. So, do focus on making sure that answer. So, do focus on making sure that answer. So, do focus on making sure that the presentation format is ideal for the the presentation format is ideal for the the presentation format is ideal for the user being able to find the right user being able to find the right user being able to find the right information. information. information. Um the way you have the model response Um the way you have the model response Um the way you have the model response structure is also very much part of the structure is also very much part of the structure is also very much part of the intelligence. And then the last one is intelligence. And then the last one is intelligence. And then the last one is make sure that you grade the loop. You make sure that you grade the loop. You make sure that you grade the loop. You have the right auto rater set up on have the right auto rater set up on have the right auto rater set up on every step of the process. Um and every step of the process. Um and every step of the process. Um and honestly, developing these auto raters honestly, developing these auto raters honestly, developing these auto raters is a is a is a is is almost like an evolving system. It is is almost like an evolving system. It is is almost like an evolving system. It starts very simple, but as and when you starts very simple, but as and when you starts very simple, but as and when you start the the system starts evolving, start the the system starts evolving, start the the system starts evolving, you want the auto raters to kind of you want the auto raters to kind of you want the auto raters to kind of gradually grow with your system and gradually grow with your system and gradually grow with your system and start start start um um um yeah, it which just gradually grow with yeah, it which just gradually grow with yeah, it which just gradually grow with your system.

  14. your system. your system. So, that's all I had. Um So, that's all I had. Um So, that's all I had. Um I can take a few questions, but I can take a few questions, but I can take a few questions, but hopefully the learnings we shared are hopefully the learnings we shared are hopefully the learnings we shared are useful for whatever you folks are useful for whatever you folks are useful for whatever you folks are building. building. building. Yeah. Yeah. Yeah. Yeah, great question. So, the question Yeah, great question. So, the question Yeah, great question. So, the question is about what should the how should the is about what should the how should the is about what should the how should the ontology be structured on the merchant ontology be structured on the merchant ontology be structured on the merchant side so that it's fair both for the side so that it's fair both for the side so that it's fair both for the agent and the merchants. So, yes, we do agent and the merchants. So, yes, we do agent and the merchants. So, yes, we do take a lot of advantage on the domain take a lot of advantage on the domain take a lot of advantage on the domain expertise of the merchant on as to what expertise of the merchant on as to what expertise of the merchant on as to what they're trying to sell and like we do they're trying to sell and like we do they're trying to sell and like we do work towards creating So, you remember work towards creating So, you remember work towards creating So, you remember how I was mentioning about the bridge how I was mentioning about the bridge how I was mentioning about the bridge between the constraints that the user between the constraints that the user between the constraints that the user might have specified and then what the might have specified and then what the might have specified and then what the agent agent agent kind of understands. That is where we do kind of understands. That is where we do kind of understands. That is where we do expect a lot of intelligence to flow expect a lot of intelligence to flow expect a lot of intelligence to flow from the merchant side where the from the merchant side where the from the merchant side where the ontology on how that constraint could ontology on how that constraint could ontology on how that constraint could map to the different metadata that the map to the different metadata that the map to the different metadata that the agent has. Sorry, the that the merchant agent has. Sorry, the that the merchant agent has. Sorry, the that the merchant has maps in. So, yes, we do partner a

  15. has maps in. So, yes, we do partner a has maps in. So, yes, we do partner a lot and then there's also like the UCP lot and then there's also like the UCP lot and then there's also like the UCP stuff that we launched recently which stuff that we launched recently which stuff that we launched recently which allows all the merchants to kind of allows all the merchants to kind of allows all the merchants to kind of start speaking the common language with start speaking the common language with start speaking the common language with the agent as well. >> Yeah, so ideally, I mean, honestly right >> Yeah, so ideally, I mean, honestly right now we focus on making sure all of this now we focus on making sure all of this now we focus on making sure all of this flows back to the agent and the agent flows back to the agent and the agent flows back to the agent and the agent makes the decisions because you want to makes the decisions because you want to makes the decisions because you want to build like a horizontal common layer build like a horizontal common layer build like a horizontal common layer across all the different merchants. So, across all the different merchants. So, across all the different merchants. So, and also like it should be a seamless and also like it should be a seamless and also like it should be a seamless experience for the for the user who's experience for the for the user who's experience for the for the user who's interacting with our apps. So, right now interacting with our apps. So, right now interacting with our apps. So, right now the response format is very much part of the response format is very much part of the response format is very much part of the agent's intelligence. It is not the agent's intelligence. It is not the agent's intelligence. It is not something that the merchant gets to something that the merchant gets to something that the merchant gets to decide. decide. decide. Yeah. Yeah. Yeah. Yes. Yes. Yes. >> I'm just curious your opinion what >> I'm just curious your opinion what >> I'm just curious your opinion what happens when your user is not a user happens when your user is not a user happens when your user is not a user anymore, it's an agent.

  16. anymore, it's an agent. anymore, it's an agent. >> That's a great question. >> That's a great question. >> That's a great question. We're in the early stages of building We're in the early stages of building We're in the early stages of building this out, but I think I mean there could this out, but I think I mean there could this out, but I think I mean there could be a case where um be a case where um be a case where um Yeah. Yeah. Yeah. Yeah, yeah, exactly. So, yeah, I would I Yeah, yeah, exactly. So, yeah, I would I Yeah, yeah, exactly. So, yeah, I would I would expect like an MCP to be the would expect like an MCP to be the would expect like an MCP to be the interface between the two for sure. interface between the two for sure. interface between the two for sure. Honestly, we haven't gotten to a point Honestly, we haven't gotten to a point Honestly, we haven't gotten to a point where we have agents interacting with where we have agents interacting with where we have agents interacting with our agents just yet. Also, like what our agents just yet. Also, like what our agents just yet. Also, like what we've realized at least from our user we've realized at least from our user we've realized at least from our user studies is studies is studies is users really like to be more involved in users really like to be more involved in users really like to be more involved in the process of choosing or even the process of choosing or even the process of choosing or even exploring the different possibilities. exploring the different possibilities. exploring the different possibilities. So, during the upper funnel journeys So, during the upper funnel journeys So, during the upper funnel journeys where users is looking more towards where users is looking more towards where users is looking more towards discovery, inspiration, that is where discovery, inspiration, that is where discovery, inspiration, that is where they would rather be interacting with they would rather be interacting with they would rather be interacting with the system than with their agent. I the system than with their agent. I the system than with their agent. I think where the agent typically comes in think where the agent typically comes in think where the agent typically comes in or even where or even where or even where what we've heard is like the was the what we've heard is like the was the what we've heard is like the was the lower end of the journey where they're lower end of the journey where they're lower end of the journey where they're just looking to compare or negotiate or just looking to compare or negotiate or just looking to compare or negotiate or compare prices across different compare prices across different compare prices across different merchants, but very much up there in the merchants, but very much up there in the merchants, but very much up there in the funnel, it's the users who kind of funnel, it's the users who kind of funnel, it's the users who kind of interact more with our systems. So, interact more with our systems. So, interact more with our systems. So, yeah. Yeah, I think Yeah, I think yeah, I can take questions outside, but yeah, I can take questions outside, but yeah, I can take questions outside, but thank you folks for coming.

Summary

The talk focuses on multimodal collaborative agents designed to assist users with fuzzy or unclear intent, particularly in shopping scenarios. It highlights the current limitation of agents acting as mere search wrappers and the need for them to proactively elicit and hand-hold users through preference discovery. The practical takeaway is the development of a framework or loop that guides users from ambiguous initial desires to achieving their goals through effective agent interaction.

View original episode ↗