Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku
Read full transcript 16 segments
-
Hi everyone. How's it going? Hey Hi everyone. How's it going? Hey Patricia, how are you? Patricia, how are you? Patricia, how are you? >> Uh so last presentation of the day, so >> Uh so last presentation of the day, so >> Uh so last presentation of the day, so let's make it count. Um, let's make it count. Um, let's make it count. Um, all right. Let's, uh, let me start with all right. Let's, uh, let me start with all right. Let's, uh, let me start with a little bit of background on myself. a little bit of background on myself. a little bit of background on myself. And, um, my background, I'm a voice And, um, my background, I'm a voice And, um, my background, I'm a voice subject matter expert. I've been working subject matter expert. I've been working subject matter expert. I've been working in voice AI for a long time across in voice AI for a long time across in voice AI for a long time across different surfaces, devices, and um, different surfaces, devices, and um, different surfaces, devices, and um, both at Alexa, at at Roku, at my own both at Alexa, at at Roku, at my own both at Alexa, at at Roku, at my own startups, you know, in the app store. startups, you know, in the app store. startups, you know, in the app store. And my perspective is a little different And my perspective is a little different And my perspective is a little different from a lot of other voice AI from a lot of other voice AI from a lot of other voice AI practitioners. I think it's a practitioners. I think it's a practitioners. I think it's a combination of um a deep um voice user combination of um a deep um voice user combination of um a deep um voice user interface expertise and intu intuition interface expertise and intu intuition interface expertise and intu intuition mixed in with new technical approaches mixed in with new technical approaches mixed in with new technical approaches uh that I think can produce really uh that I think can produce really uh that I think can produce really magical experiences. So I think it's magical experiences. So I think it's magical experiences. So I think it's both sides and I think that's especially both sides and I think that's especially both sides and I think that's especially true in this new area that we're in with true in this new area that we're in with true in this new area that we're in with frontier tech where the human interface frontier tech where the human interface frontier tech where the human interface is basically being redefined. So let me is basically being redefined. So let me is basically being redefined. So let me start with uh I'll just blast through start with uh I'll just blast through start with uh I'll just blast through the first couple of slides then get to the first couple of slides then get to the first couple of slides then get to the premise. I think everybody knows the premise. I think everybody knows the premise. I think everybody knows that voice has incredible potential.
-
that voice has incredible potential. that voice has incredible potential. There's the power of voice I think There's the power of voice I think There's the power of voice I think across everywhere. It's the most natural across everywhere. It's the most natural across everywhere. It's the most natural interface. Humans love talking. And uh interface. Humans love talking. And uh interface. Humans love talking. And uh the problem is the other half is the the problem is the other half is the the problem is the other half is the pain of voice. So it's the power and the pain of voice. So it's the power and the pain of voice. So it's the power and the pain. Voice is errorprone. And I think pain. Voice is errorprone. And I think pain. Voice is errorprone. And I think those errors are going to continue for a those errors are going to continue for a those errors are going to continue for a while. And I think the cost or while. And I think the cost or while. And I think the cost or consequence of those errors is going to consequence of those errors is going to consequence of those errors is going to grow, especially as we go fromational grow, especially as we go fromational grow, especially as we go fromational AI bots to embodied AI where rather than AI bots to embodied AI where rather than AI bots to embodied AI where rather than just giving answers that might be just giving answers that might be just giving answers that might be erroneous, erroneous, erroneous, we're going to have AI systems take we're going to have AI systems take we're going to have AI systems take physical actions or digital actions physical actions or digital actions physical actions or digital actions where, you know, if the robot throws where, you know, if the robot throws where, you know, if the robot throws your watch out with the trash, it's a your watch out with the trash, it's a your watch out with the trash, it's a lot worse than playing the wrong song. lot worse than playing the wrong song. lot worse than playing the wrong song. So I do think that a new approach is So I do think that a new approach is So I do think that a new approach is definitely needed and here's the TLDDR definitely needed and here's the TLDDR definitely needed and here's the TLDDR of the premise we're going to walk of the premise we're going to walk of the premise we're going to walk through today. Um there are two ways to through today. Um there are two ways to through today. Um there are two ways to improve customer or user satisfaction of improve customer or user satisfaction of improve customer or user satisfaction of a voice AI assistant and that is by a voice AI assistant and that is by a voice AI assistant and that is by increasing accuracy which people know increasing accuracy which people know increasing accuracy which people know about I mean technically accuracy and about I mean technically accuracy and about I mean technically accuracy and the other is a different knob that we the other is a different knob that we the other is a different knob that we have that we are not using adequately have that we are not using adequately have that we are not using adequately and I'll call that a system decision and I'll call that a system decision and I'll call that a system decision which we will define which is orthogonal which we will define which is orthogonal which we will define which is orthogonal which is different from accuracy and I which is different from accuracy and I which is different from accuracy and I believe This approach which I have used believe This approach which I have used believe This approach which I have used in several different environments and
-
in several different environments and in several different environments and seen some success I think is a promising seen some success I think is a promising seen some success I think is a promising area that we should consider developing. area that we should consider developing. area that we should consider developing. Um let me walk through this with a Um let me walk through this with a Um let me walk through this with a simple smart speaker example and we'll simple smart speaker example and we'll simple smart speaker example and we'll go step by step with this approach but go step by step with this approach but go step by step with this approach but it is a scalable approach that I think it is a scalable approach that I think it is a scalable approach that I think uh can apply across different surfaces uh can apply across different surfaces uh can apply across different surfaces and devices. So let's get started. So and devices. So let's get started. So and devices. So let's get started. So suppose we all you know are making a suppose we all you know are making a suppose we all you know are making a smart speaker coincidentally called uh smart speaker coincidentally called uh smart speaker coincidentally called uh Alexa and Alexa is very simple. It just Alexa and Alexa is very simple. It just Alexa and Alexa is very simple. It just allows you to you know ask for music and allows you to you know ask for music and allows you to you know ask for music and it'll play a song and of course it will it'll play a song and of course it will it'll play a song and of course it will play either the song you wanted or a play either the song you wanted or a play either the song you wanted or a different song. So it'll be right or different song. So it'll be right or different song. So it'll be right or it'll be wrong. This isn't that it'll be wrong. This isn't that it'll be wrong. This isn't that different from what you've seen out different from what you've seen out different from what you've seen out there. Um now let's to first talk about there. Um now let's to first talk about there. Um now let's to first talk about accuracy. Accuracy. Let's say we define accuracy. Accuracy. Let's say we define accuracy. Accuracy. Let's say we define it as we you know take a thousand spoken it as we you know take a thousand spoken it as we you know take a thousand spoken requests. We observe the input and the requests. We observe the input and the requests. We observe the input and the output. We label it and we look at this.
-
output. We label it and we look at this. output. We label it and we look at this. This is the map of a thousand points and This is the map of a thousand points and This is the map of a thousand points and 79% of the time 790 dots here were 79% of the time 790 dots here were 79% of the time 790 dots here were actually the correct song. This is let's actually the correct song. This is let's actually the correct song. This is let's say human annotated 20% 21% wrong song. say human annotated 20% 21% wrong song. say human annotated 20% 21% wrong song. So that's the accuracy. Now, like I So that's the accuracy. Now, like I So that's the accuracy. Now, like I said, knob one is to spend a lot of time said, knob one is to spend a lot of time said, knob one is to spend a lot of time working on improving the accuracy, you working on improving the accuracy, you working on improving the accuracy, you know, um, percentage point by percentage know, um, percentage point by percentage know, um, percentage point by percentage point at any layer in the stack. point at any layer in the stack. point at any layer in the stack. There's, if it's a cascaded system, you There's, if it's a cascaded system, you There's, if it's a cascaded system, you know, there's a perhaps a wakeword layer know, there's a perhaps a wakeword layer know, there's a perhaps a wakeword layer and a speech ASR layer and a NLU layer and a speech ASR layer and a NLU layer and a speech ASR layer and a NLU layer which might have intent classification, which might have intent classification, which might have intent classification, entity extraction, a lot of different entity extraction, a lot of different entity extraction, a lot of different layers, VAD, etc. And any of those can layers, VAD, etc. And any of those can layers, VAD, etc. And any of those can contribute to errors. So we spent time contribute to errors. So we spent time contribute to errors. So we spent time we might be able to reduce that 210 to a we might be able to reduce that 210 to a we might be able to reduce that 210 to a smaller number that is I think a known smaller number that is I think a known smaller number that is I think a known area that we're tackling but I think area that we're tackling but I think area that we're tackling but I think knob 2 which is what I was talking about knob 2 which is what I was talking about knob 2 which is what I was talking about is what we'll go through here which is is what we'll go through here which is is what we'll go through here which is keeping the accuracy exactly the same.
-
keeping the accuracy exactly the same. keeping the accuracy exactly the same. So 79% what could we do in conditions of So 79% what could we do in conditions of So 79% what could we do in conditions of uncertainty to improve user satisfaction uncertainty to improve user satisfaction uncertainty to improve user satisfaction apparent and I I think we can do a lot. apparent and I I think we can do a lot. apparent and I I think we can do a lot. So let's start first with the original So let's start first with the original So let's start first with the original system is just acting like I said user system is just acting like I said user system is just acting like I said user says something system plays a song it's says something system plays a song it's says something system plays a song it's either the right song or the wrong song either the right song or the wrong song either the right song or the wrong song immediately I think just common sense immediately I think just common sense immediately I think just common sense tells us that we could introduce at tells us that we could introduce at tells us that we could introduce at least one system behavior to stop or least one system behavior to stop or least one system behavior to stop or rather to reject the hypothesis and do rather to reject the hypothesis and do rather to reject the hypothesis and do nothing. So uh there is now one more nothing. So uh there is now one more nothing. So uh there is now one more option to decide the system may decide option to decide the system may decide option to decide the system may decide and say sorry I didn't get that or sorry and say sorry I didn't get that or sorry and say sorry I didn't get that or sorry could you repeat that? Uh the challenge could you repeat that? Uh the challenge could you repeat that? Uh the challenge of course is how how when do we decide of course is how how when do we decide of course is how how when do we decide to stop and I mean quantitatively. Um to stop and I mean quantitatively. Um to stop and I mean quantitatively. Um here's one approach to kind of here's one approach to kind of here's one approach to kind of visualizing this because if we don't visualizing this because if we don't visualizing this because if we don't we'll just take probably some swag like we'll just take probably some swag like we'll just take probably some swag like some guesstimate and I'll prove that if some guesstimate and I'll prove that if some guesstimate and I'll prove that if we just took a guesstimate we would end we just took a guesstimate we would end we just took a guesstimate we would end up with a worse situation than a more up with a worse situation than a more up with a worse situation than a more rigorous approach. So let's just assume rigorous approach. So let's just assume rigorous approach. So let's just assume I took those thousand data points and I took those thousand data points and I took those thousand data points and like I said they've been annotated and like I said they've been annotated and like I said they've been annotated and we assign a confidence score a single we assign a confidence score a single we assign a confidence score a single confidence score to the hypothesis that confidence score to the hypothesis that confidence score to the hypothesis that was generated by the system you know was generated by the system you know was generated by the system you know between zero and one and let's say it's between zero and one and let's say it's between zero and one and let's say it's reasonably calibrated. This is a reasonably calibrated. This is a reasonably calibrated. This is a simplification of if it's a cascaded simplification of if it's a cascaded simplification of if it's a cascaded system there are multiple layers and system there are multiple layers and system there are multiple layers and multiple you know confidence scores but multiple you know confidence scores but multiple you know confidence scores but let's just assume that for now. Whoops.
-
let's just assume that for now. Whoops. let's just assume that for now. Whoops. So we're going to have 790 points 200 So we're going to have 790 points 200 So we're going to have 790 points 200 that are correct 210 wrong. Each one has that are correct 210 wrong. Each one has that are correct 210 wrong. Each one has a confidence score and we're going to a confidence score and we're going to a confidence score and we're going to plot it, you know, plot the plot it, you know, plot the plot it, you know, plot the distributions. Uh on the x-axis, I've distributions. Uh on the x-axis, I've distributions. Uh on the x-axis, I've just converted from 0ero to one to just converted from 0ero to one to just converted from 0ero to one to percentages. And the question is how do percentages. And the question is how do percentages. And the question is how do we choose a threshold t such that we choose a threshold t such that we choose a threshold t such that whatever that percentage is um to the whatever that percentage is um to the whatever that percentage is um to the left of it meaning if when the system um left of it meaning if when the system um left of it meaning if when the system um forms a hypothesis if the confidence forms a hypothesis if the confidence forms a hypothesis if the confidence score c is less than that t stop and say score c is less than that t stop and say score c is less than that t stop and say sorry otherwise play question is how do sorry otherwise play question is how do sorry otherwise play question is how do we choose a t so far everything I'm we choose a t so far everything I'm we choose a t so far everything I'm saying is fairly common sensical but saying is fairly common sensical but saying is fairly common sensical but this is where um intuition will fail us this is where um intuition will fail us this is where um intuition will fail us we might say something like okay I don't we might say something like okay I don't we might say something like okay I don't know let's do 65%. It seems you know gut know let's do 65%. It seems you know gut know let's do 65%. It seems you know gut feeling like okay it's kind of confident feeling like okay it's kind of confident feeling like okay it's kind of confident that's probably when we should speak. Um that's probably when we should speak. Um that's probably when we should speak. Um now here's where we start coming out now here's where we start coming out now here's where we start coming out with some sophistication. with some sophistication. with some sophistication. Any tea we choose is producing bad Any tea we choose is producing bad Any tea we choose is producing bad outcomes. Bad in the in two fields. One outcomes. Bad in the in two fields. One outcomes. Bad in the in two fields. One is obviously on the left side anytime is obviously on the left side anytime is obviously on the left side anytime you stop it's bad. The user doesn't want you stop it's bad. The user doesn't want you stop it's bad. The user doesn't want it to stop. He wants to they want to it to stop. He wants to they want to it to stop. He wants to they want to hear their song. The other bad is if you hear their song. The other bad is if you hear their song. The other bad is if you do play a wrong song, of course that's do play a wrong song, of course that's do play a wrong song, of course that's bad as well. So these are two two kinds bad as well. So these are two two kinds bad as well. So these are two two kinds of bad outcomes. But here's the of bad outcomes. But here's the of bad outcomes. But here's the important part. Now I've like elaborated important part. Now I've like elaborated important part. Now I've like elaborated on the um tree diagram on the right hand on the um tree diagram on the right hand on the um tree diagram on the right hand side. The bad outcomes are not equally
-
side. The bad outcomes are not equally side. The bad outcomes are not equally bad. They're not the same thing from a bad. They're not the same thing from a bad. They're not the same thing from a user perspective. And obviously let's user perspective. And obviously let's user perspective. And obviously let's let's think about it. If the wrong song let's think about it. If the wrong song let's think about it. If the wrong song plays, you said play kiss and it starts plays, you said play kiss and it starts plays, you said play kiss and it starts playing kiss by Chris Brown instead of playing kiss by Chris Brown instead of playing kiss by Chris Brown instead of the one by Prince. That's going to be um the one by Prince. That's going to be um the one by Prince. That's going to be um the highest user cost. Now I'm defining the highest user cost. Now I'm defining the highest user cost. Now I'm defining user cost from the user's perspective. user cost from the user's perspective. user cost from the user's perspective. First I have to like hear music and First I have to like hear music and First I have to like hear music and realize that is not Prince. Then I have realize that is not Prince. Then I have realize that is not Prince. Then I have to shout over my Alexa and um you know to shout over my Alexa and um you know to shout over my Alexa and um you know get it to stop and then I have to get it to stop and then I have to get it to stop and then I have to re-request. All of that is a lot of re-request. All of that is a lot of re-request. All of that is a lot of effort. that is definitely a worse effort. that is definitely a worse effort. that is definitely a worse outcome than the system stopping and outcome than the system stopping and outcome than the system stopping and saying sorry I didn't understand that saying sorry I didn't understand that saying sorry I didn't understand that however we should go further and try to however we should go further and try to however we should go further and try to quantify that relative badness and there quantify that relative badness and there quantify that relative badness and there many ways to do it and I think this is many ways to do it and I think this is many ways to do it and I think this is an area to be explored for now let's an area to be explored for now let's an area to be explored for now let's just consider this a heristic of if that just consider this a heristic of if that just consider this a heristic of if that outcome happens how many more seconds outcome happens how many more seconds outcome happens how many more seconds additional seconds will it take for the additional seconds will it take for the additional seconds will it take for the user to get back to success which is to user to get back to success which is to user to get back to success which is to play the song they wanted kiss by Prince play the song they wanted kiss by Prince play the song they wanted kiss by Prince and I'm I just put down some numbers and I'm I just put down some numbers and I'm I just put down some numbers here. Let's say in the case of a bad here. Let's say in the case of a bad here. Let's say in the case of a bad song, it's 10 seconds if you add up all song, it's 10 seconds if you add up all song, it's 10 seconds if you add up all the things I got to do. And if it's a I the things I got to do. And if it's a I the things I got to do. And if it's a I didn't understand you, it's 4 seconds didn't understand you, it's 4 seconds didn't understand you, it's 4 seconds because that's how long it would take because that's how long it would take because that's how long it would take you to respe and and the extra latency.
-
you to respe and and the extra latency. you to respe and and the extra latency. And now here's where we can start And now here's where we can start And now here's where we can start utilizing that. If we go back to our utilizing that. If we go back to our utilizing that. If we go back to our distribution curve on trying to find out distribution curve on trying to find out distribution curve on trying to find out where is T. Now we've basically turned where is T. Now we've basically turned where is T. Now we've basically turned this into a problem of minimizing a cost this into a problem of minimizing a cost this into a problem of minimizing a cost function. It's a user cost function. It function. It's a user cost function. It function. It's a user cost function. It is the number of bad acts wherever that is the number of bad acts wherever that is the number of bad acts wherever that whatever the t causes times 10 because whatever the t causes times 10 because whatever the t causes times 10 because that was a unit cost we gave plus the that was a unit cost we gave plus the that was a unit cost we gave plus the number of stops times four because number of stops times four because number of stops times four because that's the the unit cost we gave. By the that's the the unit cost we gave. By the that's the the unit cost we gave. By the way, one thing I should have elaborated way, one thing I should have elaborated way, one thing I should have elaborated because I work in voice and we like because I work in voice and we like because I work in voice and we like language and we like puns. So this whole language and we like puns. So this whole language and we like puns. So this whole thing is called an outcome user cost thing is called an outcome user cost thing is called an outcome user cost heruristic. So that spells the word ouch heruristic. So that spells the word ouch heruristic. So that spells the word ouch and that is some expression of pain. and that is some expression of pain. and that is some expression of pain. Yes, we are you know language nerds. So Yes, we are you know language nerds. So Yes, we are you know language nerds. So these kinds of things amuse us. Um so these kinds of things amuse us. Um so these kinds of things amuse us. Um so now let's consider that is the cost now let's consider that is the cost now let's consider that is the cost function is to minimize the ouch. And function is to minimize the ouch. And function is to minimize the ouch. And now um that let's see if uh I'm going to now um that let's see if uh I'm going to now um that let's see if uh I'm going to bring up a tool. Let's see if this bring up a tool. Let's see if this bring up a tool. Let's see if this works.
-
works. works. Where I have actually gotten or with one Where I have actually gotten or with one Where I have actually gotten or with one of my coding assistants gotten uh an of my coding assistants gotten uh an of my coding assistants gotten uh an interactive interactive interactive um graph where we have actually plotted um graph where we have actually plotted um graph where we have actually plotted those thousand points and as we vary the those thousand points and as we vary the those thousand points and as we vary the threshold t you can see that the total threshold t you can see that the total threshold t you can see that the total user cost here which is that function of user cost here which is that function of user cost here which is that function of you know x * y + a * b actually changes. you know x * y + a * b actually changes. you know x * y + a * b actually changes. So let's in the very beginning when we So let's in the very beginning when we So let's in the very beginning when we said the system was just playing said the system was just playing said the system was just playing the the cost across those thousand the the cost across those thousand the the cost across those thousand points was 2100 or divided by a,000 is points was 2100 or divided by a,000 is points was 2100 or divided by a,000 is 2.1 ouch points per turn. Then we said 2.1 ouch points per turn. Then we said 2.1 ouch points per turn. Then we said okay let's insert a stop behavior and okay let's insert a stop behavior and okay let's insert a stop behavior and let's like wing it and say 65%. That's let's like wing it and say 65%. That's let's like wing it and say 65%. That's when I want the threshold. If we brought when I want the threshold. If we brought when I want the threshold. If we brought this up to 65 yeah that's better. Now this up to 65 yeah that's better. Now this up to 65 yeah that's better. Now it's 1904 or 1.9 per turn, but it's not it's 1904 or 1.9 per turn, but it's not it's 1904 or 1.9 per turn, but it's not optimal. As it turns out, if we do optimal. As it turns out, if we do optimal. As it turns out, if we do actually um ask for the AI to solve the actually um ask for the AI to solve the actually um ask for the AI to solve the uh the problem across this curve, it uh the problem across this curve, it uh the problem across this curve, it turns out 43%. So I'll drag it now to 43 turns out 43%. So I'll drag it now to 43 turns out 43%. So I'll drag it now to 43 is in fact is in fact is in fact the optimal the optimal the optimal optimal point of t. This minimizes the optimal point of t. This minimizes the optimal point of t. This minimizes the cost function. You can see it's the cost function. You can see it's the cost function. You can see it's the lowest point on this graph down here to lowest point on this graph down here to lowest point on this graph down here to 1 1 1 27. So effectively we haven't changed 27. So effectively we haven't changed 27. So effectively we haven't changed the accuracy at all. The system is not the accuracy at all. The system is not the accuracy at all. The system is not any smarter in that sense. But with some
-
any smarter in that sense. But with some any smarter in that sense. But with some clever system behavior, conversational clever system behavior, conversational clever system behavior, conversational behavior is what we'd call it and some behavior is what we'd call it and some behavior is what we'd call it and some optimization and a cost function called optimization and a cost function called optimization and a cost function called ouch. Um we have from the user's ouch. Um we have from the user's ouch. Um we have from the user's perspective produced a more satisfactory perspective produced a more satisfactory perspective produced a more satisfactory assistant. And this is not a trivial you assistant. And this is not a trivial you assistant. And this is not a trivial you know accomplishment. Okay. Now, let me know accomplishment. Okay. Now, let me know accomplishment. Okay. Now, let me go back to this. [clears throat] Let me go back to this. [clears throat] Let me go back to this. [clears throat] Let me see if I can get this. Oh, great. Okay, see if I can get this. Oh, great. Okay, see if I can get this. Oh, great. Okay, let's continue this. Let's continue this let's continue this. Let's continue this let's continue this. Let's continue this with by now adding one more behavior. with by now adding one more behavior. with by now adding one more behavior. Let's call it the confirm behavior. So, Let's call it the confirm behavior. So, Let's call it the confirm behavior. So, there was play obviously, then stop, there was play obviously, then stop, there was play obviously, then stop, confirm. Confirm is basically the system confirm. Confirm is basically the system confirm. Confirm is basically the system after you said something saying uh kiss after you said something saying uh kiss after you said something saying uh kiss play kiss by Prince or maybe play kiss play kiss by Prince or maybe play kiss play kiss by Prince or maybe play kiss by Chris Brown. And uh you know the user by Chris Brown. And uh you know the user by Chris Brown. And uh you know the user can either confirm like affirm it or can either confirm like affirm it or can either confirm like affirm it or they can correct it. It is a different they can correct it. It is a different they can correct it. It is a different kind of behavior and again this is kind kind of behavior and again this is kind kind of behavior and again this is kind of how humans behave. Um that's of how humans behave. Um that's of how humans behave. Um that's obviously the inspiration. Now if we go obviously the inspiration. Now if we go obviously the inspiration. Now if we go back to our problem of optimization, back to our problem of optimization, back to our problem of optimization, we have a third obviously um option we have a third obviously um option we have a third obviously um option which is to confirm. And so this would which is to confirm. And so this would which is to confirm. And so this would translate to two thresholds translate to two thresholds translate to two thresholds um two thresholds which are separating um two thresholds which are separating um two thresholds which are separating the distribution into three spaces of the distribution into three spaces of the distribution into three spaces of stop, confirm and uh act. And the stop, confirm and uh act. And the stop, confirm and uh act. And the question is now where are these T's? and question is now where are these T's? and question is now where are these T's? and we have now given up on guesstimating we have now given up on guesstimating we have now given up on guesstimating because we know it doesn't work. So because we know it doesn't work. So because we know it doesn't work. So we're going to be a lot smarter and go we're going to be a lot smarter and go we're going to be a lot smarter and go back to the concept of user outcome cost
-
back to the concept of user outcome cost back to the concept of user outcome cost and then you know use it go look for and then you know use it go look for and then you know use it go look for some optimization in that graph. So some optimization in that graph. So some optimization in that graph. So let's uh define what are the what are let's uh define what are the what are let's uh define what are the what are all the possible bad outcomes that t1 all the possible bad outcomes that t1 all the possible bad outcomes that t1 and t2 um make for. So good you can see and t2 um make for. So good you can see and t2 um make for. So good you can see my cursor. So uh of course any stops are my cursor. So uh of course any stops are my cursor. So uh of course any stops are still bad. Then in the middle are still bad. Then in the middle are still bad. Then in the middle are confirmations. Confirmations are bad confirmations. Confirmations are bad confirmations. Confirmations are bad because they slow the user down. There because they slow the user down. There because they slow the user down. There is a confirmation outcome called confirm is a confirmation outcome called confirm is a confirmation outcome called confirm yes where they just affirmed it by yes where they just affirmed it by yes where they just affirmed it by saying yeah or no where they had to saying yeah or no where they had to saying yeah or no where they had to correct it. And going back to our correct it. And going back to our correct it. And going back to our formula these outcomes are not equally formula these outcomes are not equally formula these outcomes are not equally bad. And in fact, nobody will, I think, bad. And in fact, nobody will, I think, bad. And in fact, nobody will, I think, argue here from a user's perspective. argue here from a user's perspective. argue here from a user's perspective. Affirming, just saying yes is obviously Affirming, just saying yes is obviously Affirming, just saying yes is obviously less painful than saying no and then less painful than saying no and then less painful than saying no and then having to restate whatever it is that having to restate whatever it is that having to restate whatever it is that you wanted in the first place. So now we you wanted in the first place. So now we you wanted in the first place. So now we I've assigned values of two or six. And I've assigned values of two or six. And I've assigned values of two or six. And again, I said it was a heristic. This again, I said it was a heristic. This again, I said it was a heristic. This would be roughly the amount of time it would be roughly the amount of time it would be roughly the amount of time it would take for the extra for the user to would take for the extra for the user to would take for the extra for the user to get to the song they want. Saying get to the song they want. Saying get to the song they want. Saying listening and then saying yes is like listening and then saying yes is like listening and then saying yes is like two seconds. Um and then now two seconds. Um and then now two seconds. Um and then now uh we restate the cost function for this uh we restate the cost function for this uh we restate the cost function for this you know added behavior as this number you know added behavior as this number you know added behavior as this number of you know bad type one times unit cost of you know bad type one times unit cost of you know bad type one times unit cost bad type plus bad type two times unit bad type plus bad type two times unit bad type plus bad type two times unit cost etc. And now we try to minimize cost etc. And now we try to minimize cost etc. And now we try to minimize this user cost function and minimize the this user cost function and minimize the this user cost function and minimize the ouch.
-
ouch. ouch. Yes, I'm going to keep doing that pun. Yes, I'm going to keep doing that pun. Yes, I'm going to keep doing that pun. Um let's go back. So this is now the Um let's go back. So this is now the Um let's go back. So this is now the interactive graph but interactive graph but interactive graph but with with with um the cost values the unit costs here um the cost values the unit costs here um the cost values the unit costs here 10264 10264 10264 and uh you know we're just going to ask and uh you know we're just going to ask and uh you know we're just going to ask the AI to tell us here's the heat map the AI to tell us here's the heat map the AI to tell us here's the heat map because it's now two dimensions saying because it's now two dimensions saying because it's now two dimensions saying that the optimal values are 41 for the that the optimal values are 41 for the that the optimal values are 41 for the the T1 and 49 for the T2 and if we the T1 and 49 for the T2 and if we the T1 and 49 for the T2 and if we employed that then we would go to 1464. employed that then we would go to 1464. employed that then we would go to 1464. Uh, by the way, whatever numbers I put Uh, by the way, whatever numbers I put Uh, by the way, whatever numbers I put in here, like let's say I thought wrong in here, like let's say I thought wrong in here, like let's say I thought wrong act was 20. It's really irritating and act was 20. It's really irritating and act was 20. It's really irritating and painful and takes way longer to actually painful and takes way longer to actually painful and takes way longer to actually correct it when you hear a wrong song. correct it when you hear a wrong song. correct it when you hear a wrong song. That would change you know all these That would change you know all these That would change you know all these numbers uh and the optim optimal point. numbers uh and the optim optimal point. numbers uh and the optim optimal point. So again it is about how what is the So again it is about how what is the So again it is about how what is the relative badness of these outcomes also relative badness of these outcomes also relative badness of these outcomes also of course the distribution curves of course the distribution curves of course the distribution curves naturally. Uh let's go back here. Okay.
-
naturally. Uh let's go back here. Okay. naturally. Uh let's go back here. Okay. So, um I'm gonna So, um I'm gonna So, um I'm gonna speed up a little bit. Uh let's go back speed up a little bit. Uh let's go back speed up a little bit. Uh let's go back here. here. here. Presentation mode. Okay. So, what have Presentation mode. Okay. So, what have Presentation mode. Okay. So, what have we shown that if we did the super naive we shown that if we did the super naive we shown that if we did the super naive approach, it's 2.1 act and stop 1.9 then approach, it's 2.1 act and stop 1.9 then approach, it's 2.1 act and stop 1.9 then 1.27 then 1.26. We are able to bring 1.27 then 1.26. We are able to bring 1.27 then 1.26. We are able to bring this with every added layer of this with every added layer of this with every added layer of sophistication, adding more behaviors, sophistication, adding more behaviors, sophistication, adding more behaviors, being smart about outcome, uh, user cost being smart about outcome, uh, user cost being smart about outcome, uh, user cost and optimizing. Um, we have made a and optimizing. Um, we have made a and optimizing. Um, we have made a tremendous difference without changing tremendous difference without changing tremendous difference without changing the accuracy at all. Um, this was a the accuracy at all. Um, this was a the accuracy at all. Um, this was a super simplified example. In real super simplified example. In real super simplified example. In real systems, you're not going to have systems, you're not going to have systems, you're not going to have obviously some offline decision obviously some offline decision obviously some offline decision threshold or two. It's going to be a threshold or two. It's going to be a threshold or two. It's going to be a real time, you know, learned decision real time, you know, learned decision real time, you know, learned decision model. But the principle is the same. model. But the principle is the same. model. But the principle is the same. And I believe this is uh scalable across And I believe this is uh scalable across And I believe this is uh scalable across all voice AI surfaces. Obviously this is all voice AI surfaces. Obviously this is all voice AI surfaces. Obviously this is a smart speaker but if we go across any a smart speaker but if we go across any a smart speaker but if we go across any of these surfaces you will find the of these surfaces you will find the of these surfaces you will find the equivalence. If we um we will find the equivalence. If we um we will find the equivalence. If we um we will find the analogies with some differences but the analogies with some differences but the analogies with some differences but the spirit and the I think the the gain will spirit and the I think the the gain will spirit and the I think the the gain will be similar. So just for example in the be similar. So just for example in the be similar. So just for example in the TV AI assistant space if you employ it TV AI assistant space if you employ it TV AI assistant space if you employ it here it's going to you're going to have here it's going to you're going to have here it's going to you're going to have the same thing when users express the same thing when users express the same thing when users express intents like on TV it's you know open a intents like on TV it's you know open a intents like on TV it's you know open a channel that's one of the most common
-
channel that's one of the most common channel that's one of the most common obviously um requests on a TV voice obviously um requests on a TV voice obviously um requests on a TV voice assistant same thing you're going to assistant same thing you're going to assistant same thing you're going to find you'll have exactly the same find you'll have exactly the same find you'll have exactly the same approach but the difference will be approach but the difference will be approach but the difference will be maybe in the the assignments of the user maybe in the the assignments of the user maybe in the the assignments of the user outcomes because the UI and the outcomes because the UI and the outcomes because the UI and the modalities are different when you have a modalities are different when you have a modalities are different when you have a TV you have a multimodal interface where TV you have a multimodal interface where TV you have a multimodal interface where choices can be shown. So instead of you choices can be shown. So instead of you choices can be shown. So instead of you know asking did you mean ABC you know uh know asking did you mean ABC you know uh know asking did you mean ABC you know uh news live by speech that you will the news live by speech that you will the news live by speech that you will the system would display choices and not system would display choices and not system would display choices and not just one it show ABC News live this that just one it show ABC News live this that just one it show ABC News live this that would be the confirm step and if it's would be the confirm step and if it's would be the confirm step and if it's visual and you can use your remote visual and you can use your remote visual and you can use your remote control to select something it's less control to select something it's less control to select something it's less pain so you would change some of these pain so you would change some of these pain so you would change some of these values or if in fact launching the values or if in fact launching the values or if in fact launching the channel would kick you out of your channel would kick you out of your channel would kick you out of your current state then it would go in the current state then it would go in the current state then it would go in the other direction than cost of you know a other direction than cost of you know a other direction than cost of you know a bad act would go much higher. So it's bad act would go much higher. So it's bad act would go much higher. So it's the same concept but in this new the same concept but in this new the same concept but in this new modalities modalities modalities um variables can change, values can um variables can change, values can um variables can change, values can change, arguments can change but the change, arguments can change but the change, arguments can change but the premise still holds and you can improve premise still holds and you can improve premise still holds and you can improve from the user's perspective because from the user's perspective because from the user's perspective because we're all about you know making humans we're all about you know making humans we're all about you know making humans happy. Um you can make them happier and happy. Um you can make them happier and happy. Um you can make them happier and this as I said in conclusion can be this as I said in conclusion can be this as I said in conclusion can be applied across all surfaces. I did say applied across all surfaces. I did say applied across all surfaces. I did say at the very beginning, just to recap for at the very beginning, just to recap for at the very beginning, just to recap for us, that voice is great when it works, us, that voice is great when it works, us, that voice is great when it works, bad when it doesn't. And as we get into bad when it doesn't. And as we get into bad when it doesn't. And as we get into embodied AI, where these AI assistants
-
embodied AI, where these AI assistants embodied AI, where these AI assistants are taking actions, physical or even are taking actions, physical or even are taking actions, physical or even digital, like making a phone call or digital, like making a phone call or digital, like making a phone call or sending an email, it is getting more and sending an email, it is getting more and sending an email, it is getting more and more difficult just to rely on accuracy more difficult just to rely on accuracy more difficult just to rely on accuracy to improve user satisfaction. I believe to improve user satisfaction. I believe to improve user satisfaction. I believe there's a whole knob the second knob there's a whole knob the second knob there's a whole knob the second knob called smarter conversational behavior called smarter conversational behavior called smarter conversational behavior under uncertainty under uncertainty under uncertainty and um if we actually exploit that we and um if we actually exploit that we and um if we actually exploit that we can uh very much help these AI systems can uh very much help these AI systems can uh very much help these AI systems reach a acceptable user experience reach a acceptable user experience reach a acceptable user experience otherwise I think this will continue to otherwise I think this will continue to otherwise I think this will continue to be a bottleneck like a lot of things be a bottleneck like a lot of things be a bottleneck like a lot of things will get better but if the voice will get better but if the voice will get better but if the voice interface as experienced by user does interface as experienced by user does interface as experienced by user does not improve it is going to be a a a not improve it is going to be a a a not improve it is going to be a a a choke point. And um if you just remember choke point. And um if you just remember choke point. And um if you just remember one word or two words from this whole um one word or two words from this whole um one word or two words from this whole um presentation, it would be to minimize presentation, it would be to minimize presentation, it would be to minimize the ouch of the experience. Um so thank the ouch of the experience. Um so thank the ouch of the experience. Um so thank you. I'll stick around for questions if you. I'll stick around for questions if you. I'll stick around for questions if you guys got any. Thanks a lot.
-
you guys got any. Thanks a lot. you guys got any. Thanks a lot. [applause]
No summary available yet.
View original episode ↗