← Back
AI Engineer September 15, 2026 19m

Your Voice Agent is Just a Walkie Talkie — Neil Zeghidour, Gradium

Read full transcript 17 segments
  1. >> Okay, hi everyone. >> Okay, hi everyone. I'm Nel I'm Nel I'm Nel co-founder and CEO of of Gradio. co-founder and CEO of of Gradio. co-founder and CEO of of Gradio. So Gradio is a startup based in Paris. So Gradio is a startup based in Paris. So Gradio is a startup based in Paris. Most of our background is from research. Most of our background is from research. Most of our background is from research. In particular, we have invented In particular, we have invented In particular, we have invented algorithms such as audio LLMs, speech algorithms such as audio LLMs, speech algorithms such as audio LLMs, speech speech-to-speech models, neural codex, speech-to-speech models, neural codex, speech-to-speech models, neural codex, and so on and so forth. and so on and so forth. and so on and so forth. And And And basically basically basically we started from a research project we started from a research project we started from a research project called QTI, a non-profit research lab called QTI, a non-profit research lab called QTI, a non-profit research lab that has been focusing on voice since that has been focusing on voice since that has been focusing on voice since day one. So in particular, we released day one. So in particular, we released day one. So in particular, we released in 2024 the first in 2024 the first in 2024 the first full duplex speech-to-speech model full duplex speech-to-speech model full duplex speech-to-speech model called Moshi, the first real-time called Moshi, the first real-time called Moshi, the first real-time speech-to-speech translation system speech-to-speech translation system speech-to-speech translation system called Hibiki, and the first called Hibiki, and the first called Hibiki, and the first TTS model that can run locally on a on a TTS model that can run locally on a on a TTS model that can run locally on a on a smartphone. smartphone. smartphone. And basically I will just say a few And basically I will just say a few And basically I will just say a few words about what we do, but we are words about what we do, but we are words about what we do, but we are a model company that trains models for a model company that trains models for a model company that trains models for building voice agents and voice building voice agents and voice building voice agents and voice applications. So we do TTS, API, and applications. So we do TTS, API, and applications. So we do TTS, API, and on-device speech-to-text, on-device speech-to-text, on-device speech-to-text, speech-to-speech translation, and much speech-to-speech translation, and much speech-to-speech translation, and much more to come. What we do is that we more to come. What we do is that we more to come. What we do is that we train foundation models for audio, and train foundation models for audio, and train foundation models for audio, and then we can apply them for a lot of then we can apply them for a lot of then we can apply them for a lot of different tasks.

  2. different tasks. different tasks. So I go quickly on voice agents because So I go quickly on voice agents because So I go quickly on voice agents because it's the fourth talk it's the fourth talk it's the fourth talk about the topic, but basically now we about the topic, but basically now we about the topic, but basically now we have these have these have these voice interfaces that we can use to do a voice interfaces that we can use to do a voice interfaces that we can use to do a lot of things across a variety of lot of things across a variety of lot of things across a variety of products and types of interactions with products and types of interactions with products and types of interactions with NPCs, with customer agents, language NPCs, with customer agents, language NPCs, with customer agents, language learners, coach, and so on and so forth. learners, coach, and so on and so forth. learners, coach, and so on and so forth. And in this talk I tried to go through And in this talk I tried to go through And in this talk I tried to go through the history of this technology and the history of this technology and the history of this technology and where I see it going in the next years. where I see it going in the next years. where I see it going in the next years. And maybe to start, I think we can take And maybe to start, I think we can take And maybe to start, I think we can take a look at the announcement of Siri back a look at the announcement of Siri back a look at the announcement of Siri back in 2011. in 2011. in 2011. And you'll see that it's actually, you And you'll see that it's actually, you And you'll see that it's actually, you know, I think it it aged pretty well. >> Here's the forecast for today. >> Here's the forecast for today. >> It is THAT EASY. >> It is THAT EASY. >> It is THAT EASY. >> [cheering] >> [cheering] >> [cheering] [applause] [applause] [applause] >> LOTS OF THINGS. We've integrated with >> LOTS OF THINGS. We've integrated with >> LOTS OF THINGS. We've integrated with the stocks. So, you can ask it about the the stocks. So, you can ask it about the the stocks. So, you can ask it about the stock market. Something like stock market. Something like stock market. Something like How is the NASDAQ doing today?

  3. >> NASDAQ composite is down right now at >> NASDAQ composite is down right now at 2,321.70. >> Again, you can ask this from the lock >> Again, you can ask this from the lock screen anywhere. Just press the button screen anywhere. Just press the button screen anywhere. Just press the button and ask. You can ask about, you know, and ask. You can ask about, you know, and ask. You can ask about, you know, the NASDAQ, the Dow. the NASDAQ, the Dow. the NASDAQ, the Dow. >> So, >> So, >> So, what you just saw is what kind of a what you just saw is what kind of a what you just saw is what kind of a voice agent. It was a bit constrained, voice agent. It was a bit constrained, voice agent. It was a bit constrained, but it was technically a voice agent. but it was technically a voice agent. but it was technically a voice agent. And the architecture behind it, so you And the architecture behind it, so you And the architecture behind it, so you have seen a thousand times today the STT have seen a thousand times today the STT have seen a thousand times today the STT LLM TTS. Back then, it was even worse, LLM TTS. Back then, it was even worse, LLM TTS. Back then, it was even worse, right? So, there was no LLM, obviously. right? So, there was no LLM, obviously. right? So, there was no LLM, obviously. So, there was what was called natural So, there was what was called natural So, there was what was called natural language understanding. So, you would go language understanding. So, you would go language understanding. So, you would go from the transcript and uh try to do from the transcript and uh try to do from the transcript and uh try to do basic classification of what is basic classification of what is basic classification of what is expected, uh what is the app that is expected, uh what is the app that is expected, uh what is the app that is supposed to be uh controlled, what is supposed to be uh controlled, what is supposed to be uh controlled, what is the action to trigger, and so on. So, it the action to trigger, and so on. So, it the action to trigger, and so on. So, it was a very complex pipeline and very was a very complex pipeline and very was a very complex pipeline and very constrained to very specific use cases. constrained to very specific use cases. constrained to very specific use cases. So, it was what was called close So, it was what was called close So, it was what was called close uh ended um uh ended um uh ended um uh dialogue back then. uh dialogue back then. uh dialogue back then. Then, the next generation of models, uh Then, the next generation of models, uh Then, the next generation of models, uh I could say, I think was I could say, I think was I could say, I think was is illustrated pretty well by the OpenAI is illustrated pretty well by the OpenAI is illustrated pretty well by the OpenAI voice mode, not the advanced voice mode, voice mode, not the advanced voice mode, voice mode, not the advanced voice mode, the original one. In which case, you the original one. In which case, you the original one. In which case, you just have um a conversational experience just have um a conversational experience just have um a conversational experience that is open-ended. So, it's much less that is open-ended. So, it's much less that is open-ended. So, it's much less constrained than Siri. You can talk constrained than Siri. You can talk constrained than Siri. You can talk about pretty much anything. At the same about pretty much anything. At the same about pretty much anything. At the same time, you don't have um time, you don't have um time, you don't have um as the same agentic capabilities, let's as the same agentic capabilities, let's as the same agentic capabilities, let's say, in the sense that it cannot really say, in the sense that it cannot really say, in the sense that it cannot really give you the NASDAQ or uh give you the give you the NASDAQ or uh give you the give you the NASDAQ or uh give you the weather.

  4. >> We want to hear a bedtime story. >> We want to hear a bedtime story. Tell us a story about the super-duper Tell us a story about the super-duper Tell us a story about the super-duper sunflower hedgehog named Larry. sunflower hedgehog named Larry. sunflower hedgehog named Larry. Start with telling us a little bit about Start with telling us a little bit about Start with telling us a little bit about him. Larry was a unique hedgehog unlike any Larry was a unique hedgehog unlike any other. other. other. He had bright sunflower petals instead He had bright sunflower petals instead He had bright sunflower petals instead of spines. Residing in Meadowville, of spines. Residing in Meadowville, of spines. Residing in Meadowville, Larry spread joy and color wherever he Larry spread joy and color wherever he Larry spread joy and color wherever he rolled. rolled. rolled. >> So, in that context, basically that was, >> So, in that context, basically that was, >> So, in that context, basically that was, you know, kind of second generation of you know, kind of second generation of you know, kind of second generation of models where all the logic, the very models where all the logic, the very models where all the logic, the very complex logic was replaced complex logic was replaced complex logic was replaced uh by an LLM. uh by an LLM. uh by an LLM. And And And if we compare Siri and the voice mode if we compare Siri and the voice mode if we compare Siri and the voice mode from OpenAI, uh from OpenAI, uh from OpenAI, uh not the Ad Ons one, Ad Ons one, the not the Ad Ons one, Ad Ons one, the not the Ad Ons one, Ad Ons one, the original one, original one, original one, uh uh uh the big strength is opening the the big strength is opening the the big strength is opening the dialogue, but the agentic capabilities dialogue, but the agentic capabilities dialogue, but the agentic capabilities were inexistent. The naturalness was, were inexistent. The naturalness was, were inexistent. The naturalness was, obviously, much better than with Siri. obviously, much better than with Siri. obviously, much better than with Siri. The latency was still pretty bad. So, we The latency was still pretty bad. So, we The latency was still pretty bad. So, we are talking about several seconds, and are talking about several seconds, and are talking about several seconds, and that was documented in the in the that was documented in the in the that was documented in the in the previous presentation. previous presentation. previous presentation. And so, then, you know, we're in this And so, then, you know, we're in this And so, then, you know, we're in this third phase, which is a voice, what we third phase, which is a voice, what we third phase, which is a voice, what we call voice agents. Uh and here I'm going call voice agents. Uh and here I'm going call voice agents. Uh and here I'm going to show a quick demo of uh of a to show a quick demo of uh of a to show a quick demo of uh of a drive-thru bot uh live-coded with our drive-thru bot uh live-coded with our drive-thru bot uh live-coded with our with our voice models.

  5. with our voice models. with our voice models. >> Hi there. What can I get started for you >> Hi there. What can I get started for you >> Hi there. What can I get started for you today? today? today? >> Hi, I'm actually in a hurry and I'd like >> Hi, I'm actually in a hurry and I'd like >> Hi, I'm actually in a hurry and I'd like to order a chicken sandwich quickly. to order a chicken sandwich quickly. to order a chicken sandwich quickly. >> We have the classic, spicy, deluxe, or >> We have the classic, spicy, deluxe, or >> We have the classic, spicy, deluxe, or grilled chicken sandwich. Which one grilled chicken sandwich. Which one grilled chicken sandwich. Which one would you like? would you like? would you like? >> Uh let's go with the classic chicken >> Uh let's go with the classic chicken >> Uh let's go with the classic chicken sandwich. sandwich. sandwich. >> Anything else for you? >> Anything else for you? >> Anything else for you? >> Uh what else do you have on the menu? >> Uh what else do you have on the menu? >> Uh what else do you have on the menu? >> We have sandwiches, sides, drinks, and >> We have sandwiches, sides, drinks, and >> We have sandwiches, sides, drinks, and desserts. What can I get for you? desserts. What can I get for you? desserts. What can I get for you? >> Can I get some mac and cheese, maybe? >> Can I get some mac and cheese, maybe? >> Can I get some mac and cheese, maybe? >> You have a classic chicken sandwich and >> You have a classic chicken sandwich and >> You have a classic chicken sandwich and mac and cheese. mac and cheese. mac and cheese. Your total comes to eight. Your total comes to eight. Your total comes to eight. >> So, now you know, unlike what we just >> So, now you know, unlike what we just >> So, now you know, unlike what we just saw, it's an actual agent. It's taking saw, it's an actual agent. It's taking saw, it's an actual agent. It's taking actions. It's keeping track of the actions. It's keeping track of the actions. It's keeping track of the order. It's then going to make you pay. order. It's then going to make you pay. order. It's then going to make you pay. So, it's it's an actual voice agent that So, it's it's an actual voice agent that So, it's it's an actual voice agent that can do uh real tasks. So, here instead can do uh real tasks. So, here instead can do uh real tasks. So, here instead of having an LLM that is just a of having an LLM that is just a of having an LLM that is just a conversational interface, conversational interface, conversational interface, we have a real agent that is empowered we have a real agent that is empowered we have a real agent that is empowered with tool call, reasoning, planning, and with tool call, reasoning, planning, and with tool call, reasoning, planning, and and all this stuff. and all this stuff. and all this stuff. So, So, So, what we see now is we have gained back what we see now is we have gained back what we see now is we have gained back agentic capabilities, and actually they agentic capabilities, and actually they agentic capabilities, and actually they are much more are much more are much more uh powerful and generic than before, uh powerful and generic than before, uh powerful and generic than before, while keeping a very good level of uh of while keeping a very good level of uh of while keeping a very good level of uh of naturalness.

  6. naturalness. naturalness. And And And that's where speech-to-speech LLM came. that's where speech-to-speech LLM came. that's where speech-to-speech LLM came. In particular, what we could see here is In particular, what we could see here is In particular, what we could see here is the latency, it's better with cascaded the latency, it's better with cascaded the latency, it's better with cascaded system, but it's still higher than you system, but it's still higher than you system, but it's still higher than you will have with human conversation. And will have with human conversation. And will have with human conversation. And as also was explained before, the as also was explained before, the as also was explained before, the naturalness is fundamentally limited by naturalness is fundamentally limited by naturalness is fundamentally limited by the fact that you go through text, so the fact that you go through text, so the fact that you go through text, so you lose a lot of information about what you lose a lot of information about what you lose a lot of information about what uh is said, the tone, the emotion of the uh is said, the tone, the emotion of the uh is said, the tone, the emotion of the user, and so on and so forth. user, and so on and so forth. user, and so on and so forth. So, now that we have tackled So, now that we have tackled So, now that we have tackled intelligence and agentic capabilities, intelligence and agentic capabilities, intelligence and agentic capabilities, speech-to-speech seems like a natural speech-to-speech seems like a natural speech-to-speech seems like a natural next step for naturalness and latency. next step for naturalness and latency. next step for naturalness and latency. And so here it's the announcement from And so here it's the announcement from And so here it's the announcement from the uh OpenAI advanced voice mode. the uh OpenAI advanced voice mode. the uh OpenAI advanced voice mode. >> [clears throat] >> [clears throat] >> [clears throat] >> Hey, ChatGPT. I'm Mark. How are you? >> Hey, ChatGPT. I'm Mark. How are you? >> Hey, ChatGPT. I'm Mark. How are you? >> Oh, Mark. >> Oh, Mark. >> Oh, Mark. I'm doing great. Thanks for asking. How I'm doing great. Thanks for asking. How I'm doing great. Thanks for asking. How about you? about you? about you? >> Hey, so I'm on stage right now. I'm >> Hey, so I'm on stage right now. I'm >> Hey, so I'm on stage right now. I'm doing a live demo, and frankly I'm doing a live demo, and frankly I'm doing a live demo, and frankly I'm feeling a little bit nervous. Can you feeling a little bit nervous. Can you feeling a little bit nervous. Can you help me calm my nerves a little bit? help me calm my nerves a little bit? help me calm my nerves a little bit? >> Oh, you're doing a live demo right now? >> Oh, you're doing a live demo right now? >> Oh, you're doing a live demo right now? That's awesome. That's awesome. That's awesome. Just Just Just >> I think we all remember it was very >> I think we all remember it was very >> I think we all remember it was very impressive very impressive release.

  7. impressive very impressive release. impressive very impressive release. And in that context now, all the steps And in that context now, all the steps And in that context now, all the steps of STT, LLM, and TTS have been absorbed of STT, LLM, and TTS have been absorbed of STT, LLM, and TTS have been absorbed into a a single one. into a a single one. into a a single one. And so now, And so now, And so now, intelligence, you know, like naturalness intelligence, you know, like naturalness intelligence, you know, like naturalness is is is still very good. Uh still very good. Uh still very good. Uh actually it can be better because it can actually it can be better because it can actually it can be better because it can understand non-linguistic information. understand non-linguistic information. understand non-linguistic information. Latency is really, really nice. Latency is really, really nice. Latency is really, really nice. Honestly, it doesn't make sense to go uh Honestly, it doesn't make sense to go uh Honestly, it doesn't make sense to go uh better than that. better than that. better than that. Interestingly and everyone was used any Interestingly and everyone was used any Interestingly and everyone was used any uh speech-to-speech model can uh attest uh speech-to-speech model can uh attest uh speech-to-speech model can uh attest that that that the intelligence the intelligence the intelligence is still much more limited in that is still much more limited in that is still much more limited in that context than uh the cascaded context than uh the cascaded context than uh the cascaded counterpart. So, the speech-to-speech counterpart. So, the speech-to-speech counterpart. So, the speech-to-speech models are fundamentally still limited models are fundamentally still limited models are fundamentally still limited compared to the textual models. compared to the textual models. compared to the textual models. Another limitation is turn-taking. So, Another limitation is turn-taking. So, Another limitation is turn-taking. So, people tend to people tend to people tend to mix speech-to-speech and full duplex. mix speech-to-speech and full duplex. mix speech-to-speech and full duplex. But basically, But basically, But basically, when you do have a speech-to-speech when you do have a speech-to-speech when you do have a speech-to-speech model like GPT-3 time, it's still based model like GPT-3 time, it's still based model like GPT-3 time, it's still based on fundamental turn-taking. In the sense on fundamental turn-taking. In the sense on fundamental turn-taking. In the sense that it's going to segment the that it's going to segment the that it's going to segment the conversation into as long as the model conversation into as long as the model conversation into as long as the model is speaking or the model is listening.

  8. is speaking or the model is listening. is speaking or the model is listening. And to give to show you how this can And to give to show you how this can And to give to show you how this can make an interaction unnatural, I'm going make an interaction unnatural, I'm going make an interaction unnatural, I'm going to show a a small demo with what is to show a a small demo with what is to show a a small demo with what is called backchanneling, which is this called backchanneling, which is this called backchanneling, which is this very human thing that you do when very human thing that you do when very human thing that you do when someone talks to you is that you say, someone talks to you is that you say, someone talks to you is that you say, "Mhm, yeah." and so on. "Mhm, yeah." and so on. "Mhm, yeah." and so on. >> Hey, how's it going? Just like to >> Hey, how's it going? Just like to >> Hey, how's it going? Just like to brainstorm a bit about it with you. brainstorm a bit about it with you. brainstorm a bit about it with you. >> Oh, that's a great topic. Yeah, I'd love >> Oh, that's a great topic. Yeah, I'd love >> Oh, that's a great topic. Yeah, I'd love to help you brainstorm. Are you thinking to help you brainstorm. Are you thinking to help you brainstorm. Are you thinking what what what Exactly. Yeah, I was thinking Exactly. Yeah, I was thinking Exactly. Yeah, I was thinking >> No, no, I didn't mean to interrupt, you >> No, no, I didn't mean to interrupt, you >> No, no, I didn't mean to interrupt, you know, I was just saying, "Yeah." like know, I was just saying, "Yeah." like know, I was just saying, "Yeah." like that. You can just keep going. I know, that. You can just keep going. I know, that. You can just keep going. I know, don't mind me. It's just something I don't mind me. It's just something I don't mind me. It's just something I typically do. typically do. typically do. >> Uh no worries at all. Yeah, I was just >> Uh no worries at all. Yeah, I was just >> Uh no worries at all. Yeah, I was just going to say we could break it down into going to say we could break it down into going to say we could break it down into a few aspects. Like a few aspects. Like a few aspects. Like Yeah, exactly. Yeah, exactly. Yeah, exactly. >> please stop stop interrupting. You know, >> please stop stop interrupting. You know, >> please stop stop interrupting. You know, it's called backchanneling. Humans do it it's called backchanneling. Humans do it it's called backchanneling. Humans do it all the time. It shows that you're just all the time. It shows that you're just all the time. It shows that you're just following the conversation. That you following the conversation. That you following the conversation. That you don't, don't, don't, you know, like interrupt you in your you know, like interrupt you in your you know, like interrupt you in your flow. Just just going. flow. Just just going. flow. Just just going. >> Ah, got it. Thanks for letting me know. >> Ah, got it. Thanks for letting me know. >> Ah, got it. Thanks for letting me know. >> No problem. >> No problem. >> No problem. Oh, come on. Oh, come on. Oh, come on. >> Yeah, so you see, you know, it's it's >> Yeah, so you see, you know, it's it's >> Yeah, so you see, you know, it's it's still very annoying. Uh you can have still very annoying. Uh you can have still very annoying. Uh you can have lightning speed latency. Fundamentally, lightning speed latency. Fundamentally, lightning speed latency. Fundamentally, this is this is this is uh an issue that can not be resolved uh an issue that can not be resolved uh an issue that can not be resolved when you're using turn-taking. So, here when you're using turn-taking. So, here when you're using turn-taking. So, here that's the walkie-talkie.

  9. that's the walkie-talkie. that's the walkie-talkie. Um any real-time model today, I mean, Um any real-time model today, I mean, Um any real-time model today, I mean, now there is a bidirectional one that now there is a bidirectional one that now there is a bidirectional one that will come from OpenAI, but it's called will come from OpenAI, but it's called will come from OpenAI, but it's called half duplex. So, the model is listening half duplex. So, the model is listening half duplex. So, the model is listening or speaking. A human conversation or speaking. A human conversation or speaking. A human conversation has a constant flow between two people. has a constant flow between two people. has a constant flow between two people. People do back channeling. People People do back channeling. People People do back channeling. People interrupt one another, talk on one interrupt one another, talk on one interrupt one another, talk on one another, and so on. another, and so on. another, and so on. If you have If you're having a relative If you have If you're having a relative If you have If you're having a relative on the phone, there is up to 20% of the on the phone, there is up to 20% of the on the phone, there is up to 20% of the time where you are both speaking at the time where you are both speaking at the time where you are both speaking at the same time. same time. same time. And that makes, you know, this very And that makes, you know, this very And that makes, you know, this very flexible dynamics in the conversation flexible dynamics in the conversation flexible dynamics in the conversation makes it much more comfortable for makes it much more comfortable for makes it much more comfortable for humans. humans. humans. And so, to understand how And so, to understand how And so, to understand how we can make a model full duplex, I'll we can make a model full duplex, I'll we can make a model full duplex, I'll give a very short give a very short give a very short presentation of how we train such presentation of how we train such presentation of how we train such models. So, the way you create a models. So, the way you create a models. So, the way you create a speech-to-speech model half duplex or speech-to-speech model half duplex or speech-to-speech model half duplex or full duplex is the following one. So, full duplex is the following one. So, full duplex is the following one. So, you you start from a text LLM, which is you you start from a text LLM, which is you you start from a text LLM, which is a probabilistic models over over words. a probabilistic models over over words. a probabilistic models over over words. And instead of predicting the next word And instead of predicting the next word And instead of predicting the next word based on the past, based on the past, based on the past, what you want to do is rather predict what you want to do is rather predict what you want to do is rather predict the next audio based based on the past the next audio based based on the past the next audio based based on the past audio.

  10. audio. audio. The issue now is that if you pass a raw The issue now is that if you pass a raw The issue now is that if you pass a raw audio to your model, which is, you know, audio to your model, which is, you know, audio to your model, which is, you know, a waveform, it's a waveform, it's a waveform, it's air pressure variations. air pressure variations. air pressure variations. Uh Uh Uh basically, you take this sentence, it's basically, you take this sentence, it's basically, you take this sentence, it's eight words. eight words. eight words. It takes around 3 seconds to pronounce It takes around 3 seconds to pronounce It takes around 3 seconds to pronounce it. And so, at 24 kHz audio, instead of it. And so, at 24 kHz audio, instead of it. And so, at 24 kHz audio, instead of having eight words, the audio form is having eight words, the audio form is having eight words, the audio form is 72,000 time steps that you would need to 72,000 time steps that you would need to 72,000 time steps that you would need to feed to your LLM. Given that LLMs have feed to your LLM. Given that LLMs have feed to your LLM. Given that LLMs have quadratic complexity with sequence quadratic complexity with sequence quadratic complexity with sequence length, so the complexity is the square length, so the complexity is the square length, so the complexity is the square of the sequence length. A 10,000 times of the sequence length. A 10,000 times of the sequence length. A 10,000 times longer sequence is 100 million times longer sequence is 100 million times longer sequence is 100 million times more expensive to to process. So, there more expensive to to process. So, there more expensive to to process. So, there is no way you can train an LLM on raw is no way you can train an LLM on raw is no way you can train an LLM on raw audio. So, the way you address it is by audio. So, the way you address it is by audio. So, the way you address it is by creating neural codecs, or you can also creating neural codecs, or you can also creating neural codecs, or you can also call them audio tokenizers. And call them audio tokenizers. And call them audio tokenizers. And basically, it's an encoder that takes an basically, it's an encoder that takes an basically, it's an encoder that takes an audio and compresses it in a very dense audio and compresses it in a very dense audio and compresses it in a very dense compressed representation, a bit similar compressed representation, a bit similar compressed representation, a bit similar to text. And then you have a decoder to text. And then you have a decoder to text. And then you have a decoder that can reconstruct high-quality audio that can reconstruct high-quality audio that can reconstruct high-quality audio from it. So, now from it. So, now from it. So, now you have gone from the audio domain into you have gone from the audio domain into you have gone from the audio domain into a abstract representation domain, where a abstract representation domain, where a abstract representation domain, where you can train an LLM exactly like you you can train an LLM exactly like you you can train an LLM exactly like you would train it on text.

  11. would train it on text. would train it on text. And the speech-to-speech model from And the speech-to-speech model from And the speech-to-speech model from ElevenLabs, as I was showing before, ElevenLabs, as I was showing before, ElevenLabs, as I was showing before, works in this fashion. So, instead of works in this fashion. So, instead of works in this fashion. So, instead of having text tokens into your model, you having text tokens into your model, you having text tokens into your model, you have audio tokens that represent either have audio tokens that represent either have audio tokens that represent either the LLM or the user, and you put them the LLM or the user, and you put them the LLM or the user, and you put them one after the other, and the model one after the other, and the model one after the other, and the model predicts the audio tokens uh that should predicts the audio tokens uh that should predicts the audio tokens uh that should be said by the model, being given the be said by the model, being given the be said by the model, being given the context from both sides of the context from both sides of the context from both sides of the conversation. conversation. conversation. However, you can see that it's still a However, you can see that it's still a However, you can see that it's still a sequence between user and system, which sequence between user and system, which sequence between user and system, which is still half duplex. So, how did we is still half duplex. So, how did we is still half duplex. So, how did we make the first full duplex model ever? make the first full duplex model ever? make the first full duplex model ever? Very simple. We call it multi-stream Very simple. We call it multi-stream Very simple. We call it multi-stream language models. That's the technology language models. That's the technology language models. That's the technology now used also by Thinking Machines for now used also by Thinking Machines for now used also by Thinking Machines for their interaction model, and most likely their interaction model, and most likely their interaction model, and most likely by the for the by the directional model by the for the by the directional model by the for the by the directional model of OpenAI. Is that instead of having a of OpenAI. Is that instead of having a of OpenAI. Is that instead of having a transformer that models one sequence of transformer that models one sequence of transformer that models one sequence of tokens, it models two of them, so that tokens, it models two of them, so that tokens, it models two of them, so that both parties can be active at the same both parties can be active at the same both parties can be active at the same time, inactive at the same time, one time, inactive at the same time, one time, inactive at the same time, one active and one inactive. And active and one inactive. And active and one inactive. And I just show a very quick demo of uh of I just show a very quick demo of uh of I just show a very quick demo of uh of how it sounds [snorts] like, but that's how it sounds [snorts] like, but that's how it sounds [snorts] like, but that's the release of machine August 2024, the release of machine August 2024, the release of machine August 2024, uh where we did an announcement live on uh where we did an announcement live on uh where we did an announcement live on stage talking to it for the first time.

  12. stage talking to it for the first time. stage talking to it for the first time. And you'll see that the model often And you'll see that the model often And you'll see that the model often guesses the end of the question, answers guesses the end of the question, answers guesses the end of the question, answers over the speaker, over the speaker, over the speaker, and both speaking at the same time is and both speaking at the same time is and both speaking at the same time is not breaking the flow like we saw with not breaking the flow like we saw with not breaking the flow like we saw with GPT. The whole thing is just extremely GPT. The whole thing is just extremely GPT. The whole thing is just extremely resilient to the most chaotic uh resilient to the most chaotic uh resilient to the most chaotic uh situations. situations. situations. >> So, the planet is serious 22. Can you >> So, the planet is serious 22. Can you >> So, the planet is serious 22. Can you plot a trajectory course to it, please? plot a trajectory course to it, please? plot a trajectory course to it, please? >> Yes, sir. >> Yes, sir. >> Yes, sir. >> Okay. How long is it going to take us to >> Okay. How long is it going to take us to >> Okay. How long is it going to take us to get there? get there? get there? >> it out. It's approximately 5 months to >> it out. It's approximately 5 months to >> it out. It's approximately 5 months to get there. get there. get there. >> Okay, that's that's not too bad. Uh do >> Okay, that's that's not too bad. Uh do >> Okay, that's that's not too bad. Uh do you think we have all we need on board you think we have all we need on board you think we have all we need on board the ship to start the mission? the ship to start the mission? the ship to start the mission? >> We have everything we need. >> We have everything we need. >> We have everything we need. >> So, back then it was even a bit >> So, back then it was even a bit >> So, back then it was even a bit irritating to people because it was irritating to people because it was irritating to people because it was interrupting you all the time. But the interrupting you all the time. But the interrupting you all the time. But the thing is that you can use you could thing is that you can use you could thing is that you can use you could still use it in extremely noisy still use it in extremely noisy still use it in extremely noisy environments with a lot of noise, people environments with a lot of noise, people environments with a lot of noise, people coughing, and so on. And you know, the coughing, and so on. And you know, the coughing, and so on. And you know, the flow is just constant. You don't get flow is just constant. You don't get flow is just constant. You don't get this very irritating break of the this very irritating break of the this very irritating break of the conversational flow. So, conversational flow. So, conversational flow. So, these full duplex models, they are the these full duplex models, they are the these full duplex models, they are the highest level of naturalness you can highest level of naturalness you can highest level of naturalness you can expect. That's the same conversation expect. That's the same conversation expect. That's the same conversation with a human.

  13. with a human. with a human. The thing is, in with our models, it was The thing is, in with our models, it was The thing is, in with our models, it was even more stupid than even more stupid than even more stupid than speech-to-speech models that were speech-to-speech models that were speech-to-speech models that were already less intelligent than cascaded already less intelligent than cascaded already less intelligent than cascaded systems. systems. systems. It's probably fine for some use cases if It's probably fine for some use cases if It's probably fine for some use cases if you just want to have a chit-chat. You you just want to have a chit-chat. You you just want to have a chit-chat. You know, the model doesn't need to be very know, the model doesn't need to be very know, the model doesn't need to be very intelligent. But make an actual full intelligent. But make an actual full intelligent. But make an actual full duplex voice agent, duplex voice agent, duplex voice agent, there is no way we can give up on on there is no way we can give up on on there is no way we can give up on on intelligence just to gain uh intelligence just to gain uh intelligence just to gain uh speech-to-speech abilities. speech-to-speech abilities. speech-to-speech abilities. So, how do we finally make models that So, how do we finally make models that So, how do we finally make models that tackle all these aspects jointly? tackle all these aspects jointly? tackle all these aspects jointly? And I think interestingly, if you if you And I think interestingly, if you if you And I think interestingly, if you if you look at the history I showed, there is a look at the history I showed, there is a look at the history I showed, there is a tension between naturalness and tension between naturalness and tension between naturalness and intelligence. So, every time we improve intelligence. So, every time we improve intelligence. So, every time we improve naturalness or humanness of the of the naturalness or humanness of the of the naturalness or humanness of the of the models, they were less intelligent than models, they were less intelligent than models, they were less intelligent than the cascaded system. The cascaded the cascaded system. The cascaded the cascaded system. The cascaded agents, they are basically as smart as agents, they are basically as smart as agents, they are basically as smart as the best text models. So, if you have a the best text models. So, if you have a the best text models. So, if you have a voice agent that is powered by the voice agent that is powered by the voice agent that is powered by the latest model from Anthropic or OpenAI, latest model from Anthropic or OpenAI, latest model from Anthropic or OpenAI, it's going to be extremely smart, have it's going to be extremely smart, have it's going to be extremely smart, have all the same reliability for tool call, all the same reliability for tool call, all the same reliability for tool call, and so on. Speech-to-speech has this and so on. Speech-to-speech has this and so on. Speech-to-speech has this naturalness naturalness naturalness aspect. However, you give up aspect. However, you give up aspect. However, you give up intelligence to get that. And the reason intelligence to get that. And the reason intelligence to get that. And the reason why you give up intelligence is remember why you give up intelligence is remember why you give up intelligence is remember that the LLM is a model that has a that the LLM is a model that has a that the LLM is a model that has a certain number of weights that we call certain number of weights that we call certain number of weights that we call the capacity.

  14. the capacity. the capacity. And if you take a text model and now it And if you take a text model and now it And if you take a text model and now it not only has to handle text, but it also not only has to handle text, but it also not only has to handle text, but it also needs to understand speech and produce needs to understand speech and produce needs to understand speech and produce speech, speech, speech, it's taking some of its capacity, and it's taking some of its capacity, and it's taking some of its capacity, and this capacity now this capacity now this capacity now is taken from the intelligence. So, is taken from the intelligence. So, is taken from the intelligence. So, fundamentally, there is a cost of adding fundamentally, there is a cost of adding fundamentally, there is a cost of adding a new modality to a text model that is a new modality to a text model that is a new modality to a text model that is going to be paid in in intelligence. going to be paid in in intelligence. going to be paid in in intelligence. So, where do we go from here? There are So, where do we go from here? There are So, where do we go from here? There are two paths that are in front of us, and two paths that are in front of us, and two paths that are in front of us, and both are going to be explored at the both are going to be explored at the both are going to be explored at the same time. The first one is scaling the same time. The first one is scaling the same time. The first one is scaling the model. So, model. So, model. So, making your speech-to-speech model making your speech-to-speech model making your speech-to-speech model bigger, better pre-trained, better bigger, better pre-trained, better bigger, better pre-trained, better post-trained, and so on. We likely post-trained, and so on. We likely post-trained, and so on. We likely progressively increase its intelligence progressively increase its intelligence progressively increase its intelligence until it it's good enough for a lot of until it it's good enough for a lot of until it it's good enough for a lot of use cases. use cases. use cases. The second one is splitting the model The second one is splitting the model The second one is splitting the model between naturalness and intelligence. between naturalness and intelligence. between naturalness and intelligence. The first one, The first one, The first one, I'm not at OpenAI, so I don't know I'm not at OpenAI, so I don't know I'm not at OpenAI, so I don't know because they don't release their model. because they don't release their model. because they don't release their model. I guess OpenAI is the path one, so it's I guess OpenAI is the path one, so it's I guess OpenAI is the path one, so it's a frontier text LLM with a lot of a frontier text LLM with a lot of a frontier text LLM with a lot of science around post-training, in- science around post-training, in- science around post-training, in- instruct tuning to fine-tune it on instruct tuning to fine-tune it on instruct tuning to fine-tune it on audio, and teaching it to be quite smart audio, and teaching it to be quite smart audio, and teaching it to be quite smart while using audio.

  15. while using audio. while using audio. The nice thing about that is you have a The nice thing about that is you have a The nice thing about that is you have a single model to orchestrate, so it's single model to orchestrate, so it's single model to orchestrate, so it's quite easy to deploy. quite easy to deploy. quite easy to deploy. Um Um Um and and and one one one big aspect, however, is that it's a big aspect, however, is that it's a big aspect, however, is that it's a extremely complex and costly process to extremely complex and costly process to extremely complex and costly process to go from the text model to the go from the text model to the go from the text model to the speech-to-speech model. speech-to-speech model. speech-to-speech model. The second path is to split it. It's an The second path is to split it. It's an The second path is to split it. It's an approach that we introduced in one of approach that we introduced in one of approach that we introduced in one of our recent papers called Moushiraq, and our recent papers called Moushiraq, and our recent papers called Moushiraq, and that has been reused by, in particular, that has been reused by, in particular, that has been reused by, in particular, the Thinking Machine Interaction models. the Thinking Machine Interaction models. the Thinking Machine Interaction models. Where, basically, the idea is that now Where, basically, the idea is that now Where, basically, the idea is that now you have two models. The first one is a you have two models. The first one is a you have two models. The first one is a small, maybe even on-device, small, maybe even on-device, small, maybe even on-device, full-duplex, extremely natural full-duplex, extremely natural full-duplex, extremely natural speech-to-speech interface. And its only speech-to-speech interface. And its only speech-to-speech interface. And its only role role role is to keep a very natural is to keep a very natural is to keep a very natural conversation and be able to delegate conversation and be able to delegate conversation and be able to delegate all the thinking, tool calling, all the thinking, tool calling, all the thinking, tool calling, reasoning, agentic capabilities to a reasoning, agentic capabilities to a reasoning, agentic capabilities to a background text model. And so, the way background text model. And so, the way background text model. And so, the way to see it is you have a background text to see it is you have a background text to see it is you have a background text LLM LLM LLM that receives asynchronously queries that receives asynchronously queries that receives asynchronously queries from hundreds to thousands of small from hundreds to thousands of small from hundreds to thousands of small voice interfaces and just give them voice interfaces and just give them voice interfaces and just give them their text, you know? And, basically, their text, you know? And, basically, their text, you know? And, basically, what we did is um what we did is um what we did is um very small full-duplex model that just very small full-duplex model that just very small full-duplex model that just needs to know when it doesn't know, so needs to know when it doesn't know, so needs to know when it doesn't know, so that it can delegate to the background that it can delegate to the background that it can delegate to the background model.

  16. model. model. And the reason why we believe mostly in And the reason why we believe mostly in And the reason why we believe mostly in this approach, this approach, this approach, um um um and I go back to to it later, it's a and I go back to to it later, it's a and I go back to to it later, it's a our our our uh uh uh let's say our culture is more of first let's say our culture is more of first let's say our culture is more of first one, the bitter lesson. So, every time one, the bitter lesson. So, every time one, the bitter lesson. So, every time we've been pushing for end-to-end we've been pushing for end-to-end we've been pushing for end-to-end systems and so on. But now I think the systems and so on. But now I think the systems and so on. But now I think the hybrid approach has two main hybrid approach has two main hybrid approach has two main advantages. The The first one is cost. advantages. The The first one is cost. advantages. The The first one is cost. So, speech-to-speech models are So, speech-to-speech models are So, speech-to-speech models are notoriously quite expensive. notoriously quite expensive. notoriously quite expensive. And when you think about it, And when you think about it, And when you think about it, it's a loss of money to do chit-chat it's a loss of money to do chit-chat it's a loss of money to do chit-chat with gigantic speech-to-speech models with gigantic speech-to-speech models with gigantic speech-to-speech models that can resolve differential equations that can resolve differential equations that can resolve differential equations and so on. So, it doesn't really and so on. So, it doesn't really and so on. So, it doesn't really make sense economically to get all your make sense economically to get all your make sense economically to get all your workflow through this gigantic workflow through this gigantic workflow through this gigantic multimodal mixture of experts. multimodal mixture of experts. multimodal mixture of experts. At the same time, we see that people are At the same time, we see that people are At the same time, we see that people are very attached to their ability to very attached to their ability to very attached to their ability to control the backend, to be able to control the backend, to be able to control the backend, to be able to switch So they So they 5 was released a switch So they So they 5 was released a switch So they So they 5 was released a few minutes ago. People want to switch few minutes ago. People want to switch few minutes ago. People want to switch the backend and the intelligence and get the backend and the intelligence and get the backend and the intelligence and get a lot of optionality on that, right? a lot of optionality on that, right? a lot of optionality on that, right? When you're using a speech-to-speech When you're using a speech-to-speech When you're using a speech-to-speech model, your your hands are a bit tied model, your your hands are a bit tied model, your your hands are a bit tied with this model provider. And to give with this model provider. And to give with this model provider. And to give you an idea of that, you an idea of that, you an idea of that, until recently the AdSense voice mode until recently the AdSense voice mode until recently the AdSense voice mode from OpenAI was powered by GPT-4o, from OpenAI was powered by GPT-4o, from OpenAI was powered by GPT-4o, despite the fact that there have been despite the fact that there have been despite the fact that there have been several generations of the text model several generations of the text model several generations of the text model since then, because this process is so since then, because this process is so since then, because this process is so expensive and so long.

  17. expensive and so long. expensive and so long. For this reason, we rather bet on the For this reason, we rather bet on the For this reason, we rather bet on the hybrid approach because that will give hybrid approach because that will give hybrid approach because that will give something that is not only very natural something that is not only very natural something that is not only very natural and very nice for demos and impressive, and very nice for demos and impressive, and very nice for demos and impressive, but also will be a viable alternative but also will be a viable alternative but also will be a viable alternative from a economic point of view and from a economic point of view and from a economic point of view and agentic capabilities point of view agentic capabilities point of view agentic capabilities point of view to the best cascaded systems that are to the best cascaded systems that are to the best cascaded systems that are still most of the market today in voice. still most of the market today in voice. still most of the market today in voice. So, So, So, what now? what now? what now? Uh you can use our models on gradium.ai. Uh you can use our models on gradium.ai. Uh you can use our models on gradium.ai. You can apply to gradium. We are You can apply to gradium. We are You can apply to gradium. We are recruiting research scientists and recruiting research scientists and recruiting research scientists and engineers. And thanks for your engineers. And thanks for your engineers. And thanks for your attention.

Summary

This tech talk by the CEO of Gradio explores the evolution and future of voice agents, referencing early examples like Siri and highlighting advancements in speech-to-speech translation and on-device speech processing. The practical takeaway emphasizes the current era of sophisticated voice interfaces powered by foundation models for audio, enabling diverse applications and interactions.

View original episode ↗