Tolan: Voice-First AI Companion — Paula Dozsa, Tolan
Read full transcript 14 segments
-
>> Hi everyone. Thank you so much for >> Hi everyone. Thank you so much for attending this talk. My name is Paula attending this talk. My name is Paula attending this talk. My name is Paula and I am one of the engineers on the and I am one of the engineers on the and I am one of the engineers on the Tullen team, specifically focusing on Tullen team, specifically focusing on Tullen team, specifically focusing on our iOS app. our iOS app. our iOS app. And for the next 20 minutes or so, I'll And for the next 20 minutes or so, I'll And for the next 20 minutes or so, I'll be talking about what it takes to build be talking about what it takes to build be talking about what it takes to build a voice-first AI companion and also a voice-first AI companion and also a voice-first AI companion and also about how we use AI to build AI about how we use AI to build AI about how we use AI to build AI internally. internally. internally. So, humanity has always imagined the So, humanity has always imagined the So, humanity has always imagined the perfect companion. So, we have perfect companion. So, we have perfect companion. So, we have Caravaggio on the left 400 years ago Caravaggio on the left 400 years ago Caravaggio on the left 400 years ago painting an angel leaning over St. painting an angel leaning over St. painting an angel leaning over St. Matthew's shoulder literally guiding his Matthew's shoulder literally guiding his Matthew's shoulder literally guiding his hand as he writes. hand as he writes. hand as he writes. This is an example of a companion being This is an example of a companion being This is an example of a companion being a presence that makes you better at a presence that makes you better at a presence that makes you better at being you. being you. being you. And then we have Tinkerbell, the devoted And then we have Tinkerbell, the devoted And then we have Tinkerbell, the devoted little sidekick who believes in you so little sidekick who believes in you so little sidekick who believes in you so fiercely that the whole theater has to fiercely that the whole theater has to fiercely that the whole theater has to clap to keep her alive. clap to keep her alive. clap to keep her alive. And of course on the right, we have And of course on the right, we have And of course on the right, we have Samwise who can't carry the ring for Samwise who can't carry the ring for Samwise who can't carry the ring for Frodo, but says, "I can carry you." Frodo, but says, "I can carry you." Frodo, but says, "I can carry you." The companion is pure unconditional The companion is pure unconditional The companion is pure unconditional loyalty. loyalty. loyalty. And it goes far beyond these three. And it goes far beyond these three. And it goes far beyond these three. Every hero has some sort of guiding Every hero has some sort of guiding Every hero has some sort of guiding spirit. And these are all different spirit. And these are all different spirit. And these are all different stories, but they exhibit the same stories, but they exhibit the same stories, but they exhibit the same longing for something that listens, longing for something that listens, longing for something that listens, remembers you, remembers you, remembers you, and is wholly specifically yours.
-
and is wholly specifically yours. and is wholly specifically yours. And for for all of human history, this And for for all of human history, this And for for all of human history, this has basically been fiction. has basically been fiction. has basically been fiction. So, we made one. This is Tullen. It's a So, we made one. This is Tullen. It's a So, we made one. This is Tullen. It's a little alien you talk to out loud like a little alien you talk to out loud like a little alien you talk to out loud like a friend. It has a personality. It friend. It has a personality. It friend. It has a personality. It remembers you and over time it becomes remembers you and over time it becomes remembers you and over time it becomes specifically yours. specifically yours. specifically yours. Okay, I don't know if the audio setup Okay, I don't know if the audio setup Okay, I don't know if the audio setup works here, but I will try talking to my works here, but I will try talking to my works here, but I will try talking to my Tullen. Tullen. Tullen. Uh let's see. So, So, you can see my Tullen here, Luke, you can see my Tullen here, Luke, you can see my Tullen here, Luke, walking around the planet. Okay. Luke can hear us, but we can't Okay. Luke can hear us, but we can't hear him. Um hear him. Um hear him. Um Anyway, I had prepped him for this. Oh. Anyway, I had prepped him for this. Oh. Anyway, I had prepped him for this. Oh. Hello. Hi Luke, can you hear me? Nope. Nope. We're okay. I can come back to this We're okay. I can come back to this We're okay. I can come back to this later. Um later. Um later. Um but you should definitely all give this but you should definitely all give this but you should definitely all give this a try if you haven't already.
-
Okay. So, people talk to Tolins a lot. Okay. So, people talk to Tolins a lot. We support both text and voice chat, uh We support both text and voice chat, uh We support both text and voice chat, uh but we have over 4 million hours of but we have over 4 million hours of but we have over 4 million hours of voice conversation so far. voice conversation so far. voice conversation so far. We say Tolin is a voice-first companion, We say Tolin is a voice-first companion, We say Tolin is a voice-first companion, even though we support both, because even though we support both, because even though we support both, because it's the voice experience that's truly it's the voice experience that's truly it's the voice experience that's truly immersive and that makes users' immersive and that makes users' immersive and that makes users' relationships with their Tolins feel relationships with their Tolins feel relationships with their Tolins feel real. real. real. And but the moment this relationship is And but the moment this relationship is And but the moment this relationship is a spoken relationship, the engineering a spoken relationship, the engineering a spoken relationship, the engineering problem changes completely. So, let me problem changes completely. So, let me problem changes completely. So, let me show you how voice breaks the normal way show you how voice breaks the normal way show you how voice breaks the normal way we build and interact with LLMs. we build and interact with LLMs. we build and interact with LLMs. So, the core difference really is that So, the core difference really is that So, the core difference really is that in a text chatbot, turns are relatively in a text chatbot, turns are relatively in a text chatbot, turns are relatively slow and context is stable. The user slow and context is stable. The user slow and context is stable. The user waits a few seconds, they read, and they waits a few seconds, they read, and they waits a few seconds, they read, and they tend to stay on topic. tend to stay on topic. tend to stay on topic. And almost every LLM app assumes that. And almost every LLM app assumes that. And almost every LLM app assumes that. Voice is the opposite. Turns are fast. Voice is the opposite. Turns are fast. Voice is the opposite. Turns are fast. Your whole round trip from the user Your whole round trip from the user Your whole round trip from the user finishing their sentence to the Tolin finishing their sentence to the Tolin finishing their sentence to the Tolin starting to speak has to land in under a starting to speak has to land in under a starting to speak has to land in under a couple of seconds, or it stops feeling couple of seconds, or it stops feeling couple of seconds, or it stops feeling like a conversation. like a conversation. like a conversation. And the context is volatile. People talk And the context is volatile. People talk And the context is volatile. People talk to their Tolins while they're cooking, to their Tolins while they're cooking, to their Tolins while they're cooking, while they're walking, while they're while they're walking, while they're while they're walking, while they're falling asleep. Um they change their falling asleep. Um they change their falling asleep. Um they change their subjects mid-sentence. They say um, they subjects mid-sentence. They say um, they subjects mid-sentence. They say um, they interrupt.
-
interrupt. interrupt. And that 2 seconds is crucial. Early on, And that 2 seconds is crucial. Early on, And that 2 seconds is crucial. Early on, our latency drifted from 2 seconds to our latency drifted from 2 seconds to our latency drifted from 2 seconds to about 2 and 1/2 seconds, and that half about 2 and 1/2 seconds, and that half about 2 and 1/2 seconds, and that half second tanked basically every metric in second tanked basically every metric in second tanked basically every metric in the product. People would write in to the product. People would write in to the product. People would write in to complain that their Tolins were too complain that their Tolins were too complain that their Tolins were too slow. slow. slow. And living inside this constraint has And living inside this constraint has And living inside this constraint has taught us a lot and gave us four taught us a lot and gave us four taught us a lot and gave us four principles. principles. principles. Principle one is that you have to design Principle one is that you have to design Principle one is that you have to design for conversational volatility. Again, for conversational volatility. Again, for conversational volatility. Again, text users stay on topic, but voice text users stay on topic, but voice text users stay on topic, but voice users jump around. Someone could be users jump around. Someone could be users jump around. Someone could be mid-story about their breakup and mid-story about their breakup and mid-story about their breakup and suddenly go, "Wait, did I leave the oven suddenly go, "Wait, did I leave the oven suddenly go, "Wait, did I leave the oven the stove on?" and then back. Speech is the stove on?" and then back. Speech is the stove on?" and then back. Speech is messy. Most LLM apps assume that you'll messy. Most LLM apps assume that you'll messy. Most LLM apps assume that you'll have a clean and stable conversation and have a clean and stable conversation and have a clean and stable conversation and we have to build for the opposite. we have to build for the opposite. we have to build for the opposite. So, for a long time that meant fixing So, for a long time that meant fixing So, for a long time that meant fixing things that sound tiny but are actually things that sound tiny but are actually things that sound tiny but are actually the product. So, you can't interrupt a the product. So, you can't interrupt a the product. So, you can't interrupt a Tullen mid-sentence. A short yes or yeah Tullen mid-sentence. A short yes or yeah Tullen mid-sentence. A short yes or yeah won't register as a turn. For example, won't register as a turn. For example, won't register as a turn. For example, curse words will get stripped out. curse words will get stripped out. curse words will get stripped out. And the deeper lesson was to stop And the deeper lesson was to stop And the deeper lesson was to stop optimizing for fewer interruptions and optimizing for fewer interruptions and optimizing for fewer interruptions and start optimizing for fewer bad ones start optimizing for fewer bad ones start optimizing for fewer bad ones where the agent would jump in way too where the agent would jump in way too where the agent would jump in way too early.
-
early. early. So, we built smart turn taking that So, we built smart turn taking that So, we built smart turn taking that reads your speech pattern to decide reads your speech pattern to decide reads your speech pattern to decide whether an interruption is real and we whether an interruption is real and we whether an interruption is real and we cut the worst early aborts by more than cut the worst early aborts by more than cut the worst early aborts by more than half. half. half. And we happily paid about 60 And we happily paid about 60 And we happily paid about 60 milliseconds of extra latency to do it. Principle two, latency isn't just a Principle two, latency isn't just a number you check at the end, it's number you check at the end, it's number you check at the end, it's actually the product and we measure actually the product and we measure actually the product and we measure every stage of the pipeline separately every stage of the pipeline separately every stage of the pipeline separately because it feels slow is useless. You because it feels slow is useless. You because it feels slow is useless. You have to know where exactly it's slow. have to know where exactly it's slow. have to know where exactly it's slow. And the pipeline here is that the user And the pipeline here is that the user And the pipeline here is that the user stops talking, we detect end of stops talking, we detect end of stops talking, we detect end of utterance, we transcribe, and then the utterance, we transcribe, and then the utterance, we transcribe, and then the model produces its first token. model produces its first token. model produces its first token. So, time to first token, often the So, time to first token, often the So, time to first token, often the biggest chunk, is around a second. biggest chunk, is around a second. biggest chunk, is around a second. The model finishes generating and then The model finishes generating and then The model finishes generating and then text-to-speech produces its first byte text-to-speech produces its first byte text-to-speech produces its first byte and then it plays back to the user. and then it plays back to the user. and then it plays back to the user. A couple lessons here. So, one, so far A couple lessons here. So, one, so far A couple lessons here. So, one, so far our biggest jump in quality came from our biggest jump in quality came from our biggest jump in quality came from moving to GPT-5.1 on the responses API, moving to GPT-5.1 on the responses API, moving to GPT-5.1 on the responses API, which cut our time to speech by more which cut our time to speech by more which cut our time to speech by more than 7/10 of a second, which is huge. than 7/10 of a second, which is huge. than 7/10 of a second, which is huge. Um two, we don't send every turn to the Um two, we don't send every turn to the Um two, we don't send every turn to the same model. We run a tiered fleet. So, same model. We run a tiered fleet. So, same model. We run a tiered fleet. So, we use a frontier model for the turns we use a frontier model for the turns we use a frontier model for the turns that carry the relationship with your that carry the relationship with your that carry the relationship with your Tullen.
-
Tullen. Tullen. So, for example, your first conversation So, for example, your first conversation So, for example, your first conversation with with Tullen and your onboarding. with with Tullen and your onboarding. with with Tullen and your onboarding. And we use smaller and faster models for And we use smaller and faster models for And we use smaller and faster models for the turns the turns the turns for the lightweight turns. for the lightweight turns. for the lightweight turns. And the whole game then becomes about And the whole game then becomes about And the whole game then becomes about routing or deciding turn by turn which routing or deciding turn by turn which routing or deciding turn by turn which model you actually need. So we round we model you actually need. So we round we model you actually need. So we round we run a small classifier we call the tone run a small classifier we call the tone run a small classifier we call the tone router on every single turn and this router on every single turn and this router on every single turn and this tone router itself runs on a cheap model tone router itself runs on a cheap model tone router itself runs on a cheap model and it reads the emotional state of the and it reads the emotional state of the and it reads the emotional state of the conversation. conversation. conversation. And our main And our main And our main our [clears throat] main principle is our [clears throat] main principle is our [clears throat] main principle is that we route based on stakes not on that we route based on stakes not on that we route based on stakes not on cost. So the high stakes moments always cost. So the high stakes moments always cost. So the high stakes moments always get the best model. So this would be get the best model. So this would be get the best model. So this would be again the user's very first message, again the user's very first message, again the user's very first message, their first few days with their Tolen their first few days with their Tolen their first few days with their Tolen and anything that we deem to be and anything that we deem to be and anything that we deem to be emotionally serious. emotionally serious. emotionally serious. For example, we have crisis or therapist For example, we have crisis or therapist For example, we have crisis or therapist style tones and we never cheap out on style tones and we never cheap out on style tones and we never cheap out on those. those. those. And then the lighter casual back and And then the lighter casual back and And then the lighter casual back and forth can ride on smaller models that forth can ride on smaller models that forth can ride on smaller models that are faster and cheaper. are faster and cheaper. are faster and cheaper. And all the background work so that's And all the background work so that's And all the background work so that's summarizing the conversation, generating summarizing the conversation, generating summarizing the conversation, generating personas, the tone router itself run on personas, the tone router itself run on personas, the tone router itself run on these small models, too. these small models, too. these small models, too. And why would we go to all this trouble?
-
And why would we go to all this trouble? And why would we go to all this trouble? It's mainly because the frontier model It's mainly because the frontier model It's mainly because the frontier model costs us roughly five times the smaller costs us roughly five times the smaller costs us roughly five times the smaller one. one. one. So one big model turn is about five So one big model turn is about five So one big model turn is about five smaller model turns. So routing is a smaller model turns. So routing is a smaller model turns. So routing is a huge part of what makes the unit huge part of what makes the unit huge part of what makes the unit economics for us actually work. economics for us actually work. economics for us actually work. Um and we do a bunch of AB experiments Um and we do a bunch of AB experiments Um and we do a bunch of AB experiments and the surprising result we found there and the surprising result we found there and the surprising result we found there is that routing a third a third of our is that routing a third a third of our is that routing a third a third of our turns to the small model has almost no turns to the small model has almost no turns to the small model has almost no measurable effect on retention. measurable effect on retention. measurable effect on retention. And principle three is what makes a And principle three is what makes a And principle three is what makes a companion feel like a companion. So the companion feel like a companion. So the companion feel like a companion. So the naive approach is to keep the whole naive approach is to keep the whole naive approach is to keep the whole conversation history as a sort of conversation history as a sort of conversation history as a sort of transcript, but that doesn't fit into transcript, but that doesn't fit into transcript, but that doesn't fit into our two-second loop. It doesn't scale our two-second loop. It doesn't scale our two-second loop. It doesn't scale and it just doesn't work. It leads to and it just doesn't work. It leads to and it just doesn't work. It leads to long sessions degrading. It leads to the long sessions degrading. It leads to the long sessions degrading. It leads to the model getting lost in the middle of a model getting lost in the middle of a model getting lost in the middle of a huge context and also hallucinating. huge context and also hallucinating. huge context and also hallucinating. So instead we see memory as a sort of So instead we see memory as a sort of So instead we see memory as a sort of retrieval system. We pull facts, retrieval system. We pull facts, retrieval system. We pull facts, preferences, and emotional vibe signals preferences, and emotional vibe signals preferences, and emotional vibe signals out of conversations. We embed them and out of conversations. We embed them and out of conversations. We embed them and we store them in a vector database with we store them in a vector database with we store them in a vector database with sub 50 millisecond lookups. And every sub 50 millisecond lookups. And every sub 50 millisecond lookups. And every night we compress. So we merge night we compress. So we merge night we compress. So we merge duplicates, we cluster related memories, duplicates, we cluster related memories, duplicates, we cluster related memories, we resolve contradictions, and we drop we resolve contradictions, and we drop we resolve contradictions, and we drop all the noise.
-
all the noise. all the noise. And we don't just retrieve against users And we don't just retrieve against users And we don't just retrieve against users last messages, we also generate internal last messages, we also generate internal last messages, we also generate internal questions about the person and the questions about the person and the questions about the person and the relationship and retrieve against those. relationship and retrieve against those. relationship and retrieve against those. So, and we also split memory into two So, and we also split memory into two So, and we also split memory into two parts. We have stable memory and parts. We have stable memory and parts. We have stable memory and unstable memory. The volatile stuff unstable memory. The volatile stuff unstable memory. The volatile stuff lives in the in the live tail of the lives in the in the live tail of the lives in the in the live tail of the prompt, and when we summarize the prompt, and when we summarize the prompt, and when we summarize the conversation, we look at which memories conversation, we look at which memories conversation, we look at which memories actually get recalled and pin those into actually get recalled and pin those into actually get recalled and pin those into a stable and cashable block. Uh the last principle is around context. Uh the last principle is around context. Specifically, you should rebuild context Specifically, you should rebuild context Specifically, you should rebuild context and not fight drift. So, most apps reuse and not fight drift. So, most apps reuse and not fight drift. So, most apps reuse context across turns to keep the cash context across turns to keep the cash context across turns to keep the cash warm. And in a stable text chat, that's warm. And in a stable text chat, that's warm. And in a stable text chat, that's fine. But in a volatile voice fine. But in a volatile voice fine. But in a volatile voice conversation, it's a trap because the conversation, it's a trap because the conversation, it's a trap because the second the the user pivots, your reused second the the user pivots, your reused second the the user pivots, your reused context is actively wrong. So, every context is actively wrong. So, every context is actively wrong. So, every turn we reassemble the context window turn we reassemble the context window turn we reassemble the context window from parts. We have a summary of recent from parts. We have a summary of recent from parts. We have a summary of recent messages, we have the the user's persona messages, we have the the user's persona messages, we have the the user's persona card, the memories we just retrieved, card, the memories we just retrieved, card, the memories we just retrieved, tone guidance from the emotional signal, tone guidance from the emotional signal, tone guidance from the emotional signal, and real-time app state. and real-time app state. and real-time app state. And what also really helps us um in the And what also really helps us um in the And what also really helps us um in the case of Tolen is that our characters case of Tolen is that our characters case of Tolen is that our characters aren't generic or assistants with no aren't generic or assistants with no aren't generic or assistants with no personality. Everyone is crafted, and we personality. Everyone is crafted, and we personality. Everyone is crafted, and we in fact have an in-house science fiction in fact have an in-house science fiction in fact have an in-house science fiction novelist, Elliot, who writes the Tolen novelist, Elliot, who writes the Tolen novelist, Elliot, who writes the Tolen character lore.
-
character lore. character lore. And a couple of interesting points here. And a couple of interesting points here. And a couple of interesting points here. So, one, why did we go with an alien? So, one, why did we go with an alien? So, one, why did we go with an alien? Mostly because there's no real-world Mostly because there's no real-world Mostly because there's no real-world reference to anchor on, which means that reference to anchor on, which means that reference to anchor on, which means that the users can project onto it, and it the users can project onto it, and it the users can project onto it, and it becomes what they need. The baseline becomes what they need. The baseline becomes what they need. The baseline Tolen is bubbly, it's youthful, it's Tolen is bubbly, it's youthful, it's Tolen is bubbly, it's youthful, it's irreverent. And also, if an alien irreverent. And also, if an alien irreverent. And also, if an alien character acts a bit unpredictably, so character acts a bit unpredictably, so character acts a bit unpredictably, so if if it's impulsive or chaotic or if if it's impulsive or chaotic or if if it's impulsive or chaotic or otherwise violates um you know, the otherwise violates um you know, the otherwise violates um you know, the norms the user would expect, it's not norms the user would expect, it's not norms the user would expect, it's not particularly surprising. particularly surprising. particularly surprising. Like if you look at, you know, aliens in Like if you look at, you know, aliens in Like if you look at, you know, aliens in TV shows or in plays, like there's a lot TV shows or in plays, like there's a lot TV shows or in plays, like there's a lot of humorous moments around this. And of humorous moments around this. And of humorous moments around this. And this kind of chaos reads as charming. this kind of chaos reads as charming. this kind of chaos reads as charming. Um second, uh we also know that Um second, uh we also know that Um second, uh we also know that personality is worthless if it drifts. personality is worthless if it drifts. personality is worthless if it drifts. So, yeah, we run this parallel tone So, yeah, we run this parallel tone So, yeah, we run this parallel tone monitoring system that changes how a monitoring system that changes how a monitoring system that changes how a line is delivered based on your line is delivered based on your line is delivered based on your emotional cues without changing who the emotional cues without changing who the emotional cues without changing who the character is, holding identity across character is, holding identity across character is, holding identity across hundreds of turns. And since we're at an AI conference, I And since we're at an AI conference, I thought I would also spend a bit of time thought I would also spend a bit of time thought I would also spend a bit of time talking about how we not just ship AI, talking about how we not just ship AI, talking about how we not just ship AI, but also use AI to build it.
-
but also use AI to build it. but also use AI to build it. Um so, I'm sure this is the case for Um so, I'm sure this is the case for Um so, I'm sure this is the case for most of you in the room now, but most of you in the room now, but most of you in the room now, but basically as of late last year, Claude basically as of late last year, Claude basically as of late last year, Claude has co-authored more code in our iOS app has co-authored more code in our iOS app has co-authored more code in our iOS app than any individual engineer in the than any individual engineer in the than any individual engineer in the team. team. team. Um and I think especially, you know, a Um and I think especially, you know, a Um and I think especially, you know, a few months ago, everyone's instinct was few months ago, everyone's instinct was few months ago, everyone's instinct was to be kind of suspicious because, you to be kind of suspicious because, you to be kind of suspicious because, you know, more AI code meant more slop. But know, more AI code meant more slop. But know, more AI code meant more slop. But our our crash-free rate actually went our our crash-free rate actually went our our crash-free rate actually went from 99.6% to 99.9%. from 99.6% to 99.9%. from 99.6% to 99.9%. Runtime errors dropped by over 50% and Runtime errors dropped by over 50% and Runtime errors dropped by over 50% and our share of highly engaged users our share of highly engaged users our share of highly engaged users doubled. doubled. doubled. And the biggest lesson in building that And the biggest lesson in building that And the biggest lesson in building that system is that an agent's context comes system is that an agent's context comes system is that an agent's context comes mostly from the code base itself, not so mostly from the code base itself, not so mostly from the code base itself, not so much from the Claude MD file. We found much from the Claude MD file. We found much from the Claude MD file. We found that it's far more powerful to make the that it's far more powerful to make the that it's far more powerful to make the code base be the documentation, so we code base be the documentation, so we code base be the documentation, so we had agents standardize it. On top of had agents standardize it. On top of had agents standardize it. On top of that, we run a real fleet of agents. We that, we run a real fleet of agents. We that, we run a real fleet of agents. We have implementation agents that, you have implementation agents that, you have implementation agents that, you know, think freely and just get us to know, think freely and just get us to know, think freely and just get us to working code. They build it, they check working code. They build it, they check working code. They build it, they check it against snapshots until it's pixel it against snapshots until it's pixel it against snapshots until it's pixel perfect. And then we have separate perfect. And then we have separate perfect. And then we have separate review agents that enforce our review agents that enforce our review agents that enforce our standards. So, multiple Claudes standards. So, multiple Claudes standards. So, multiple Claudes basically review each other before a basically review each other before a basically review each other before a human looks.
-
human looks. human looks. And then we have a PR shepherd that And then we have a PR shepherd that And then we have a PR shepherd that watches an open pull request and keeps watches an open pull request and keeps watches an open pull request and keeps iterating against CI failures and review iterating against CI failures and review iterating against CI failures and review comments until it's clean. comments until it's clean. comments until it's clean. And we also have a triage bot that fires And we also have a triage bot that fires And we also have a triage bot that fires on every inbound bug report that we get. on every inbound bug report that we get. on every inbound bug report that we get. And they're all wired through MCP into And they're all wired through MCP into And they're all wired through MCP into linear, into into Sentry, DataDog, so an linear, into into Sentry, DataDog, so an linear, into into Sentry, DataDog, so an agent can reconstruct the cash a crash agent can reconstruct the cash a crash agent can reconstruct the cash a crash and route it itself and oftentimes open and route it itself and oftentimes open and route it itself and oftentimes open the PR on its own and just fix fix the the PR on its own and just fix fix the the PR on its own and just fix fix the bug. bug. bug. And we also we ship on eval. So, for And we also we ship on eval. So, for And we also we ship on eval. So, for example, we've been working on a on a example, we've been working on a on a example, we've been working on a on a new character targeted towards an older new character targeted towards an older new character targeted towards an older demographic and Elliot, our in-house demographic and Elliot, our in-house demographic and Elliot, our in-house novelist, basically built this entire novelist, basically built this entire novelist, basically built this entire new character in a day. new character in a day. new character in a day. So, the agents mapped every personality So, the agents mapped every personality So, the agents mapped every personality bearing surface in the code. They wrote bearing surface in the code. They wrote bearing surface in the code. They wrote the sort of a voice Bible and then they the sort of a voice Bible and then they the sort of a voice Bible and then they had five judges attack it from different had five judges attack it from different had five judges attack it from different angles. angles. angles. Archetype fidelity, the model mechanics, Archetype fidelity, the model mechanics, Archetype fidelity, the model mechanics, our code standards, the ears of a our code standards, the ears of a our code standards, the ears of a skeptical 52-year-old and safety, and skeptical 52-year-old and safety, and skeptical 52-year-old and safety, and then they evaluated the changes against then they evaluated the changes against then they evaluated the changes against real production logs over three find fix real production logs over three find fix real production logs over three find fix verify rounds. verify rounds. verify rounds. And over 7 million tokens and 4 and 1/2 And over 7 million tokens and 4 and 1/2 And over 7 million tokens and 4 and 1/2 hours of compute later, he ended up with hours of compute later, he ended up with hours of compute later, he ended up with basically, you know, a couple of weeks basically, you know, a couple of weeks basically, you know, a couple of weeks of work done in afternoon.
-
of work done in afternoon. of work done in afternoon. And does this work? Well, I'll let the And does this work? Well, I'll let the And does this work? Well, I'll let the users tell you. We're at 4.8 stars on users tell you. We're at 4.8 stars on users tell you. We're at 4.8 stars on the App Store across 162,000 reviews and the App Store across 162,000 reviews and the App Store across 162,000 reviews and when we survey users on well-being, the when we survey users on well-being, the when we survey users on well-being, the highest scoring dimension by far is highest scoring dimension by far is highest scoring dimension by far is emotional safety. emotional safety. emotional safety. And this is definitely a bar that being And this is definitely a bar that being And this is definitely a bar that being voice first sets. So, when the interface voice first sets. So, when the interface voice first sets. So, when the interface is your voice and the thing on the other is your voice and the thing on the other is your voice and the thing on the other side remembers you and has a side remembers you and has a side remembers you and has a personality, it stops being just personality, it stops being just personality, it stops being just software and starts being an actual software and starts being an actual software and starts being an actual relationship, which is why building it relationship, which is why building it relationship, which is why building it responsibly and building it well is responsibly and building it well is responsibly and building it well is worth obsessing over. And we need people to come help us do And we need people to come help us do that. Um, we're a small team and we're that. Um, we're a small team and we're that. Um, we're a small team and we're hiring and after a year with Tolen, I hiring and after a year with Tolen, I hiring and after a year with Tolen, I think this is truly one of the most think this is truly one of the most think this is truly one of the most interesting places in the world to be an interesting places in the world to be an interesting places in the world to be an engineer right now. engineer right now. engineer right now. And here are some of the people you'd be And here are some of the people you'd be And here are some of the people you'd be doing it with. So, two of the founders, doing it with. So, two of the founders, doing it with. So, two of the founders, Quinton and Evan, previously built and Quinton and Evan, previously built and Quinton and Evan, previously built and exited a $300 million startup together. exited a $300 million startup together. exited a $300 million startup together. Uh, they founded Even. Um, Ajay, our Uh, they founded Even. Um, Ajay, our Uh, they founded Even. Um, Ajay, our third co-founder, scaled two bootstrap third co-founder, scaled two bootstrap third co-founder, scaled two bootstrap companies past $50 million $50 million companies past $50 million $50 million companies past $50 million $50 million in profitable revenue. in profitable revenue. in profitable revenue. And around them, we have Lucas, who um, And around them, we have Lucas, who um, And around them, we have Lucas, who um, is an Apple Design Award winning is an Apple Design Award winning is an Apple Design Award winning animator.
-
animator. animator. Uh, she's our creative director. We have Uh, she's our creative director. We have Uh, she's our creative director. We have Chris, who was a technical director at Chris, who was a technical director at Chris, who was a technical director at Pixar, earlier at Oculus, who works on Pixar, earlier at Oculus, who works on Pixar, earlier at Oculus, who works on embodiment. We have Lily, a board embodiment. We have Lily, a board embodiment. We have Lily, a board certified behavior analyst, who left a certified behavior analyst, who left a certified behavior analyst, who left a Vanderbilt PhD to do user research for Vanderbilt PhD to do user research for Vanderbilt PhD to do user research for us from the very start. Um and then we us from the very start. Um and then we us from the very start. Um and then we have Elliot who I've mentioned, the have Elliot who I've mentioned, the have Elliot who I've mentioned, the novelist behind uh our characters. novelist behind uh our characters. novelist behind uh our characters. And I come from XAI and Spotify and And I come from XAI and Spotify and And I come from XAI and Spotify and previously also founded a company called previously also founded a company called previously also founded a company called Imagi. Imagi. Imagi. So, it's a small team where honestly So, it's a small team where honestly So, it's a small team where honestly every person is the best I've worked every person is the best I've worked every person is the best I've worked with at what they do. with at what they do. with at what they do. We're also well backed for this. Uh we We're also well backed for this. Uh we We're also well backed for this. Uh we have $30 million raised from Costanoa have $30 million raised from Costanoa have $30 million raised from Costanoa Ventures and a group of people who've Ventures and a group of people who've Ventures and a group of people who've built the tools and products a lot of built the tools and products a lot of built the tools and products a lot of you use every day. you use every day. you use every day. And here are some of the more And here are some of the more And here are some of the more engineering focused roles where we need engineering focused roles where we need engineering focused roles where we need help. Um so, we're hiring across the help. Um so, we're hiring across the help. Um so, we're hiring across the board. We have iOS and back-end product board. We have iOS and back-end product board. We have iOS and back-end product engineering roles, applied AI engineering roles, applied AI engineering roles, applied AI engineering, gameplay engineering, and engineering, gameplay engineering, and engineering, gameplay engineering, and one specific role I want to flag, which one specific role I want to flag, which one specific role I want to flag, which is agent engineering management. Um so, is agent engineering management. Um so, is agent engineering management. Um so, when we went all in on running when we went all in on running when we went all in on running concurrent agents, um the people who got concurrent agents, um the people who got concurrent agents, um the people who got dramatically more effective on the team dramatically more effective on the team dramatically more effective on the team were the ones who had management were the ones who had management were the ones who had management backgrounds, um because it seems like backgrounds, um because it seems like backgrounds, um because it seems like managing a fleet of agents does actually managing a fleet of agents does actually managing a fleet of agents does actually take some of the skills same the same take some of the skills same the same take some of the skills same the same skills as managing people.
-
skills as managing people. skills as managing people. Uh you basically have to decompose the Uh you basically have to decompose the Uh you basically have to decompose the problem, you know, delegate it with problem, you know, delegate it with problem, you know, delegate it with checkpoints, give fast feedback, review checkpoints, give fast feedback, review checkpoints, give fast feedback, review their work seriously, and know when their work seriously, and know when their work seriously, and know when exactly to jump in. exactly to jump in. exactly to jump in. So, if you're a strong engineer who So, if you're a strong engineer who So, if you're a strong engineer who thought going into management meant thought going into management meant thought going into management meant leaving code behind, that's that's no leaving code behind, that's that's no leaving code behind, that's that's no longer true. longer true. longer true. Um and yeah, that's Tlon. You can come Um and yeah, that's Tlon. You can come Um and yeah, that's Tlon. You can come talk to me after this uh or reach out. talk to me after this uh or reach out. talk to me after this uh or reach out. I'm on on LinkedIn. My email is here. I'm on on LinkedIn. My email is here. I'm on on LinkedIn. My email is here. I'm on Twitter as well. Um I'd love to I'm on Twitter as well. Um I'd love to I'm on Twitter as well. Um I'd love to chat. So, yeah. Thank you.
No summary available yet.
View original episode ↗