Shipping AI to a Million Patients Without an A/B Test — Jared Joselowitz, Ufonia
Read full transcript 18 segments
-
>> Uh hello everyone. It's really nice to >> Uh hello everyone. It's really nice to see you all. see you all. see you all. Um my name is Jared and I'm going to Um my name is Jared and I'm going to Um my name is Jared and I'm going to share some of the work that we do on share some of the work that we do on share some of the work that we do on shipping healthcare AI safely basically. shipping healthcare AI safely basically. shipping healthcare AI safely basically. Um so just a little bit about me. Um I Um so just a little bit about me. Um I Um so just a little bit about me. Um I come from South Africa where I actually come from South Africa where I actually come from South Africa where I actually studied electrical engineering studied electrical engineering studied electrical engineering um before making the very unique um before making the very unique um before making the very unique decision to transition to AI a few years decision to transition to AI a few years decision to transition to AI a few years ago. ago. ago. Um I now work as a research engineer for Um I now work as a research engineer for Um I now work as a research engineer for Euphony which is basically a a Euphony which is basically a a Euphony which is basically a a healthcare company based in the UK. healthcare company based in the UK. healthcare company based in the UK. And um the work we do I work within the And um the work we do I work within the And um the work we do I work within the science team is we built the safety and science team is we built the safety and science team is we built the safety and evaluation stack behind Dora which is a evaluation stack behind Dora which is a evaluation stack behind Dora which is a clinical conversational agent. clinical conversational agent. clinical conversational agent. And my job and our job within the And my job and our job within the And my job and our job within the science team is proving that the product science team is proving that the product science team is proving that the product is safe before a patient ever actually is safe before a patient ever actually is safe before a patient ever actually hears it. hears it. hears it. So shipping to patients takes away the So shipping to patients takes away the So shipping to patients takes away the normal safety nets you would normally normal safety nets you would normally normal safety nets you would normally ship with. Um three of them could be ship with. Um three of them could be ship with. Um three of them could be that you can't actually AB test on that you can't actually AB test on that you can't actually AB test on patients of course. Randomizing patients patients of course. Randomizing patients patients of course. Randomizing patients into a worse variant is unethical and into a worse variant is unethical and into a worse variant is unethical and often illegal. often illegal. often illegal. Um you can't undo a call. Once Dora says Um you can't undo a call. Once Dora says Um you can't undo a call. Once Dora says it, it's been said and there is no it, it's been said and there is no it, it's been said and there is no rollback.
-
rollback. rollback. And very importantly the model card And very importantly the model card And very importantly the model card won't save you. Um you can't claim like won't save you. Um you can't claim like won't save you. Um you can't claim like some model vendors said that they have some model vendors said that they have some model vendors said that they have 92% on some benchmark. Um it's not a 92% on some benchmark. Um it's not a 92% on some benchmark. Um it's not a defense at a post instant review. So defense at a post instant review. So defense at a post instant review. So everything that we do at Euphony comes everything that we do at Euphony comes everything that we do at Euphony comes from living within constraints. So what is Dora? Um a little bit about So what is Dora? Um a little bit about Dora. Dora is a a voice AI agent that Dora. Dora is a a voice AI agent that Dora. Dora is a a voice AI agent that phones patients and has real clinical phones patients and has real clinical phones patients and has real clinical conversations like post-op follow-ups or conversations like post-op follow-ups or conversations like post-op follow-ups or pre-op checks. Um these are calls that a pre-op checks. Um these are calls that a pre-op checks. Um these are calls that a clinician would usually have to make clinician would usually have to make clinician would usually have to make themselves but Dora doesn't So Dora themselves but Dora doesn't So Dora themselves but Dora doesn't So Dora obviously doesn't replace doctors but it obviously doesn't replace doctors but it obviously doesn't replace doctors but it just takes a hugely time-consuming task just takes a hugely time-consuming task just takes a hugely time-consuming task off their plate and gives them their off their plate and gives them their off their plate and gives them their time back. time back. time back. Um so far we've done around 200,000 real Um so far we've done around 200,000 real Um so far we've done around 200,000 real clinical calls within the UK across 20 clinical calls within the UK across 20 clinical calls within the UK across 20 hospitals. hospitals. hospitals. And we are contracted to scale to a And we are contracted to scale to a And we are contracted to scale to a million patients in the next 2 years. million patients in the next 2 years. million patients in the next 2 years. Um, we also um, released our product in Um, we also um, released our product in Um, we also um, released our product in the US this year and we are live in in the US this year and we are live in in the US this year and we are live in in in two clinics and we've signed up for in two clinics and we've signed up for in two clinics and we've signed up for six more clinics across four different six more clinics across four different six more clinics across four different states. So, I'm just going to for for a bit of So, I'm just going to for for a bit of better context, I just wanted to do a better context, I just wanted to do a better context, I just wanted to do a quick demo of what the experience with quick demo of what the experience with quick demo of what the experience with Dora is actually like.
-
Dora is actually like. Dora is actually like. >> That's not uncommon after cataract >> That's not uncommon after cataract >> That's not uncommon after cataract surgery. Are you having trouble seeing surgery. Are you having trouble seeing surgery. Are you having trouble seeing clearly at a distance, up close, or clearly at a distance, up close, or clearly at a distance, up close, or both? both? both? >> Um, at a distance is fine. It's just uh, >> Um, at a distance is fine. It's just uh, >> Um, at a distance is fine. It's just uh, up close that's um, it's a bit it's I'm up close that's um, it's a bit it's I'm up close that's um, it's a bit it's I'm having a bit of trouble with that. having a bit of trouble with that. having a bit of trouble with that. >> So, when exactly did you first notice it >> So, when exactly did you first notice it >> So, when exactly did you first notice it was a bit blurry when trying to read or was a bit blurry when trying to read or was a bit blurry when trying to read or see things up close? see things up close? see things up close? >> Um, >> Um, >> Um, well, really just after the surgery. well, really just after the surgery. well, really just after the surgery. And I was wondering how long it would And I was wondering how long it would And I was wondering how long it would last. last. last. >> So, you're asking about the duration. >> So, you're asking about the duration. >> So, you're asking about the duration. That's understandable. Blurry vision That's understandable. Blurry vision That's understandable. Blurry vision should improve rapidly in the first few should improve rapidly in the first few should improve rapidly in the first few days after surgery. days after surgery. days after surgery. >> Okay. Um, also, um, is it all right if I >> Okay. Um, also, um, is it all right if I >> Okay. Um, also, um, is it all right if I go swimming again? go swimming again? go swimming again? >> You should avoid swimming for a month >> You should avoid swimming for a month >> You should avoid swimming for a month after surgery. That's just to be on the after surgery. That's just to be on the after surgery. That's just to be on the safe side and help your eye heal safe side and help your eye heal safe side and help your eye heal properly. properly. properly. >> All right. Okay. >> All right. Okay. >> All right. Okay. >> Awesome.
-
So, as you just heard, sometimes Dora So, as you just heard, sometimes Dora asks about symptoms, it gives advice, asks about symptoms, it gives advice, asks about symptoms, it gives advice, sometimes it it answers questions and sometimes it it answers questions and sometimes it it answers questions and when it does this, this makes it what's when it does this, this makes it what's when it does this, this makes it what's called a medical device and therefore it called a medical device and therefore it called a medical device and therefore it has to of course be regulated. has to of course be regulated. has to of course be regulated. Um, and we have to prove that it's safe. Um, and we have to prove that it's safe. Um, and we have to prove that it's safe. And regulation essentially bottles down And regulation essentially bottles down And regulation essentially bottles down to three main questions. What does your to three main questions. What does your to three main questions. What does your software do? What could go wrong? And software do? What could go wrong? And software do? What could go wrong? And how do you ensure that it doesn't? how do you ensure that it doesn't? how do you ensure that it doesn't? For normal software, it's quite For normal software, it's quite For normal software, it's quite tractable that question, but for a model tractable that question, but for a model tractable that question, but for a model that talks to actual patients, what that talks to actual patients, what that talks to actual patients, what could go wrong is quite huge. could go wrong is quite huge. could go wrong is quite huge. So, where do we start? We start from So, where do we start? We start from So, where do we start? We start from what could go wrong. We start from the what could go wrong. We start from the what could go wrong. We start from the harm. What could actually harm a harm. What could actually harm a harm. What could actually harm a patient? And let's look at some patient? And let's look at some patient? And let's look at some examples. examples. examples. Um Dora could miss a red flag symptom Um Dora could miss a red flag symptom Um Dora could miss a red flag symptom such as sudden vision loss or severe such as sudden vision loss or severe such as sudden vision loss or severe pain. pain. pain. A patient could ask a medical question A patient could ask a medical question A patient could ask a medical question and Dora invents an answer, hallucinates and Dora invents an answer, hallucinates and Dora invents an answer, hallucinates something. The patient could be something. The patient could be something. The patient could be distressed and Dora just ignores it and distressed and Dora just ignores it and distressed and Dora just ignores it and carries on without actually carries on without actually carries on without actually acknowledging the distress. acknowledging the distress. acknowledging the distress. There's many, many, many documented There's many, many, many documented There's many, many, many documented hazards of these, 20, 30, 40. And we hazards of these, 20, 30, 40. And we hazards of these, 20, 30, 40. And we have to ensure that none of them have to ensure that none of them have to ensure that none of them actually happen in real life.
-
actually happen in real life. actually happen in real life. So, how would we actually normally catch So, how would we actually normally catch So, how would we actually normally catch a problem like this before it actually a problem like this before it actually a problem like this before it actually spreads? We would lean usually on the spreads? We would lean usually on the spreads? We would lean usually on the playbook that most softwares ships on. playbook that most softwares ships on. playbook that most softwares ships on. You ship to a small percentage of You ship to a small percentage of You ship to a small percentage of people. You watch the dashboard. You people. You watch the dashboard. You people. You watch the dashboard. You roll back if it breaks and you iterate roll back if it breaks and you iterate roll back if it breaks and you iterate from there. from there. from there. This is a very good playbook. It's This is a very good playbook. It's This is a very good playbook. It's reactive. It's fast. It's very safe. And reactive. It's fast. It's very safe. And reactive. It's fast. It's very safe. And it's how the industry usually de-risks a it's how the industry usually de-risks a it's how the industry usually de-risks a launch. But there's a hidden assumption launch. But there's a hidden assumption launch. But there's a hidden assumption here that it only works because you can here that it only works because you can here that it only works because you can afford to be wrong for an instance. afford to be wrong for an instance. afford to be wrong for an instance. A bad change hits a few users. You can A bad change hits a few users. You can A bad change hits a few users. You can quickly catch it. You can roll back and quickly catch it. You can roll back and quickly catch it. You can roll back and then no one's actually literally harmed. then no one's actually literally harmed. then no one's actually literally harmed. This is the one assumption is why that This is the one assumption is why that This is the one assumption is why that it breaks when when the actual user is a it breaks when when the actual user is a it breaks when when the actual user is a patient. patient. patient. For 5% that could be a hundreds if not For 5% that could be a hundreds if not For 5% that could be a hundreds if not thousands of of patients that have got thousands of of patients that have got thousands of of patients that have got unproven changes and undue care. unproven changes and undue care. unproven changes and undue care. Roll back. You can't really roll back. Roll back. You can't really roll back. Roll back. You can't really roll back. The call has already happened. The The call has already happened. The The call has already happened. The person has already been harmed. By person has already been harmed. By person has already been harmed. By watching the dashboards, the dashboards watching the dashboards, the dashboards watching the dashboards, the dashboards are just going to going red means that a are just going to going red means that a are just going to going red means that a patient was actually hurt. patient was actually hurt. patient was actually hurt. So, the reactive loop is actually gone So, the reactive loop is actually gone So, the reactive loop is actually gone now. So, how do you iterate at all when now. So, how do you iterate at all when now. So, how do you iterate at all when you can't touch a patient until you're you can't touch a patient until you're you can't touch a patient until you're sure?
-
sure? sure? Well, for this we started looking at Well, for this we started looking at Well, for this we started looking at other at at examples from other high other at at examples from other high other at at examples from other high reliable industries. Uh the most obvious reliable industries. Uh the most obvious reliable industries. Uh the most obvious one is self-driving cars. Obviously, one is self-driving cars. Obviously, one is self-driving cars. Obviously, we're in SF now. There's a lot of Waymos we're in SF now. There's a lot of Waymos we're in SF now. There's a lot of Waymos driving around. They've only just come driving around. They've only just come driving around. They've only just come to London unfortunately, very, very late to London unfortunately, very, very late to London unfortunately, very, very late to the to the party. But what did what to the to the party. But what did what to the to the party. But what did what did self-driving cars do? Well, they did self-driving cars do? Well, they did self-driving cars do? Well, they didn't just drive around crashing into didn't just drive around crashing into didn't just drive around crashing into walls and say, "We won't do that again." walls and say, "We won't do that again." walls and say, "We won't do that again." and then doing another RL loop. and then doing another RL loop. and then doing another RL loop. They put millions of miles of They put millions of miles of They put millions of miles of simulations first before they actually simulations first before they actually simulations first before they actually got any got any got any um passengers in into the car. um passengers in into the car. um passengers in into the car. For us, we believe in the same thing. For us, we believe in the same thing. For us, we believe in the same thing. Simulation is only the real ethical Simulation is only the real ethical Simulation is only the real ethical option we can go with. You can't run all option we can go with. You can't run all option we can go with. You can't run all the hazard the hazards I just mentioned the hazard the hazards I just mentioned the hazard the hazards I just mentioned on real people as a first grasp. on real people as a first grasp. on real people as a first grasp. So, for our clinical history taking, we So, for our clinical history taking, we So, for our clinical history taking, we built a built a built a simulation framework called Matrix. simulation framework called Matrix. simulation framework called Matrix. And I'm going to work through how it And I'm going to work through how it And I'm going to work through how it works and how we use it to prove that works and how we use it to prove that works and how we use it to prove that our product is safe. our product is safe. our product is safe. And uh the paper's on archive if you And uh the paper's on archive if you And uh the paper's on archive if you want to read it along with some of the want to read it along with some of the want to read it along with some of the other research that we do. other research that we do. other research that we do. At its core, Matrix recreates a real At its core, Matrix recreates a real At its core, Matrix recreates a real clinical work conversation but with no clinical work conversation but with no clinical work conversation but with no real patient in it.
-
real patient in it. real patient in it. We use an LLM to play the patient. We We use an LLM to play the patient. We We use an LLM to play the patient. We call it PatBot. call it PatBot. call it PatBot. And what does it do? We use a simulated And what does it do? We use a simulated And what does it do? We use a simulated patient and not a hired actor because patient and not a hired actor because patient and not a hired actor because hired actors don't scale. If you want to hired actors don't scale. If you want to hired actors don't scale. If you want to iterate very fast and simulate different iterate very fast and simulate different iterate very fast and simulate different things at the same time while also things at the same time while also things at the same time while also updating our system, um hiring actors updating our system, um hiring actors updating our system, um hiring actors would just be too slow of a process. So, would just be too slow of a process. So, would just be too slow of a process. So, as a first version, we just use a as a first version, we just use a as a first version, we just use a simulated patient. simulated patient. simulated patient. The simulated patient is conditioned on The simulated patient is conditioned on The simulated patient is conditioned on the actual scenario we want to test. The the actual scenario we want to test. The the actual scenario we want to test. The scenario defines exactly what the scenario defines exactly what the scenario defines exactly what the patient should try and do when talking patient should try and do when talking patient should try and do when talking to our agent. to our agent. to our agent. For example, asking whether the agent is For example, asking whether the agent is For example, asking whether the agent is a human or a or an AI. a human or a or an AI. a human or a or an AI. PatBot then has a conversation with PatBot then has a conversation with PatBot then has a conversation with Dora, our target system, and then Dora, our target system, and then Dora, our target system, and then generates simulated dialogues. generates simulated dialogues. generates simulated dialogues. Very importantly, this all happens under Very importantly, this all happens under Very importantly, this all happens under a very specific clinical use case a very specific clinical use case a very specific clinical use case context. So, the scenarios are grounded context. So, the scenarios are grounded context. So, the scenarios are grounded in real clinical workflows and not in real clinical workflows and not in real clinical workflows and not abstract situations. abstract situations. abstract situations. So, how do we actually make sure that So, how do we actually make sure that So, how do we actually make sure that the the patient is realistic? the the patient is realistic? the the patient is realistic? If if PatBot is is sounding robotic, the If if PatBot is is sounding robotic, the If if PatBot is is sounding robotic, the test aren't really worth much. So, we test aren't really worth much. So, we test aren't really worth much. So, we had to of course validate it.
-
had to of course validate it. had to of course validate it. The first thing we did was just a pure The first thing we did was just a pure The first thing we did was just a pure um script adherence check. If we told um script adherence check. If we told um script adherence check. If we told PatBot to do something, does PatBot do PatBot to do something, does PatBot do PatBot to do something, does PatBot do it? Yes or no? it? Yes or no? it? Yes or no? This helped us filter out a lot of maybe This helped us filter out a lot of maybe This helped us filter out a lot of maybe weaker models that that were that didn't weaker models that that were that didn't weaker models that that were that didn't listen to instructions properly, but listen to instructions properly, but listen to instructions properly, but just purely um following instructions just purely um following instructions just purely um following instructions does not make a realistic patient. We does not make a realistic patient. We does not make a realistic patient. We want a patient that flows more want a patient that flows more want a patient that flows more realistically like a like a real person. realistically like a like a real person. realistically like a like a real person. So, we set up what's called a PPI study, So, we set up what's called a PPI study, So, we set up what's called a PPI study, a patient and public involvement study. a patient and public involvement study. a patient and public involvement study. We took real patients and we showed them We took real patients and we showed them We took real patients and we showed them two sets of conversations. One two sets of conversations. One two sets of conversations. One conversation was with between a real conversation was with between a real conversation was with between a real doctor and a real patient, doctor and a real patient, doctor and a real patient, and one conversation was between Dora and one conversation was between Dora and one conversation was between Dora and PatBot within our Matrix framework. and PatBot within our Matrix framework. and PatBot within our Matrix framework. And we showed them these two examples And we showed them these two examples And we showed them these two examples side by side and said, "Looking at the side by side and said, "Looking at the side by side and said, "Looking at the patient, can you tell which one is the patient, can you tell which one is the patient, can you tell which one is the real person and which one is the real person and which one is the real person and which one is the simulated person?" simulated person?" simulated person?" So, I'm going to just wait for a few So, I'm going to just wait for a few So, I'm going to just wait for a few seconds here if you guys want to quickly seconds here if you guys want to quickly seconds here if you guys want to quickly read the two conversations. read the two conversations. read the two conversations. Um maybe we can do a do a hands up who Um maybe we can do a do a hands up who Um maybe we can do a do a hands up who thinks conversation A is the real thinks conversation A is the real thinks conversation A is the real person? person? person? Who thinks Who thinks Who thinks Who thinks conversation B is the real Who thinks conversation B is the real Who thinks conversation B is the real person?
-
person? person? Okay. I think us as as engineers Okay. I think us as as engineers Okay. I think us as as engineers sometimes are pretty good at at finding sometimes are pretty good at at finding sometimes are pretty good at at finding these things, but um it was actually these things, but um it was actually these things, but um it was actually much more difficult than we thought. And much more difficult than we thought. And much more difficult than we thought. And we did this with four conversations we did this with four conversations we did this with four conversations that's basically. In three out of the that's basically. In three out of the that's basically. In three out of the four, the majority of people actually four, the majority of people actually four, the majority of people actually thought that the simulated patient was thought that the simulated patient was thought that the simulated patient was more realistic. more realistic. more realistic. But the most important thing that we But the most important thing that we But the most important thing that we found was of course there is no single found was of course there is no single found was of course there is no single realistic patient. That doesn't really realistic patient. That doesn't really realistic patient. That doesn't really make sense. Some people prefer to speak make sense. Some people prefer to speak make sense. Some people prefer to speak more verbosely, a lot of ums and ahs. more verbosely, a lot of ums and ahs. more verbosely, a lot of ums and ahs. Some people are more straight to the Some people are more straight to the Some people are more straight to the point, a lot of yeses and nos. But the point, a lot of yeses and nos. But the point, a lot of yeses and nos. But the point is that we actually want to point is that we actually want to point is that we actually want to simulate all these different scenarios. simulate all these different scenarios. simulate all these different scenarios. We want to simulate people with very We want to simulate people with very We want to simulate people with very diverse personas. Um but what it did diverse personas. Um but what it did diverse personas. Um but what it did show us is that at least our PatBot was show us is that at least our PatBot was show us is that at least our PatBot was realistic enough um for this simulation. Okay, so now you've got thousands and Okay, so now you've got thousands and thousands of simulated dialogues. Are us thousands of simulated dialogues. Are us thousands of simulated dialogues. Are us as the engineers going to go read as the engineers going to go read as the engineers going to go read through them one by one and see if a through them one by one and see if a through them one by one and see if a hazard has happened? Of course not, it hazard has happened? Of course not, it hazard has happened? Of course not, it doesn't scale at all. For one, and doesn't scale at all. For one, and doesn't scale at all. For one, and number two, we aren't clinicians, so we number two, we aren't clinicians, so we number two, we aren't clinicians, so we You actually know if an actual hazard You actually know if an actual hazard You actually know if an actual hazard has really occurred. has really occurred. has really occurred. So, we use another LLM as a judge, of So, we use another LLM as a judge, of So, we use another LLM as a judge, of course, and we call it BevJudge.
-
course, and we call it BevJudge. course, and we call it BevJudge. It takes the simulator dialogue, it a It takes the simulator dialogue, it a It takes the simulator dialogue, it a set of expected behaviors, and the set of expected behaviors, and the set of expected behaviors, and the hazardous scenarios that we talked hazardous scenarios that we talked hazardous scenarios that we talked through with clinicians, and it makes a through with clinicians, and it makes a through with clinicians, and it makes a judgment pass or fail. judgment pass or fail. judgment pass or fail. If it fails, it gives us a reason why it If it fails, it gives us a reason why it If it fails, it gives us a reason why it gave that answer. gave that answer. gave that answer. So, we get we get a structured output of So, we get we get a structured output of So, we get we get a structured output of which hazards were triggered and what which hazards were triggered and what which hazards were triggered and what actually went wrong in that scenario. So, how did we validate BevJudge? Uh we So, how did we validate BevJudge? Uh we validated BevJudge against expert validated BevJudge against expert validated BevJudge against expert clinicians. We had we created a corpus clinicians. We had we created a corpus clinicians. We had we created a corpus of 240 examples, and we had a ground of 240 examples, and we had a ground of 240 examples, and we had a ground truth of whether a hazard existed in truth of whether a hazard existed in truth of whether a hazard existed in these conversations, yes or no. Then we these conversations, yes or no. Then we these conversations, yes or no. Then we got 10 clinicians from 10 clinical got 10 clinicians from 10 clinical got 10 clinicians from 10 clinical specialties to to label them for whether specialties to to label them for whether specialties to to label them for whether they had a hazard or not, and we did the they had a hazard or not, and we did the they had a hazard or not, and we did the same thing with the judge. same thing with the judge. same thing with the judge. And the results showed that our judge is And the results showed that our judge is And the results showed that our judge is at least on par, if not slightly better, at least on par, if not slightly better, at least on par, if not slightly better, than the the real expert clinicians. than the the real expert clinicians. than the the real expert clinicians. The top model, which as of a year ago The top model, which as of a year ago The top model, which as of a year ago when we wrote the paper, was Gemini 2.5 when we wrote the paper, was Gemini 2.5 when we wrote the paper, was Gemini 2.5 Pro. Now we've maybe updated the models. Pro. Now we've maybe updated the models. Pro. Now we've maybe updated the models. Um it achieved an F1 score of of 0.96. Um it achieved an F1 score of of 0.96. Um it achieved an F1 score of of 0.96. And even maybe more importantly, it And even maybe more importantly, it And even maybe more importantly, it achieved almost perfect sensitivity. Um achieved almost perfect sensitivity. Um achieved almost perfect sensitivity. Um sensitivity being a very important sensitivity being a very important sensitivity being a very important metric to to healthcare and and to metric to to healthcare and and to metric to to healthcare and and to clinicians, of course, because you want clinicians, of course, because you want clinicians, of course, because you want to make 100% sure almost that there's to make 100% sure almost that there's to make 100% sure almost that there's that no hazards appear in a that no hazards appear in a that no hazards appear in a conversation. You would rather overcall conversation. You would rather overcall conversation. You would rather overcall hazards that aren't there than undercall hazards that aren't there than undercall hazards that aren't there than undercall hazards that are there.
-
hazards that are there. hazards that are there. So, now we have an automated judge that So, now we have an automated judge that So, now we have an automated judge that performs at expert level at at expert performs at expert level at at expert performs at expert level at at expert level, and this is actually what makes level, and this is actually what makes level, and this is actually what makes this whole process scalable. this whole process scalable. this whole process scalable. Okay. So, now Matched can grade Okay. So, now Matched can grade Okay. So, now Matched can grade thousands of conversations, but grading thousands of conversations, but grading thousands of conversations, but grading isn't technically improving the product. isn't technically improving the product. isn't technically improving the product. A pile of pass/fails tells you where A pile of pass/fails tells you where A pile of pass/fails tells you where Dora breaks and where it's not safe, but Dora breaks and where it's not safe, but Dora breaks and where it's not safe, but doesn't actually make the product doesn't actually make the product doesn't actually make the product better. So, how do you do this without better. So, how do you do this without better. So, how do you do this without experimenting on the patient? experimenting on the patient? experimenting on the patient? The answer isn't it like maybe a very The answer isn't it like maybe a very The answer isn't it like maybe a very long ago in our world, 8 months ago, we long ago in our world, 8 months ago, we long ago in our world, 8 months ago, we would manually prompt engineer this. We would manually prompt engineer this. We would manually prompt engineer this. We would look at which agents are going would look at which agents are going would look at which agents are going wrong or which prompts are going wrong, wrong or which prompts are going wrong, wrong or which prompts are going wrong, and you'd have to manually prompt and you'd have to manually prompt and you'd have to manually prompt engineer. engineer. engineer. But, we know that prompt brittleness is But, we know that prompt brittleness is But, we know that prompt brittleness is real, and it's it's quite absurd. real, and it's it's quite absurd. real, and it's it's quite absurd. Formatting changes alone have been seen Formatting changes alone have been seen Formatting changes alone have been seen to swing benchmark by 76 percentage to swing benchmark by 76 percentage to swing benchmark by 76 percentage points. And reordering few-shot examples points. And reordering few-shot examples points. And reordering few-shot examples flips a model from near random, so near flips a model from near random, so near flips a model from near random, so near 50%, to near state-of-the-art on some 50%, to near state-of-the-art on some 50%, to near state-of-the-art on some benchmarks. And hand-tuning can't benchmarks. And hand-tuning can't benchmarks. And hand-tuning can't survive that. It's very subjective, it's survive that. It's very subjective, it's survive that. It's very subjective, it's not reproducible, and very importantly, not reproducible, and very importantly, not reproducible, and very importantly, it's extremely time-consuming. it's extremely time-consuming. it's extremely time-consuming. So, over the last year or so, there's So, over the last year or so, there's So, over the last year or so, there's been these prompt optimizers that have been these prompt optimizers that have been these prompt optimizers that have started to come out, and we've focused started to come out, and we've focused started to come out, and we've focused on those. The one that we use the most on those. The one that we use the most on those. The one that we use the most is is Jeppa, which stands for genetic is is Jeppa, which stands for genetic is is Jeppa, which stands for genetic pareto. It comes from the same people pareto. It comes from the same people pareto. It comes from the same people who made DSPy, if if anyone knows about who made DSPy, if if anyone knows about who made DSPy, if if anyone knows about them.
-
them. them. And how does Jeppa work? You essentially And how does Jeppa work? You essentially And how does Jeppa work? You essentially define a metric for what's good is, define a metric for what's good is, define a metric for what's good is, which I'll I'll get into a bit later. which I'll I'll get into a bit later. which I'll I'll get into a bit later. Then, um you you pass your data through Then, um you you pass your data through Then, um you you pass your data through through um through um through um it through through Jeppa, and and it it through through Jeppa, and and it it through through Jeppa, and and it tells you which examples failed. Then, tells you which examples failed. Then, tells you which examples failed. Then, you get a very strong LLM to reflect on you get a very strong LLM to reflect on you get a very strong LLM to reflect on the failures and update the prompt the failures and update the prompt the failures and update the prompt automatically. You do this over and over automatically. You do this over and over automatically. You do this over and over and over again, and it it keeps what and over again, and it it keeps what and over again, and it it keeps what they call a pareto frontier of the best they call a pareto frontier of the best they call a pareto frontier of the best prompts until your budget has been prompts until your budget has been prompts until your budget has been exhausted, and you've now come up with exhausted, and you've now come up with exhausted, and you've now come up with with what what Jeppa comes up with the with what what Jeppa comes up with the with what what Jeppa comes up with the best prompt. best prompt. best prompt. So, we believe this is a much better So, we believe this is a much better So, we believe this is a much better process from both both a time-consuming process from both both a time-consuming process from both both a time-consuming process, you know, it takes maybe manual process, you know, it takes maybe manual process, you know, it takes maybe manual prompt engineering would take in the prompt engineering would take in the prompt engineering would take in the order of hours to days, to this is in order of hours to days, to this is in order of hours to days, to this is in the hour of minutes. Normally, between the hour of minutes. Normally, between the hour of minutes. Normally, between 30 and an hour minutes and an hour, you 30 and an hour minutes and an hour, you 30 and an hour minutes and an hour, you get an optimized prompt. get an optimized prompt. get an optimized prompt. And very importantly, it's reproducible, And very importantly, it's reproducible, And very importantly, it's reproducible, and there's a very clear order trail and and there's a very clear order trail and and there's a very clear order trail and clear feedback loop. And if anything clear feedback loop. And if anything clear feedback loop. And if anything goes wrong, you can just maybe it's goes wrong, you can just maybe it's goes wrong, you can just maybe it's purely now a data science problem. It's purely now a data science problem. It's purely now a data science problem. It's mainly focused on the data, how to make mainly focused on the data, how to make mainly focused on the data, how to make your data right, the feature your data right, the feature your data right, the feature engineering, and you define the actual engineering, and you define the actual engineering, and you define the actual metric along with the clinicians.
-
So, how do you actually know what's good So, how do you actually know what's good is? is? is? Um it's not a flat accuracy score. You Um it's not a flat accuracy score. You Um it's not a flat accuracy score. You don't want just an average of how your don't want just an average of how your don't want just an average of how your whole data set did. whole data set did. whole data set did. You give it a cost matrix. So, let's go You give it a cost matrix. So, let's go You give it a cost matrix. So, let's go back to our our sensitivity metric. back to our our sensitivity metric. back to our our sensitivity metric. Let's say it's very important for Let's say it's very important for Let's say it's very important for conditions to understand when when and conditions to understand when when and conditions to understand when when and where a red flag is present. where a red flag is present. where a red flag is present. If a red flag is present and you If a red flag is present and you If a red flag is present and you correctly catch it, that's good. If you correctly catch it, that's good. If you correctly catch it, that's good. If you if you miss it, it is it could be if you miss it, it is it could be if you miss it, it is it could be catastrophic. If there's no actual red catastrophic. If there's no actual red catastrophic. If there's no actual red flag and it over calls um um flag and it over calls um um flag and it over calls um um that there's a red flag there, it's just that there's a red flag there, it's just that there's a red flag there, it's just mildly annoying to the patient. They may mildly annoying to the patient. They may mildly annoying to the patient. They may need to ask answer a couple extra need to ask answer a couple extra need to ask answer a couple extra questions, but it's not a cast questions, but it's not a cast questions, but it's not a cast catastrophic harm situation. catastrophic harm situation. catastrophic harm situation. So, we what can we do? We can optimize So, we what can we do? We can optimize So, we what can we do? We can optimize for sensitivity. We can We can work with for sensitivity. We can We can work with for sensitivity. We can We can work with the feedback metric, and we can just the feedback metric, and we can just the feedback metric, and we can just make it um give a higher reward for make it um give a higher reward for make it um give a higher reward for finding the red flags and a lower reward finding the red flags and a lower reward finding the red flags and a lower reward for um for um for um for missing them. So, you can optimize for missing them. So, you can optimize for missing them. So, you can optimize for certain metrics. You can also for certain metrics. You can also for certain metrics. You can also optimize for something else that a that optimize for something else that a that optimize for something else that a that a that a for a clinician might want. a that a for a clinician might want. a that a for a clinician might want. They might want to optimize for accuracy They might want to optimize for accuracy They might want to optimize for accuracy or might want to optimize for some other or might want to optimize for some other or might want to optimize for some other metric. All you have to do is recompile metric. All you have to do is recompile metric. All you have to do is recompile the prompt, and then you got a a new um the prompt, and then you got a a new um the prompt, and then you got a a new um optimized prompt.
-
optimized prompt. optimized prompt. So, remember earlier when we said the we So, remember earlier when we said the we So, remember earlier when we said the we feel like the reactive loop is gone. You feel like the reactive loop is gone. You feel like the reactive loop is gone. You ship, watch, and roll back. This is what ship, watch, and roll back. This is what ship, watch, and roll back. This is what we believe replaces it. we believe replaces it. we believe replaces it. We take real calls, real data. We take real calls, real data. We take real calls, real data. We um we use then synthetic edge cases We um we use then synthetic edge cases We um we use then synthetic edge cases which may not come up in real calls, which may not come up in real calls, which may not come up in real calls, such as like rare symptoms or or such as like rare symptoms or or such as like rare symptoms or or mis-transcriptions. mis-transcriptions. mis-transcriptions. We get an optimized prompt through We get an optimized prompt through We get an optimized prompt through through Jape or some other prompt through Jape or some other prompt through Jape or some other prompt optimizer. optimizer. optimizer. Then we pass it through through Then we pass it through through Then we pass it through through something like Matrix as a simulation something like Matrix as a simulation something like Matrix as a simulation safety gate. If anything fails or safety gate. If anything fails or safety gate. If anything fails or anything doesn't look right, we can redo anything doesn't look right, we can redo anything doesn't look right, we can redo that whole process, redo the data, or or that whole process, redo the data, or or that whole process, redo the data, or or relabel, or get more data. And then relabel, or get more data. And then relabel, or get more data. And then after that and you're happy with that, after that and you're happy with that, after that and you're happy with that, then you can do some gated deploy which then you can do some gated deploy which then you can do some gated deploy which we'll get into a bit. we'll get into a bit. we'll get into a bit. Um the most important thing here is that Um the most important thing here is that Um the most important thing here is that it's a flywheel. Every single deployment it's a flywheel. Every single deployment it's a flywheel. Every single deployment and every new call produces more call and every new call produces more call and every new call produces more call data, so your system is consistently data, so your system is consistently data, so your system is consistently improving. So, as we said with with our matrix So, as we said with with our matrix framework, we use simulated patients. framework, we use simulated patients. framework, we use simulated patients. But, But, But, however realistic you think they are, of however realistic you think they are, of however realistic you think they are, of course they are not real patients.
-
course they are not real patients. course they are not real patients. Passing every test in simulation doesn't Passing every test in simulation doesn't Passing every test in simulation doesn't prove that Dora actually helps someone prove that Dora actually helps someone prove that Dora actually helps someone in real life. The things might come up in real life. The things might come up in real life. The things might come up in real life that that you can't get in in real life that that you can't get in in real life that that you can't get in simulation. It only earns the right to simulation. It only earns the right to simulation. It only earns the right to actually try carefully. Simulation is actually try carefully. Simulation is actually try carefully. Simulation is the inner loop. It's fast, it's free, the inner loop. It's fast, it's free, the inner loop. It's fast, it's free, you can do thousands of runs before you can do thousands of runs before you can do thousands of runs before anyone actually real is is exposed. But, anyone actually real is is exposed. But, anyone actually real is is exposed. But, real patients are the outer loop, and real patients are the outer loop, and real patients are the outer loop, and that's where the only real proof is. So, that's where the only real proof is. So, that's where the only real proof is. So, simulation is necessary, but it's not simulation is necessary, but it's not simulation is necessary, but it's not sufficient. sufficient. sufficient. Simulation earns the right to to to test Simulation earns the right to to to test Simulation earns the right to to to test on real people and then eventually real on real people and then eventually real on real people and then eventually real patients. But, you don't just flip a patients. But, you don't just flip a patients. But, you don't just flip a switch. You cross it in stages, and each switch. You cross it in stages, and each switch. You cross it in stages, and each stage earns the the right for the next. stage earns the the right for the next. stage earns the the right for the next. After you've done your simulations and After you've done your simulations and After you've done your simulations and you're happy with the results, you might you're happy with the results, you might you're happy with the results, you might do a round of user testing. Then you get do a round of user testing. Then you get do a round of user testing. Then you get supervised clinical evaluation based on supervised clinical evaluation based on supervised clinical evaluation based on those tests, and you and you base it on those tests, and you and you base it on those tests, and you and you base it on real patients. You can do some voice real patients. You can do some voice real patients. You can do some voice actors, but of course the the most actors, but of course the the most actors, but of course the the most realistic is to get real patients. realistic is to get real patients. realistic is to get real patients. But, in this step, it's very important But, in this step, it's very important But, in this step, it's very important that there's clinicians at every step in that there's clinicians at every step in that there's clinicians at every step in the loop. the loop. the loop. Then you can do some deployments, uh but Then you can do some deployments, uh but Then you can do some deployments, uh but it's still monitored. And how much it's still monitored. And how much it's still monitored. And how much autonomy you allow the system to do autonomy you allow the system to do autonomy you allow the system to do depends on your evidence. As the system depends on your evidence. As the system depends on your evidence. As the system gets more evidence, you can give it more gets more evidence, you can give it more gets more evidence, you can give it more independence.
-
independence. independence. And underneath all of this, um every And underneath all of this, um every And underneath all of this, um every call, every data set, and every pinned call, every data set, and every pinned call, every data set, and every pinned um um um prompt, every judge verdict traces back prompt, every judge verdict traces back prompt, every judge verdict traces back to the exact hazard that it addresses. to the exact hazard that it addresses. to the exact hazard that it addresses. That's the real deliverable. The That's the real deliverable. The That's the real deliverable. The important thing is that you don't ship important thing is that you don't ship important thing is that you don't ship the model, you ship the evidence when the model, you ship the evidence when the model, you ship the evidence when trying to regulate. trying to regulate. trying to regulate. So, what can you take back to your own So, what can you take back to your own So, what can you take back to your own stacks, both in healthcare and and other stacks, both in healthcare and and other stacks, both in healthcare and and other areas? You first have to define exactly areas? You first have to define exactly areas? You first have to define exactly what harm is for your product. You have what harm is for your product. You have what harm is for your product. You have to manufacture your rare but then but to manufacture your rare but then but to manufacture your rare but then but dangerous cases. Don't wait for them dangerous cases. Don't wait for them dangerous cases. Don't wait for them just to happen naturally. Make your just to happen naturally. Make your just to happen naturally. Make your evaluation metric your real cost evaluation metric your real cost evaluation metric your real cost function to optimize. function to optimize. function to optimize. Pin your prompt versions and keep the Pin your prompt versions and keep the Pin your prompt versions and keep the traces. These are the important things. traces. These are the important things. traces. These are the important things. The The thing also is that the the work The The thing also is that the the work The The thing also is that the the work is never done. As you move into new is never done. As you move into new is never done. As you move into new modalities or new languages, there modalities or new languages, there modalities or new languages, there always be new hazards that start to always be new hazards that start to always be new hazards that start to arise. But with something like Matrix arise. But with something like Matrix arise. But with something like Matrix and a an a prompt optimization loop, um and a an a prompt optimization loop, um and a an a prompt optimization loop, um you can use the same framework as new you can use the same framework as new you can use the same framework as new modalities arise. modalities arise. modalities arise. You know, when you move into voice for You know, when you move into voice for You know, when you move into voice for example, there's things like back example, there's things like back example, there's things like back channeling and interruptions which which channeling and interruptions which which channeling and interruptions which which breaks a lot of the the text emails.
-
breaks a lot of the the text emails. breaks a lot of the the text emails. For example, an agent might be mid For example, an agent might be mid For example, an agent might be mid safety advice like you must avoid bright safety advice like you must avoid bright safety advice like you must avoid bright lights and when the patient maybe um um lights and when the patient maybe um um lights and when the patient maybe um um cuts in with some um cuts in with some um cuts in with some um out of scope question, and weak models out of scope question, and weak models out of scope question, and weak models usually just forget about the safety usually just forget about the safety usually just forget about the safety advice and just ask answer the next advice and just ask answer the next advice and just ask answer the next question. But Matrix captures these question. But Matrix captures these question. But Matrix captures these things. things. things. Um other things Um other things Um other things that that usually go wrong that we've that that usually go wrong that we've that that usually go wrong that we've tested is that um tested is that um tested is that um the the patient could the the Dora or the the patient could the the Dora or the the patient could the the Dora or the or an agent could be halfway through the or an agent could be halfway through the or an agent could be halfway through giving some safety advice and a back giving some safety advice and a back giving some safety advice and a back channel just just cause the model to channel just just cause the model to channel just just cause the model to just completely ignore the safety advice just completely ignore the safety advice just completely ignore the safety advice and and stop right there and and wait and and stop right there and and wait and and stop right there and and wait for the patient to to say something for the patient to to say something for the patient to to say something else. else. else. The failure models The failure modes The The failure models The failure modes The The failure models The failure modes The framework doesn't. You still black box framework doesn't. You still black box framework doesn't. You still black box the system, write down the new hazards, the system, write down the new hazards, the system, write down the new hazards, and simulate and judge them exactly the and simulate and judge them exactly the and simulate and judge them exactly the that we did over text. Voice is just a that we did over text. Voice is just a that we did over text. Voice is just a new module in the same safety case, and new module in the same safety case, and new module in the same safety case, and we actually use Matrix in the voice we actually use Matrix in the voice we actually use Matrix in the voice space to do these exact same things. space to do these exact same things. space to do these exact same things. So, whatever modality you move into, So, whatever modality you move into, So, whatever modality you move into, you're not starting over. The same you're not starting over. The same you're not starting over. The same approach finds the hazards before a real approach finds the hazards before a real approach finds the hazards before a real user does.
-
user does. user does. So again, you simulate before it ever So again, you simulate before it ever So again, you simulate before it ever touches a patient, and you use an touches a patient, and you use an touches a patient, and you use an optimization loop to actually improve optimization loop to actually improve optimization loop to actually improve the system. the system. the system. Thank you very very much for coming to Thank you very very much for coming to Thank you very very much for coming to listen. Um I'm very very open to to listen. Um I'm very very open to to listen. Um I'm very very open to to talking more about this, and if you want talking more about this, and if you want talking more about this, and if you want to connect during the conference or to connect during the conference or to connect during the conference or after, there's my LinkedIn and socials. after, there's my LinkedIn and socials. after, there's my LinkedIn and socials. Thank you very much. Thank you very much. Thank you very much. >> [applause]
Summary
This presentation focuses on the safe deployment of healthcare AI, specifically the Dora clinical conversational agent used for patient calls, referencing the inability to AB test on patients, the permanence of AI-generated information, and the inadequacy of model cards as a sole defense. The key takeaway is that building safe healthcare AI requires operating within strict constraints and rigorous evaluation before patient interaction.