Don’t be data poor — Anuj Iravane, Anterior
Read full transcript 15 segments
-
>> Hello everyone. Welcome to Don't Be Data >> Hello everyone. Welcome to Don't Be Data Poor. Poor. Poor. My name's Anuj. I I lead AI at Anterior. My name's Anuj. I I lead AI at Anterior. My name's Anuj. I I lead AI at Anterior. Um just a bit about Anterior, we are a Um just a bit about Anterior, we are a Um just a bit about Anterior, we are a clinician-led AI company clinician-led AI company clinician-led AI company um built for health plans backed by um built for health plans backed by um built for health plans backed by Sequoia and NEA. Um and what we do is we Sequoia and NEA. Um and what we do is we Sequoia and NEA. Um and what we do is we run AI transformations for health plans. run AI transformations for health plans. run AI transformations for health plans. Um as part of which we build agents for Um as part of which we build agents for Um as part of which we build agents for several high-stakes healthcare several high-stakes healthcare several high-stakes healthcare administrative workflows in production. administrative workflows in production. administrative workflows in production. Um things like prior authorization, Um things like prior authorization, Um things like prior authorization, payment integrity, HEDIS measures, etc. payment integrity, HEDIS measures, etc. payment integrity, HEDIS measures, etc. Um it's it's okay if you're not familiar Um it's it's okay if you're not familiar Um it's it's okay if you're not familiar with any of these workflows with any of these workflows with any of these workflows um because a lot of the work that we do um because a lot of the work that we do um because a lot of the work that we do can actually be summarized can actually be summarized can actually be summarized in in the same way. It's uh in in the same way. It's uh in in the same way. It's uh policy-guided decision-making over policy-guided decision-making over policy-guided decision-making over highly unstructured data. And the unstructured data looks And the unstructured data looks something like this, right? You have uh something like this, right? You have uh something like this, right? You have uh you have these scanned fax bundles you have these scanned fax bundles you have these scanned fax bundles containing medical records full of containing medical records full of containing medical records full of patient information. patient information. patient information. Um a not-so-fun fact is that I think Um a not-so-fun fact is that I think Um a not-so-fun fact is that I think around 70% of medical communication around 70% of medical communication around 70% of medical communication still happens via fax. still happens via fax. still happens via fax. Um and fortunately or unfortunately, Um and fortunately or unfortunately, Um and fortunately or unfortunately, this is the data that we end up working this is the data that we end up working this is the data that we end up working with the most.
-
Um it is a very rich and Um it is a very rich and information-dense data that we see here. information-dense data that we see here. information-dense data that we see here. Um the data distribution here is it Um the data distribution here is it Um the data distribution here is it comes from a very long tail of rare um comes from a very long tail of rare um comes from a very long tail of rare um cases with a nuanced scenarios. Uh it cases with a nuanced scenarios. Uh it cases with a nuanced scenarios. Uh it models an entire clinical trajectory for models an entire clinical trajectory for models an entire clinical trajectory for a patient. Uh and every single person's a patient. Uh and every single person's a patient. Uh and every single person's journey is very different. journey is very different. journey is very different. Uh it also presents itself in various Uh it also presents itself in various Uh it also presents itself in various formats. So, you have like things like formats. So, you have like things like formats. So, you have like things like bad handwriting, tables, checkboxes, um bad handwriting, tables, checkboxes, um bad handwriting, tables, checkboxes, um key-value pairs, images, key-value pairs, images, key-value pairs, images, um a lot of stuff here to deal with. um a lot of stuff here to deal with. um a lot of stuff here to deal with. But I personally think it's a very But I personally think it's a very But I personally think it's a very fascinating source of data that we see fascinating source of data that we see fascinating source of data that we see here. Like it's it's it's like sort of here. Like it's it's it's like sort of here. Like it's it's it's like sort of like an observation through a very fuzzy like an observation through a very fuzzy like an observation through a very fuzzy lens lens lens over an entire person's lifespan. It's over an entire person's lifespan. It's over an entire person's lifespan. It's really unique. really unique. really unique. And I'm sure you must have heard this And I'm sure you must have heard this And I'm sure you must have heard this like enough times today already, but like enough times today already, but like enough times today already, but in healthcare in healthcare in healthcare the the baselines for accuracy are just the the baselines for accuracy are just the the baselines for accuracy are just exceptionally high. exceptionally high. exceptionally high. 95% is not good enough. Um 95% is not good enough. Um 95% is not good enough. Um And at Anterior, this is why we invest And at Anterior, this is why we invest And at Anterior, this is why we invest very deeply in datasets and emails. very deeply in datasets and emails. very deeply in datasets and emails. And these unstructured medical records And these unstructured medical records And these unstructured medical records are a staple source of data for these are a staple source of data for these are a staple source of data for these emails.
-
emails. emails. And we we work with this kind of data in And we we work with this kind of data in And we we work with this kind of data in almost every workflow that we try to almost every workflow that we try to almost every workflow that we try to automate. automate. automate. But the problem is we can't really keep But the problem is we can't really keep But the problem is we can't really keep this data. It's PHI, it's highly this data. It's PHI, it's highly this data. It's PHI, it's highly protected. We can't retain it, we can't protected. We can't retain it, we can't protected. We can't retain it, we can't reuse it, we can't even derive reuse it, we can't even derive reuse it, we can't even derive information from it. And most of our information from it. And most of our information from it. And most of our contracts prohibit us from from doing contracts prohibit us from from doing contracts prohibit us from from doing anything like that. anything like that. anything like that. Um even things like redacting it, Um even things like redacting it, Um even things like redacting it, anonymizing it, and keeping derivative anonymizing it, and keeping derivative anonymizing it, and keeping derivative copies like is a strict no-no, copies like is a strict no-no, copies like is a strict no-no, completely off the table. So, nothing completely off the table. So, nothing completely off the table. So, nothing really survives in any sort of dataset really survives in any sort of dataset really survives in any sort of dataset that we want to persist over a period of that we want to persist over a period of that we want to persist over a period of time. time. time. >> [snorts] >> [snorts] >> [snorts] >> So, so what this talk is about is like >> So, so what this talk is about is like >> So, so what this talk is about is like what do you do when the dataset you most what do you do when the dataset you most what do you do when the dataset you most need is also the data you're least need is also the data you're least need is also the data you're least allowed to keep. And the and the bet that we the answer And the and the bet that we the answer that we put our bets on is that we can that we put our bets on is that we can that we put our bets on is that we can kind of synthetically generate this data kind of synthetically generate this data kind of synthetically generate this data ourselves. There's been a lot of focus on synthetic There's been a lot of focus on synthetic data recently. You have like Frontier data recently. You have like Frontier data recently. You have like Frontier Labs Labs Labs striving to generate synthetic data for striving to generate synthetic data for striving to generate synthetic data for continued pre-training, for RL, for continued pre-training, for RL, for continued pre-training, for RL, for computer use, for agents. computer use, for agents. computer use, for agents. So, it's it's it's a hot topic and it's So, it's it's it's a hot topic and it's So, it's it's it's a hot topic and it's it's a hot topic on our minds as well.
-
it's a hot topic on our minds as well. it's a hot topic on our minds as well. And the moment you say generate, like And the moment you say generate, like And the moment you say generate, like the first thing that comes to mind is, the first thing that comes to mind is, the first thing that comes to mind is, okay, can we can we try to use an LLM to okay, can we can we try to use an LLM to okay, can we can we try to use an LLM to generate synthetic data? And I think you can. I personally And I think you can. I personally believe LLMs are a fantastic believe LLMs are a fantastic believe LLMs are a fantastic tool to generate synthetic data. And tool to generate synthetic data. And tool to generate synthetic data. And several teams have already demonstrated several teams have already demonstrated several teams have already demonstrated this already. There's been some papers this already. There's been some papers this already. There's been some papers in the healthcare space, outside the in the healthcare space, outside the in the healthcare space, outside the healthcare space, people have healthcare space, people have healthcare space, people have successfully used LLMs to generate successfully used LLMs to generate successfully used LLMs to generate synthetic data for for different synthetic data for for different synthetic data for for different purposes. purposes. purposes. There are some known challenges in There are some known challenges in There are some known challenges in trying to use these elements trying to use these elements trying to use these elements to create data especially when you're to create data especially when you're to create data especially when you're trying to one shot the whole process. trying to one shot the whole process. trying to one shot the whole process. Uh it's really hard to generate diverse Uh it's really hard to generate diverse Uh it's really hard to generate diverse realistic looking synthetic records. And realistic looking synthetic records. And realistic looking synthetic records. And this is even more of a problem when this is even more of a problem when this is even more of a problem when you're trying to do you're trying to do you're trying to do when you're trying to do this at scale. when you're trying to do this at scale. when you're trying to do this at scale. So, So, So, uh often times these medical records uh often times these medical records uh often times these medical records are over 300 pages long and it's like are over 300 pages long and it's like are over 300 pages long and it's like imagining if you wouldn't ask an LLM to imagining if you wouldn't ask an LLM to imagining if you wouldn't ask an LLM to write a novel for you in one shot, write a novel for you in one shot, write a novel for you in one shot, right? So, it's the same reason why you right? So, it's the same reason why you right? So, it's the same reason why you wouldn't use an LLM to just one shot a wouldn't use an LLM to just one shot a wouldn't use an LLM to just one shot a synthetic record for you. synthetic record for you. synthetic record for you. Um and LLMs seem to suffer from this Um and LLMs seem to suffer from this Um and LLMs seem to suffer from this very strange mode collapse problem when very strange mode collapse problem when very strange mode collapse problem when it comes to generating like diverse uh it comes to generating like diverse uh it comes to generating like diverse uh data, creative data. And I think there's data, creative data. And I think there's data, creative data. And I think there's two main reasons for it.
-
two main reasons for it. two main reasons for it. Uh the first one is uh Uh the first one is uh Uh the first one is uh like I just mentioned in the talk like I just mentioned in the talk like I just mentioned in the talk earlier, there's very little exposure to earlier, there's very little exposure to earlier, there's very little exposure to this data source in the pre-training this data source in the pre-training this data source in the pre-training data corpus. data corpus. data corpus. Um and today's objectives for Um and today's objectives for Um and today's objectives for pre-training and post-training are are pre-training and post-training are are pre-training and post-training are are largely uh largely uh largely uh they're only they're not incentivized they're only they're not incentivized they're only they're not incentivized for creativity or diversity really. for creativity or diversity really. for creativity or diversity really. They're incentivized to be helpful They're incentivized to be helpful They're incentivized to be helpful systems. So, with these challenges in mind, uh So, with these challenges in mind, uh I'll walk you through like one of our I'll walk you through like one of our I'll walk you through like one of our approaches in how we uh managed to build approaches in how we uh managed to build approaches in how we uh managed to build a pipeline to generate synthetic data. a pipeline to generate synthetic data. a pipeline to generate synthetic data. Um earlier I mentioned our forward tasks Um earlier I mentioned our forward tasks Um earlier I mentioned our forward tasks look something like this, right? So, you look something like this, right? So, you look something like this, right? So, you have workflows and tasks that uh start have workflows and tasks that uh start have workflows and tasks that uh start with some unstructured data and a with some unstructured data and a with some unstructured data and a policy. policy. policy. Um Um Um and you execute your policy against that and you execute your policy against that and you execute your policy against that data. You follow this reasoning trace data. You follow this reasoning trace data. You follow this reasoning trace through it through it through it uh and you arrive at some sort of an uh and you arrive at some sort of an uh and you arrive at some sort of an outcome, which is your label. So, this outcome, which is your label. So, this outcome, which is your label. So, this is our forward task. is our forward task. is our forward task. Uh and the idea we had was to try and Uh and the idea we had was to try and Uh and the idea we had was to try and reverse this process. reverse this process. reverse this process. Uh can we actually start by sampling a Uh can we actually start by sampling a Uh can we actually start by sampling a random label, uh random label, uh random label, uh figuring out a a reasoning trace for figuring out a a reasoning trace for figuring out a a reasoning trace for that label, and then trying to generate that label, and then trying to generate that label, and then trying to generate data backwards from that?
-
data backwards from that? data backwards from that? Uh the idea here being that if you can Uh the idea here being that if you can Uh the idea here being that if you can actually uh sample these two things uh actually uh sample these two things uh actually uh sample these two things uh with enough diversity, with enough diversity, with enough diversity, uh we will have we will be able to uh we will have we will be able to uh we will have we will be able to generate data that's conditioned on generate data that's conditioned on generate data that's conditioned on diverse set of inputs, allowing us to diverse set of inputs, allowing us to diverse set of inputs, allowing us to kind of kind of kind of circumvent the diversity problem a circumvent the diversity problem a circumvent the diversity problem a little bit. Uh so just a quick aside on policies. Uh so just a quick aside on policies. We've talked about policies a bit, but We've talked about policies a bit, but We've talked about policies a bit, but uh let me just clarify what these really uh let me just clarify what these really uh let me just clarify what these really mean, right? So, this is an example mean, right? So, this is an example mean, right? So, this is an example policy we have for for a CPAP device for policy we have for for a CPAP device for policy we have for for a CPAP device for patients. Uh this particular one is for patients. Uh this particular one is for patients. Uh this particular one is for a medical necessity review workflow. a medical necessity review workflow. a medical necessity review workflow. Uh and it it sort of outlines all these Uh and it it sort of outlines all these Uh and it it sort of outlines all these diverse set of conditions uh that a diverse set of conditions uh that a diverse set of conditions uh that a patient might have um in which a CPAP patient might have um in which a CPAP patient might have um in which a CPAP device should be approved or or device should be approved or or device should be approved or or rejected. rejected. rejected. Um so and this policy, as well as many Um so and this policy, as well as many Um so and this policy, as well as many other policies, uh you can think of other policies, uh you can think of other policies, uh you can think of these as uh essentially decision trees these as uh essentially decision trees these as uh essentially decision trees that outline all these sorts of that outline all these sorts of that outline all these sorts of conditions conditions conditions um um um um that that dictate how some outcomes um that that dictate how some outcomes um that that dictate how some outcomes are met. And at Antheir, actually, we we spend a And at Antheir, actually, we we spend a lot of time and energy in trying to lot of time and energy in trying to lot of time and energy in trying to model these policies explicitly as model these policies explicitly as model these policies explicitly as decision trees. decision trees. decision trees. Um uh we work with uh symbolic Um uh we work with uh symbolic Um uh we work with uh symbolic representation uh representation uh representation uh similar to decision trees, uh and it similar to decision trees, uh and it similar to decision trees, uh and it helps us achieve a better accuracy uh helps us achieve a better accuracy uh helps us achieve a better accuracy uh and consistency score when executing and consistency score when executing and consistency score when executing them in LLM-based workflow.
-
them in LLM-based workflow. them in LLM-based workflow. And and the reason why I'm bringing this And and the reason why I'm bringing this And and the reason why I'm bringing this up is that uh by having this sort of up is that uh by having this sort of up is that uh by having this sort of symbolic representation of a policy, you symbolic representation of a policy, you symbolic representation of a policy, you actually have uh a way to kind of actually have uh a way to kind of actually have uh a way to kind of deterministically sample different deterministically sample different deterministically sample different reasoning traces for a given outcome. So, back to our idea of like reversing So, back to our idea of like reversing the process, right? Uh this this the process, right? Uh this this the process, right? Uh this this sampling of reasoning traces from the sampling of reasoning traces from the sampling of reasoning traces from the policies uh what helps us get that policies uh what helps us get that policies uh what helps us get that diverse conditioning input to then diverse conditioning input to then diverse conditioning input to then generate medical records from. generate medical records from. generate medical records from. Uh Uh Uh and the key idea here is that the and the key idea here is that the and the key idea here is that the distribution here uh that we sample from distribution here uh that we sample from distribution here uh that we sample from is is a much more uniform uh and is is a much more uniform uh and is is a much more uniform uh and effective prior distribution than what effective prior distribution than what effective prior distribution than what you'd normally get from an LLM. you'd normally get from an LLM. you'd normally get from an LLM. Uh one added benefit of sampling this Uh one added benefit of sampling this Uh one added benefit of sampling this way is that, in theory, you're able to way is that, in theory, you're able to way is that, in theory, you're able to test uh for far more scenarios than you test uh for far more scenarios than you test uh for far more scenarios than you would likely get from production data would likely get from production data would likely get from production data sources. sources. sources. So, what I mean by that is like say you So, what I mean by that is like say you So, what I mean by that is like say you get a sample of uh 200 cases from your get a sample of uh 200 cases from your get a sample of uh 200 cases from your customer uh customer uh customer uh um um um and and and and you try to like have an and and and and you try to like have an and and and and you try to like have an eval that measures performance against eval that measures performance against eval that measures performance against that, and you get a 95% score. Uh it that, and you get a 95% score. Uh it that, and you get a 95% score. Uh it doesn't really tell you about uh doesn't really tell you about uh doesn't really tell you about uh what you what your performance would be what you what your performance would be what you what your performance would be in those rare edge cases that are not in in those rare edge cases that are not in in those rare edge cases that are not in that data set. There'll always be rare that data set. There'll always be rare that data set. There'll always be rare edge cases uh that are outside the edge cases uh that are outside the edge cases uh that are outside the distribution just because of the fact distribution just because of the fact distribution just because of the fact that our data is so uh highly variant.
-
that our data is so uh highly variant. that our data is so uh highly variant. So, uh for those for those family with So, uh for those for those family with So, uh for those for those family with Cynthia, like uh they follow a similar Cynthia, like uh they follow a similar Cynthia, like uh they follow a similar pattern uh of sampling scenarios from a pattern uh of sampling scenarios from a pattern uh of sampling scenarios from a symbolic causal state representation. symbolic causal state representation. symbolic causal state representation. There's a few of the folks in the space There's a few of the folks in the space There's a few of the folks in the space who are uh working with these symbolic who are uh working with these symbolic who are uh working with these symbolic representations to uh to generate representations to uh to generate representations to uh to generate diversity in synthetic data generation. So, let me walk you through the rest of So, let me walk you through the rest of the pipeline. All right? So, uh once we the pipeline. All right? So, uh once we the pipeline. All right? So, uh once we have this diverse set of samples as a have this diverse set of samples as a have this diverse set of samples as a conditioning input, conditioning input, conditioning input, what we did was we built an LLM-based what we did was we built an LLM-based what we did was we built an LLM-based pipeline that uh follows uh a pipeline that uh follows uh a pipeline that uh follows uh a coarse-to-fine pattern to progressively coarse-to-fine pattern to progressively coarse-to-fine pattern to progressively uh uh build up a medical record layer by uh uh build up a medical record layer by uh uh build up a medical record layer by layer. layer. layer. So, here we first start with creating So, here we first start with creating So, here we first start with creating some patient invariants like the some patient invariants like the some patient invariants like the biological sex, the birth date, the biological sex, the birth date, the biological sex, the birth date, the blood group. blood group. blood group. Uh we use that along with a recent trace Uh we use that along with a recent trace Uh we use that along with a recent trace uh with an LLM again to produce an uh with an LLM again to produce an uh with an LLM again to produce an ordered list of uh events and provider ordered list of uh events and provider ordered list of uh events and provider that a patient might have had, and we that a patient might have had, and we that a patient might have had, and we call this the patient journey. So, this call this the patient journey. So, this call this the patient journey. So, this is a high-level uh you can think of it is a high-level uh you can think of it is a high-level uh you can think of it as a high-level uh overview of what a as a high-level uh overview of what a as a high-level uh overview of what a patient might have gone through in their patient might have gone through in their patient might have gone through in their lifespan lifespan lifespan um um um uh captured by a list of events on a uh captured by a list of events on a uh captured by a list of events on a high uh in natural language.
-
And in the real world, it is actually And in the real world, it is actually only during these uh uh encounters only during these uh uh encounters only during these uh uh encounters provider encounters that documentation provider encounters that documentation provider encounters that documentation is really generated. At least for the is really generated. At least for the is really generated. At least for the data that we get uh uh most of our data data that we get uh uh most of our data data that we get uh uh most of our data source data is generated during these source data is generated during these source data is generated during these provider encounters. So, we model provider encounters. So, we model provider encounters. So, we model exactly that in our pipeline. exactly that in our pipeline. exactly that in our pipeline. Uh Uh Uh we first generate a document plan for we first generate a document plan for we first generate a document plan for each encounter, and then based on that each encounter, and then based on that each encounter, and then based on that and the preceding history of the uh of and the preceding history of the uh of and the preceding history of the uh of the patient, we the patient, we the patient, we uh we fan out into generating the actual uh we fan out into generating the actual uh we fan out into generating the actual documents uh documents uh documents uh um to hydrate them with actual synthetic um to hydrate them with actual synthetic um to hydrate them with actual synthetic information. information. information. Uh and this coarse-to-fine layering uh Uh and this coarse-to-fine layering uh Uh and this coarse-to-fine layering uh is actually what allows us to keep uh is actually what allows us to keep uh is actually what allows us to keep uh uh the different prompt payloads in the uh the different prompt payloads in the uh the different prompt payloads in the pipeline uh very token efficient uh from pipeline uh very token efficient uh from pipeline uh very token efficient uh from both input and output perspective. both input and output perspective. both input and output perspective. While also enabling While also enabling While also enabling this also helps us enable the scale this also helps us enable the scale this also helps us enable the scale across longer patient journey. So, you across longer patient journey. So, you across longer patient journey. So, you can scale this pipeline. You can have a can scale this pipeline. You can have a can scale this pipeline. You can have a much longer patient journey much longer patient journey much longer patient journey and you can just fan out and generate and you can just fan out and generate and you can just fan out and generate documents that way without overloading documents that way without overloading documents that way without overloading the context windows of your LLMs. Finally, we have the sort of refinement Finally, we have the sort of refinement loop in the end loop in the end loop in the end that we that uses a set of emails to that we that uses a set of emails to that we that uses a set of emails to provide feedback provide feedback provide feedback to improve specific parts of the to improve specific parts of the to improve specific parts of the generated documents.
-
generated documents. generated documents. For example, one of the emails we have For example, one of the emails we have For example, one of the emails we have is an LLM based check for consistency is an LLM based check for consistency is an LLM based check for consistency between all documents. So, this makes between all documents. So, this makes between all documents. So, this makes sure that there's no contradictions or sure that there's no contradictions or sure that there's no contradictions or inaccuracies or conflicting information inaccuracies or conflicting information inaccuracies or conflicting information between two documents that are between two documents that are between two documents that are generated. This is important because we generated. This is important because we generated. This is important because we we have a parallel fan out process we have a parallel fan out process we have a parallel fan out process that is used to generate these documents that is used to generate these documents that is used to generate these documents independently. And because we started with the labels And because we started with the labels for this particular pipeline run, for this particular pipeline run, for this particular pipeline run, uh uh uh what we actually also have is an ability what we actually also have is an ability what we actually also have is an ability to kind of use those labels run and and to kind of use those labels run and and to kind of use those labels run and and and compare those against the generated and compare those against the generated and compare those against the generated medical records to see if medical records to see if medical records to see if the tasks that we originally used the tasks that we originally used the tasks that we originally used actually matches actually matches actually matches the data is in accordance with the task the data is in accordance with the task the data is in accordance with the task inputs and outputs. So, we can do this inputs and outputs. So, we can do this inputs and outputs. So, we can do this sort of round trip check to ensure that sort of round trip check to ensure that sort of round trip check to ensure that our data is actually in sync and by our data is actually in sync and by our data is actually in sync and by default get the correct labels by default get the correct labels by default get the correct labels by construction. So, construction. So, construction. So, in theory, in theory, in theory, this is a very nice property to have. this is a very nice property to have. this is a very nice property to have. Like you can Like you can Like you can basically skip the ground truth and basically skip the ground truth and basically skip the ground truth and expensive ground truth in process you expensive ground truth in process you expensive ground truth in process you need need need for for data for fair data.
-
for for data for fair data. for for data for fair data. One thing to clarify here is that One thing to clarify here is that One thing to clarify here is that so far all the generation has been so far all the generation has been so far all the generation has been happening just in plain text and happening just in plain text and happening just in plain text and markdown text. markdown text. markdown text. It is possible to go from that to a It is possible to go from that to a It is possible to go from that to a rendered PDF. But we don't really see rendered PDF. But we don't really see rendered PDF. But we don't really see much value in doing that because much value in doing that because much value in doing that because we have state of the art PDF parsers we have state of the art PDF parsers we have state of the art PDF parsers today today today available to everyone and they just available to everyone and they just available to everyone and they just allow you to convert any sort of complex allow you to convert any sort of complex allow you to convert any sort of complex PDF into a nice markdown representation. PDF into a nice markdown representation. PDF into a nice markdown representation. So, all of this synthetic generation Um, So, all of this synthetic generation Um, So, all of this synthetic generation Um, evaluation happens in the text domain. evaluation happens in the text domain. evaluation happens in the text domain. So, this is just an example of like a So, this is just an example of like a So, this is just an example of like a pipeline that we created from scratch pipeline that we created from scratch pipeline that we created from scratch and it's it's very easy to build. It's and it's it's very easy to build. It's and it's it's very easy to build. It's largely fully LM based. largely fully LM based. largely fully LM based. Um, but but who came up with this, Um, but but who came up with this, Um, but but who came up with this, right? Like who who am I to right? Like who who am I to right? Like who who am I to uh know anything about what a good uh know anything about what a good uh know anything about what a good medical record looks like? medical record looks like? medical record looks like? Uh Uh Uh So, how do we know if this is any good? So, how do we know if this is any good? So, how do we know if this is any good? And uh I think this has been mentioned a And uh I think this has been mentioned a And uh I think this has been mentioned a few times today already, but like you few times today already, but like you few times today already, but like you really don't. Like uh no way AI engineer really don't. Like uh no way AI engineer really don't. Like uh no way AI engineer would ever would. Like you want your would ever would. Like you want your would ever would. Like you want your domain experts to be the ones telling domain experts to be the ones telling domain experts to be the ones telling you what's good, what's not good. Um, you what's good, what's not good. Um, you what's good, what's not good. Um, and which is why we believe that uh it and which is why we believe that uh it and which is why we believe that uh it is of great value to empower your domain is of great value to empower your domain is of great value to empower your domain experts to own your whole data pipeline.
-
And specifically, we uh we do this in And specifically, we uh we do this in two ways, right? Uh we enable our two ways, right? Uh we enable our two ways, right? Uh we enable our clinicians to kind of interject at each clinicians to kind of interject at each clinicians to kind of interject at each point point point uh uh uh in the generation process with a in the generation process with a in the generation process with a human-in-the-loop mechanism. So, at any human-in-the-loop mechanism. So, at any human-in-the-loop mechanism. So, at any point, a clinician can steer the point, a clinician can steer the point, a clinician can steer the generation process to make uh a medical generation process to make uh a medical generation process to make uh a medical record in the way they want it. Uh we record in the way they want it. Uh we record in the way they want it. Uh we often see our clinicians use this uh to often see our clinicians use this uh to often see our clinicians use this uh to to first look at cases that happen in to first look at cases that happen in to first look at cases that happen in production, get some interesting ideas, production, get some interesting ideas, production, get some interesting ideas, and then use that use those ideas along and then use that use those ideas along and then use that use those ideas along with this uh steering in this pipeline with this uh steering in this pipeline with this uh steering in this pipeline to make uh cases that look similar to to make uh cases that look similar to to make uh cases that look similar to what we might see in production or what we might see in production or what we might see in production or they've seen in production. they've seen in production. they've seen in production. And this is what makes the data And this is what makes the data And this is what makes the data generated from this really useful, generated from this really useful, generated from this really useful, right? Like you can actually model your right? Like you can actually model your right? Like you can actually model your uh your failure cases um beforehand or uh your failure cases um beforehand or uh your failure cases um beforehand or even after they after you see them in even after they after you see them in even after they after you see them in production. And secondly, I think most importantly, And secondly, I think most importantly, we let our clinicians also own the whole we let our clinicians also own the whole we let our clinicians also own the whole logic of the pipeline. logic of the pipeline. logic of the pipeline. Um, we do this by modeling the whole Um, we do this by modeling the whole Um, we do this by modeling the whole pipeline as a skills-based workflow pipeline as a skills-based workflow pipeline as a skills-based workflow running on a generic uh agent harness running on a generic uh agent harness running on a generic uh agent harness that we built internally. that we built internally. that we built internally. So, every every uh kind of section here So, every every uh kind of section here So, every every uh kind of section here you see uh all the way from the patient you see uh all the way from the patient you see uh all the way from the patient journey to the document generation to journey to the document generation to journey to the document generation to the document enrichment to the evals, the document enrichment to the evals, the document enrichment to the evals, all of these things are skills all of these things are skills all of these things are skills uh that run on our agent harness.
-
As an example, if a clinician wanted to As an example, if a clinician wanted to uh say maybe add support for a new uh say maybe add support for a new uh say maybe add support for a new document type, let's say for a new document type, let's say for a new document type, let's say for a new customer, they wanted their intake forms customer, they wanted their intake forms customer, they wanted their intake forms to look a certain way, to look a certain way, to look a certain way, they could easily just make a new skill they could easily just make a new skill they could easily just make a new skill file for it, attach it to the pipeline, file for it, attach it to the pipeline, file for it, attach it to the pipeline, and and and voila, there there wouldn't and and and voila, there there wouldn't and and and voila, there there wouldn't be any engineering changes required. So, be any engineering changes required. So, be any engineering changes required. So, it's completely clinician owned from it's completely clinician owned from it's completely clinician owned from that perspective. that perspective. that perspective. And just in a side generally, I feel And just in a side generally, I feel And just in a side generally, I feel like skills are really an amazing like skills are really an amazing like skills are really an amazing interface between AI engineers and interface between AI engineers and interface between AI engineers and domain experts, especially in vertical domain experts, especially in vertical domain experts, especially in vertical AI. AI. AI. We see this being We see this being We see this being we see this being modeled in several of we see this being modeled in several of we see this being modeled in several of our other workflows both for internal our other workflows both for internal our other workflows both for internal use cases and in production as well. So, some results from this, right? So, So, some results from this, right? So, even though we only really use synthetic even though we only really use synthetic even though we only really use synthetic data for evaluation at the moment, data for evaluation at the moment, data for evaluation at the moment, there's already a lot of merits that we there's already a lot of merits that we there's already a lot of merits that we get from it. get from it. get from it. Roughly 90% of our data sets are already Roughly 90% of our data sets are already Roughly 90% of our data sets are already made of synthetic data. made of synthetic data. made of synthetic data. This helps us maintain a very high This helps us maintain a very high This helps us maintain a very high production accuracy score um production accuracy score um production accuracy score um for across many customer deployments.
-
for across many customer deployments. for across many customer deployments. The pipelines that we just showed you The pipelines that we just showed you The pipelines that we just showed you already we are able to achieve a a very already we are able to achieve a a very already we are able to achieve a a very high fidelity on this generated data. high fidelity on this generated data. high fidelity on this generated data. In a blind review, clinicians were not In a blind review, clinicians were not In a blind review, clinicians were not able were only able to distinguish able were only able to distinguish able were only able to distinguish synthetic from real about 60% of the synthetic from real about 60% of the synthetic from real about 60% of the time. So, room for improvement, but time. So, room for improvement, but time. So, room for improvement, but but it's but it's it's close. but it's but it's it's close. but it's but it's it's close. And I'm I'm quite quite it's a quite And I'm I'm quite quite it's a quite And I'm I'm quite quite it's a quite promising avenue for us to invest more promising avenue for us to invest more promising avenue for us to invest more more here. And and the the the the fact more here. And and the the the the fact more here. And and the the the the fact that is the most interesting to me and that is the most interesting to me and that is the most interesting to me and what I'm really what I'm really excited what I'm really what I'm really excited what I'm really what I'm really excited about is that all of these data sets about is that all of these data sets about is that all of these data sets well, most of our data sets today then well, most of our data sets today then well, most of our data sets today then are created just in time for these are created just in time for these are created just in time for these customer deployments, right? You can you customer deployments, right? You can you customer deployments, right? You can you when you have the ability to like create when you have the ability to like create when you have the ability to like create data from scratch so quickly, data from scratch so quickly, data from scratch so quickly, you can kind of you can kind of you can kind of you don't need to depend on on on you don't need to depend on on on you don't need to depend on on on waiting for data from your customer. You waiting for data from your customer. You waiting for data from your customer. You can kind of just model all your edge can kind of just model all your edge can kind of just model all your edge cases, simulate them, and test your cases, simulate them, and test your cases, simulate them, and test your workflows before you go live with the workflows before you go live with the workflows before you go live with the production production production go live in production. So, some takeaways if you're looking to So, some takeaways if you're looking to build your own synthetic data pipeline build your own synthetic data pipeline build your own synthetic data pipeline in healthcare or even another domain, in healthcare or even another domain, in healthcare or even another domain, try reversing your inference workflow.
-
try reversing your inference workflow. try reversing your inference workflow. Diversity should always be sampled from Diversity should always be sampled from Diversity should always be sampled from a from an appropriate distribution for a from an appropriate distribution for a from an appropriate distribution for your use case. your use case. your use case. Try to emulate the process in which Try to emulate the process in which Try to emulate the process in which data was actually generated. data was actually generated. data was actually generated. So, like I showed you, we were trying to So, like I showed you, we were trying to So, like I showed you, we were trying to sort of like we were using LLMs we're sort of like we were using LLMs we're sort of like we were using LLMs we're trying to emulate how trying to emulate how trying to emulate how our medical records might actually be our medical records might actually be our medical records might actually be generated during patient encounters. generated during patient encounters. generated during patient encounters. So, and I I would highly recommend you So, and I I would highly recommend you So, and I I would highly recommend you try doing that. And the fourth most try doing that. And the fourth most try doing that. And the fourth most important thing I think is when you're important thing I think is when you're important thing I think is when you're when you're making a data pipeline like when you're making a data pipeline like when you're making a data pipeline like this, it's really important to give your this, it's really important to give your this, it's really important to give your domain experts the keys because these domain experts the keys because these domain experts the keys because these are the people who know are the people who know are the people who know about your data and and and they will about your data and and and they will about your data and and and they will help you help you help you drive towards a recursive drive towards a recursive drive towards a recursive self-improvement and not the AI self-improvement and not the AI self-improvement and not the AI engineers. Cool. So, you don't need a PHI problem Cool. So, you don't need a PHI problem for this anywhere. for this anywhere. for this anywhere. The data you need is ephemeral, The data you need is ephemeral, The data you need is ephemeral, sensitive, or even expensive to label. sensitive, or even expensive to label. sensitive, or even expensive to label. You can think about You can think about You can think about generating data yourself and hopefully generating data yourself and hopefully generating data yourself and hopefully you won't be data poor. you won't be data poor. you won't be data poor. Thank you, everyone.
Summary
The main theme is navigating and extracting insights from highly unstructured and protected healthcare data, primarily from faxes. Key subjects include administrative workflows like prior authorization and payment integrity, and the challenge of using this data due to privacy regulations. The practical takeaway is the necessity of advanced AI solutions to derive value from this data while adhering to strict privacy and accuracy standards.