← Back
AI Engineer July 29, 2026 21m

Persona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.ai

Read full transcript 18 segments
  1. >> Hello, I'm Ishan and welcome to can AI >> Hello, I'm Ishan and welcome to can AI predict people like we predict the predict people like we predict the predict people like we predict the weather. A field guide to the nascent weather. A field guide to the nascent weather. A field guide to the nascent field of synthetic personas. field of synthetic personas. field of synthetic personas. Now, I'm sure most of you in this room Now, I'm sure most of you in this room Now, I'm sure most of you in this room at some point or another have prompted a at some point or another have prompted a at some point or another have prompted a large language model with a role prompt. large language model with a role prompt. large language model with a role prompt. You are a fill in the blank and then the You are a fill in the blank and then the You are a fill in the blank and then the task. task. task. Believe it or not, that core principle Believe it or not, that core principle Believe it or not, that core principle of steering a model's outputs as if it of steering a model's outputs as if it of steering a model's outputs as if it were a particular person or persona were a particular person or persona were a particular person or persona has turned into an entire category that has turned into an entire category that has turned into an entire category that companies are using to test product companies are using to test product companies are using to test product concepts and messaging against synthetic concepts and messaging against synthetic concepts and messaging against synthetic respondents. And it has moved from a respondents. And it has moved from a respondents. And it has moved from a novelty to market momentum, as you can novelty to market momentum, as you can novelty to market momentum, as you can see from both these headlines as well as see from both these headlines as well as see from both these headlines as well as the increase in funding for the last few the increase in funding for the last few the increase in funding for the last few years. years. years. And the running analogy I want to leave And the running analogy I want to leave And the running analogy I want to leave you with is that synthetic personas are you with is that synthetic personas are you with is that synthetic personas are like weather forecasting. like weather forecasting. like weather forecasting. Like weather forecasting, they were Like weather forecasting, they were Like weather forecasting, they were unlocked thanks to an increase in unlocked thanks to an increase in unlocked thanks to an increase in compute and data. compute and data. compute and data. And like weather forecasting, they And like weather forecasting, they And like weather forecasting, they operate within a particular regime and operate within a particular regime and operate within a particular regime and going past that sometimes can go outside going past that sometimes can go outside going past that sometimes can go outside of where they're accurate. So, for of where they're accurate. So, for of where they're accurate. So, for example, you can only predict the example, you can only predict the example, you can only predict the weather a certain number of days in weather a certain number of days in weather a certain number of days in advance. Similarly with synthetic advance. Similarly with synthetic advance. Similarly with synthetic personas, there's only so far you can go personas, there's only so far you can go personas, there's only so far you can go before you'll run into issues. And before you'll run into issues. And before you'll run into issues. And understanding those issues are as understanding those issues are as understanding those issues are as important as understanding their promise important as understanding their promise important as understanding their promise and their potential.

  2. So, I'm Ishan Nand. I'm the chief AI So, I'm Ishan Nand. I'm the chief AI officer at Insight Sciences. We officer at Insight Sciences. We officer at Insight Sciences. We construct LLM synthetic personas for construct LLM synthetic personas for construct LLM synthetic personas for market research and market insights market research and market insights market research and market insights teams. teams. teams. And the reason for this talk is that And the reason for this talk is that And the reason for this talk is that most of the coverage in the space is most of the coverage in the space is most of the coverage in the space is very shallow, doesn't go into the very shallow, doesn't go into the very shallow, doesn't go into the technical details. It's either outright technical details. It's either outright technical details. It's either outright hype or outright dismissal. And hype or outright dismissal. And hype or outright dismissal. And it's really hard to separate the noise it's really hard to separate the noise it's really hard to separate the noise from what's real. So what I want to from what's real. So what I want to from what's real. So what I want to cover is that messy middle of the cover is that messy middle of the cover is that messy middle of the technical details. And you don't have to technical details. And you don't have to technical details. And you don't have to take my word for it, even though I'm a take my word for it, even though I'm a take my word for it, even though I'm a vendor in the space, because everything vendor in the space, because everything vendor in the space, because everything I'm going to talk about today is going I'm going to talk about today is going I'm going to talk about today is going to be based on published research. So to be based on published research. So to be based on published research. So we're going to cover why now for we're going to cover why now for we're going to cover why now for synthetic personas, how they fail, some synthetic personas, how they fail, some synthetic personas, how they fail, some techniques to inspire you, and then some techniques to inspire you, and then some techniques to inspire you, and then some metrics to judge whether your synthetic metrics to judge whether your synthetic metrics to judge whether your synthetic persona is accurate or not. persona is accurate or not. persona is accurate or not. Speaking of weather forecasting, another Speaking of weather forecasting, another Speaking of weather forecasting, another parallel is just like in the 1950s and parallel is just like in the 1950s and parallel is just like in the 1950s and '60s, we got computers that promised us, '60s, we got computers that promised us, '60s, we got computers that promised us, correctly, a future of accurate weather correctly, a future of accurate weather correctly, a future of accurate weather forecasts.

  3. forecasts. forecasts. We were also promised, believe it or We were also promised, believe it or We were also promised, believe it or not, people forecasts. This company, not, people forecasts. This company, not, people forecasts. This company, Simulmatics, were extensively covered by Simulmatics, were extensively covered by Simulmatics, were extensively covered by Jill Lepore, Jill Lepore, Jill Lepore, promised that they could simulate and promised that they could simulate and promised that they could simulate and predict the electorate using raw predict the electorate using raw predict the electorate using raw statistics and the computational power statistics and the computational power statistics and the computational power at the time. Fortunately, that turned at the time. Fortunately, that turned at the time. Fortunately, that turned out not to be the case. So you should out not to be the case. So you should out not to be the case. So you should approach claims like this with some approach claims like this with some approach claims like this with some humility. But we have something they did humility. But we have something they did humility. But we have something they did not have then. And that unlock is, not have then. And that unlock is, not have then. And that unlock is, again, more computational power, but again, more computational power, but again, more computational power, but also better modeling thanks to LLMs. And also better modeling thanks to LLMs. And also better modeling thanks to LLMs. And LLMs unlock a new kind of simulation. LLMs unlock a new kind of simulation. LLMs unlock a new kind of simulation. For the longest time, to simulate For the longest time, to simulate For the longest time, to simulate something meant to mathematize it in something meant to mathematize it in something meant to mathematize it in formulas or equations. But certain formulas or equations. But certain formulas or equations. But certain things, how we feel, how we act, what things, how we feel, how we act, what things, how we feel, how we act, what choices we make, aren't always choices we make, aren't always choices we make, aren't always succumbing to the equations. What LLMs succumbing to the equations. What LLMs succumbing to the equations. What LLMs offer us is a new medium, a new atomic offer us is a new medium, a new atomic offer us is a new medium, a new atomic unit of language itself that we can unit of language itself that we can unit of language itself that we can model against. Now granted, they are model against. Now granted, they are model against. Now granted, they are based on math under the hood, but it based on math under the hood, but it based on math under the hood, but it gives us this intermediary layer that we gives us this intermediary layer that we gives us this intermediary layer that we can construct and simulate against that can construct and simulate against that can construct and simulate against that we couldn't before.

  4. we couldn't before. we couldn't before. And the process can work. And the process can work. And the process can work. I want to share with you one of the most I want to share with you one of the most I want to share with you one of the most well-known demonstrations of this in the well-known demonstrations of this in the well-known demonstrations of this in the field. field. field. What they did is they took about a What they did is they took about a What they did is they took about a thousand humans. thousand humans. thousand humans. They put them through about two and a They put them through about two and a They put them through about two and a half hours of extensive interviews about half hours of extensive interviews about half hours of extensive interviews about their background and their views and their background and their views and their background and their views and their attitudes. And they put those their attitudes. And they put those their attitudes. And they put those people through a battery of personality people through a battery of personality people through a battery of personality tests and surveys. tests and surveys. tests and surveys. Then they took those transcripts and Then they took those transcripts and Then they took those transcripts and they passed it to an AI agent and they they passed it to an AI agent and they they passed it to an AI agent and they had the AI agent take the same set of had the AI agent take the same set of had the AI agent take the same set of surveys and personality tests. surveys and personality tests. surveys and personality tests. What they found was as the agents were What they found was as the agents were What they found was as the agents were basically about 83% aligned and basically about 83% aligned and basically about 83% aligned and predictive to the corresponding humans predictive to the corresponding humans predictive to the corresponding humans they were modeled against. they were modeled against. they were modeled against. Now one caveat is that number is Now one caveat is that number is Now one caveat is that number is normalized against the uncertainty and normalized against the uncertainty and normalized against the uncertainty and noise of the humans themselves. It's a noise of the humans themselves. It's a noise of the humans themselves. It's a theme we're going to come back to at the theme we're going to come back to at the theme we're going to come back to at the end of this talk. end of this talk. end of this talk. But don't get too excited because But don't get too excited because But don't get too excited because synthetic personas are different from synthetic personas are different from synthetic personas are different from regular experiments and they're liable regular experiments and they're liable regular experiments and they're liable to confuse and fool you if you don't to confuse and fool you if you don't to confuse and fool you if you don't know how they fail. So I'm going to know how they fail. So I'm going to know how they fail. So I'm going to cover three important failure modes that cover three important failure modes that cover three important failure modes that you need to know about when dealing with you need to know about when dealing with you need to know about when dealing with synthetic personas.

  5. synthetic personas. synthetic personas. To understand the first one, To understand the first one, To understand the first one, I want to consider this prompt these I want to consider this prompt these I want to consider this prompt these researchers gave. It's a very un It's a researchers gave. It's a very un It's a researchers gave. It's a very un It's a very ambiguous and very unsophisticated very ambiguous and very unsophisticated very ambiguous and very unsophisticated prompt. It basically says you're a prompt. It basically says you're a prompt. It basically says you're a customer, I'm going to show you a customer, I'm going to show you a customer, I'm going to show you a product, I'm going to tell you the product, I'm going to tell you the product, I'm going to tell you the category, I'm going to give you its category, I'm going to give you its category, I'm going to give you its price. Those are going to be the price. Those are going to be the price. Those are going to be the variables in the template and then I'm variables in the template and then I'm variables in the template and then I'm going to ask you to say whether you're going to ask you to say whether you're going to ask you to say whether you're going to purchase or not purchase. going to purchase or not purchase. going to purchase or not purchase. Willingness to pay, willingness to Willingness to pay, willingness to Willingness to pay, willingness to purchase is basically the test. purchase is basically the test. purchase is basically the test. What the researchers did is they What the researchers did is they What the researchers did is they recruited a panel of humans and put them recruited a panel of humans and put them recruited a panel of humans and put them through the same test and then they put through the same test and then they put through the same test and then they put the synthetic personas through the same the synthetic personas through the same the synthetic personas through the same test. test. test. What they found is very interesting. What they found is very interesting. What they found is very interesting. So the humans are here in red. So the humans are here in red. So the humans are here in red. They do exactly what you would expect They do exactly what you would expect They do exactly what you would expect from basic economic theory. As the from basic economic theory. As the from basic economic theory. As the purchase price increases, we see that purchase price increases, we see that purchase price increases, we see that the purchase probability goes down, the purchase probability goes down, the purchase probability goes down, slopes downward. slopes downward. slopes downward. But the LLMs did something different. But the LLMs did something different. But the LLMs did something different. They had this inverted U-shaped curve. They had this inverted U-shaped curve. They had this inverted U-shaped curve. And particularly problematic is this And particularly problematic is this And particularly problematic is this area right here, where as the price is area right here, where as the price is area right here, where as the price is increasing, the purchase probability is increasing, the purchase probability is increasing, the purchase probability is going up. That seems really bizarre.

  6. going up. That seems really bizarre. going up. That seems really bizarre. Through a series of additional Through a series of additional Through a series of additional experiments, what they discovered was experiments, what they discovered was experiments, what they discovered was that the LLM was using the price as a that the LLM was using the price as a that the LLM was using the price as a proxy for other properties about the proxy for other properties about the proxy for other properties about the product that the humans were considering product that the humans were considering product that the humans were considering were fixed. were fixed. were fixed. Things like the expiration date based on Things like the expiration date based on Things like the expiration date based on the price, what the price of competing the price, what the price of competing the price, what the price of competing products were also as the price changed. products were also as the price changed. products were also as the price changed. And those correlations, those latent And those correlations, those latent And those correlations, those latent confounders that weren't clear and confounders that weren't clear and confounders that weren't clear and immediate, were actually confusing the immediate, were actually confusing the immediate, were actually confusing the result. result. result. And the way to think about this is when And the way to think about this is when And the way to think about this is when an LLM is missing context, it has to an LLM is missing context, it has to an LLM is missing context, it has to potentially infer or invent confounders. potentially infer or invent confounders. potentially infer or invent confounders. Right? When we do a human experiment, if Right? When we do a human experiment, if Right? When we do a human experiment, if I put it like a gold watch on a table, I I put it like a gold watch on a table, I I put it like a gold watch on a table, I ask a human to walk in and estimate the ask a human to walk in and estimate the ask a human to walk in and estimate the price of it, everything about the price of it, everything about the price of it, everything about the environment is fairly fixed. The human environment is fairly fixed. The human environment is fairly fixed. The human and their decisions are the random and their decisions are the random and their decisions are the random variable. In a synthetic experiment, if variable. In a synthetic experiment, if variable. In a synthetic experiment, if you don't set it up properly, other you don't set it up properly, other you don't set it up properly, other parts of it actually become part of the parts of it actually become part of the parts of it actually become part of the random variable itself. I like to say if random variable itself. I like to say if random variable itself. I like to say if it's a poorly grounded persona, it's a it's a poorly grounded persona, it's a it's a poorly grounded persona, it's a little like the LLM is playing improv little like the LLM is playing improv little like the LLM is playing improv with you.

  7. with you. with you. It's like gold watch on a table? Oh, It's like gold watch on a table? Oh, It's like gold watch on a table? Oh, well, we must be in a jewelry store, well, we must be in a jewelry store, well, we must be in a jewelry store, right? It has to infer what's likely. right? It has to infer what's likely. right? It has to infer what's likely. And maybe this is a rich person, so And maybe this is a rich person, so And maybe this is a rich person, so they're more likely to purchase. And so they're more likely to purchase. And so they're more likely to purchase. And so the lesson is, we need to richly ground the lesson is, we need to richly ground the lesson is, we need to richly ground our personas in the personality, the our personas in the personality, the our personas in the personality, the context, and bizarrely, even the study's context, and bizarrely, even the study's context, and bizarrely, even the study's own construction. In a human subject own construction. In a human subject own construction. In a human subject experiment, you want to hide the study experiment, you want to hide the study experiment, you want to hide the study construction from the participant. But construction from the participant. But construction from the participant. But in the case of an LLM, they have no in the case of an LLM, they have no in the case of an LLM, they have no universe other than what's in the universe other than what's in the universe other than what's in the prompt, and you have to use the prompt prompt, and you have to use the prompt prompt, and you have to use the prompt to paint the world to prevent any type to paint the world to prevent any type to paint the world to prevent any type of confounders. of confounders. of confounders. Another failure mode is prompt Another failure mode is prompt Another failure mode is prompt sensitivity. So, here's a researcher sensitivity. So, here's a researcher sensitivity. So, here's a researcher that took a question, they give the same that took a question, they give the same that took a question, they give the same question, same choices, they just question, same choices, they just question, same choices, they just swapped the order of the choices. Yes swapped the order of the choices. Yes swapped the order of the choices. Yes was the first one in the first question, was the first one in the first question, was the first one in the first question, yes was the second option in the the yes was the second option in the the yes was the second option in the the second question. What they found was second question. What they found was second question. What they found was that the model had an extremely strong that the model had an extremely strong that the model had an extremely strong order bias. order bias. order bias. Basically, when they took the two Basically, when they took the two Basically, when they took the two results and they averaged them together, results and they averaged them together, results and they averaged them together, it washed out into noise, into 50/50. it washed out into noise, into 50/50. it washed out into noise, into 50/50. Now, humans do have a first order bias, Now, humans do have a first order bias, Now, humans do have a first order bias, but not to this extent. And so, the but not to this extent. And so, the but not to this extent. And so, the lesson here is that we need to lesson here is that we need to lesson here is that we need to durability test our personas to durability test our personas to durability test our personas to understand how they will change under understand how they will change under understand how they will change under reorderings, under rewordings, and even reorderings, under rewordings, and even reorderings, under rewordings, and even adversarial challenges to their adversarial challenges to their adversarial challenges to their opinions.

  8. opinions. opinions. The third and final area that I want to The third and final area that I want to The third and final area that I want to highlight is that LLMs are trained on highlight is that LLMs are trained on highlight is that LLMs are trained on what people say, what people say, what people say, and they're not trained on what people and they're not trained on what people and they're not trained on what people do. do. do. So, as a consequence, predicting stated So, as a consequence, predicting stated So, as a consequence, predicting stated attitudes tend to be easier than attitudes tend to be easier than attitudes tend to be easier than predicting actions or behaviors. Both predicting actions or behaviors. Both predicting actions or behaviors. Both because they're clear and likely to be because they're clear and likely to be because they're clear and likely to be in the text, but also because they are in the text, but also because they are in the text, but also because they are natively text themselves. So, this chart natively text themselves. So, this chart natively text themselves. So, this chart is from a bunch of researchers that used is from a bunch of researchers that used is from a bunch of researchers that used an LLM to try and predict known social an LLM to try and predict known social an LLM to try and predict known social science experiments. science experiments. science experiments. The original point of this chart is to The original point of this chart is to The original point of this chart is to show that the LLMs are about as good as show that the LLMs are about as good as show that the LLMs are about as good as the experts. LLM is in black in a the experts. LLM is in black in a the experts. LLM is in black in a circle, the experts are in blue, and you circle, the experts are in blue, and you circle, the experts are in blue, and you can see they're both doing about equally can see they're both doing about equally can see they're both doing about equally well in making the prediction. But, the well in making the prediction. But, the well in making the prediction. But, the point I want to draw you to is that point I want to draw you to is that point I want to draw you to is that there are two categories of experiments there are two categories of experiments there are two categories of experiments here. The top are surveys, those are here. The top are surveys, those are here. The top are surveys, those are natively language and text-based, and natively language and text-based, and natively language and text-based, and those reflect attitudes. And on the those reflect attitudes. And on the those reflect attitudes. And on the whole, the models tend to do better whole, the models tend to do better whole, the models tend to do better there. there. there. The bottom half is field experiments. The bottom half is field experiments. The bottom half is field experiments. Those are behaviors, and those are Those are behaviors, and those are Those are behaviors, and those are things that need to be transcribed into things that need to be transcribed into things that need to be transcribed into actions. They're less likely to be in actions. They're less likely to be in actions. They're less likely to be in the training data, and correspondingly, the training data, and correspondingly, the training data, and correspondingly, the LLM doesn't do as well.

  9. the LLM doesn't do as well. the LLM doesn't do as well. So, So, So, the lesson we often tell our clients is the lesson we often tell our clients is the lesson we often tell our clients is consider questions that triangulate to consider questions that triangulate to consider questions that triangulate to behavior from attitudes. As a behavior from attitudes. As a behavior from attitudes. As a hypothetical example, if you want to hypothetical example, if you want to hypothetical example, if you want to know about gym attendance, you might be know about gym attendance, you might be know about gym attendance, you might be better off, well, you can ask about better off, well, you can ask about better off, well, you can ask about both, but asking about attitudes towards both, but asking about attitudes towards both, but asking about attitudes towards working out rather than asking about working out rather than asking about working out rather than asking about attendance and see if that's a suitable attendance and see if that's a suitable attendance and see if that's a suitable proxy. proxy. proxy. Okay. Okay. Okay. Now, let's talk about three example Now, let's talk about three example Now, let's talk about three example techniques to kind of inspire your own techniques to kind of inspire your own techniques to kind of inspire your own synthetic personas. synthetic personas. synthetic personas. So, the first one is just prompting the So, the first one is just prompting the So, the first one is just prompting the model. Uh this right here is from the model. Uh this right here is from the model. Uh this right here is from the Argyle paper, which is really one of the Argyle paper, which is really one of the Argyle paper, which is really one of the seminal papers in this field. In fact, seminal papers in this field. In fact, seminal papers in this field. In fact, it's so early that the model they used it's so early that the model they used it's so early that the model they used was a text completion model. That's why was a text completion model. That's why was a text completion model. That's why this prompt isn't in the form of a chat. this prompt isn't in the form of a chat. this prompt isn't in the form of a chat. It's a statement of I am. So, they gave It's a statement of I am. So, they gave It's a statement of I am. So, they gave it a prompt that said, and for example, it a prompt that said, and for example, it a prompt that said, and for example, the middle column is basically where the the middle column is basically where the the middle column is basically where the context is. I am a strong liberal. I context is. I am a strong liberal. I context is. I am a strong liberal. I support progressive values, etc., etc. support progressive values, etc., etc. support progressive values, etc., etc. And at the end it says, "In 2016, I And at the end it says, "In 2016, I And at the end it says, "In 2016, I voted for." And they basically voted for." And they basically voted for." And they basically let the model sample its completions and let the model sample its completions and let the model sample its completions and it says, uh it says, uh it says, uh Hillary Clinton, Bernie Sanders, Hillary Hillary Clinton, Bernie Sanders, Hillary Hillary Clinton, Bernie Sanders, Hillary Clinton, and so forth. And you can see Clinton, and so forth. And you can see Clinton, and so forth. And you can see what happens with the conservative case what happens with the conservative case what happens with the conservative case on the top.

  10. on the top. on the top. Since this time, obviously, there've Since this time, obviously, there've Since this time, obviously, there've been a lot more prompting techniques. been a lot more prompting techniques. been a lot more prompting techniques. And a lot more models. And I can't tell And a lot more models. And I can't tell And a lot more models. And I can't tell you which prompting technique and which you which prompting technique and which you which prompting technique and which model is going to work best for your use model is going to work best for your use model is going to work best for your use case. What you are going to have to do case. What you are going to have to do case. What you are going to have to do is figure it out empirically by is figure it out empirically by is figure it out empirically by validating against some known human validating against some known human validating against some known human ground truth data. You'll have to do ground truth data. You'll have to do ground truth data. You'll have to do what these guys did. So, for example, what these guys did. So, for example, what these guys did. So, for example, here in this research, they're trying to here in this research, they're trying to here in this research, they're trying to figure out how well they can construct figure out how well they can construct figure out how well they can construct personas to represent voting patterns. personas to represent voting patterns. personas to represent voting patterns. What they found was they compared here What they found was they compared here What they found was they compared here on the left is reality and on the right on the left is reality and on the right on the left is reality and on the right is their four different types of persona is their four different types of persona is their four different types of persona constructions. And they didn't realize constructions. And they didn't realize constructions. And they didn't realize it at the time, but their persona it at the time, but their persona it at the time, but their persona construction was actually amplifying construction was actually amplifying construction was actually amplifying bias within the model as they got more bias within the model as they got more bias within the model as they got more and more detailed. And they found it was and more detailed. And they found it was and more detailed. And they found it was actually throwing it further and further actually throwing it further and further actually throwing it further and further astray from reality. So, you probably astray from reality. So, you probably astray from reality. So, you probably have a bunch of different ideas. The have a bunch of different ideas. The have a bunch of different ideas. The answer is you're going to have to test answer is you're going to have to test answer is you're going to have to test it and validate it against ground truth. it and validate it against ground truth. it and validate it against ground truth. The other natural thing you might expect The other natural thing you might expect The other natural thing you might expect is well, hey, we can fine-tune it, is well, hey, we can fine-tune it, is well, hey, we can fine-tune it, especially if it's missing data that especially if it's missing data that especially if it's missing data that isn't there, especially for example, if isn't there, especially for example, if isn't there, especially for example, if it's behaviors or something that it's behaviors or something that it's behaviors or something that wouldn't be in the training text. And wouldn't be in the training text. And wouldn't be in the training text. And this is a great paper to be inspired by this is a great paper to be inspired by this is a great paper to be inspired by for this. This is the Subpop paper.

  11. for this. This is the Subpop paper. for this. This is the Subpop paper. Basically, they construct a prompt Basically, they construct a prompt Basically, they construct a prompt template, which is the demographic template, which is the demographic template, which is the demographic information, then the survey question information, then the survey question information, then the survey question they want to ask, and then they compare they want to ask, and then they compare they want to ask, and then they compare the known human data distribution to the the known human data distribution to the the known human data distribution to the distribution that comes out of the distribution that comes out of the distribution that comes out of the model, and they do fine-tuning until the model, and they do fine-tuning until the model, and they do fine-tuning until the model and the human data align. model and the human data align. model and the human data align. Now, here's the interesting thing. Now, here's the interesting thing. Now, here's the interesting thing. When they did this, as you'd expect, When they did this, as you'd expect, When they did this, as you'd expect, the results that were from the the results that were from the the results that were from the populations they gave to the model, populations they gave to the model, populations they gave to the model, that's the ones in blue, improved. that's the ones in blue, improved. that's the ones in blue, improved. But very interestingly, the ones in But very interestingly, the ones in But very interestingly, the ones in white also improved by almost the same white also improved by almost the same white also improved by almost the same degree. Alignment improved even for the degree. Alignment improved even for the degree. Alignment improved even for the unseen groups. That seems almost unseen groups. That seems almost unseen groups. That seems almost magical. magical. magical. And some subsequent research has hinted And some subsequent research has hinted And some subsequent research has hinted that what might be really happening here that what might be really happening here that what might be really happening here is that the model itself has a latent is that the model itself has a latent is that the model itself has a latent understanding of these groups. It just understanding of these groups. It just understanding of these groups. It just didn't know how to express it in the didn't know how to express it in the didn't know how to express it in the format of surveys. And if you think format of surveys. And if you think format of surveys. And if you think about it, about it, about it, LLMs aren't used to doing surveys as a LLMs aren't used to doing surveys as a LLMs aren't used to doing surveys as a task. And so they aren't going to be as task. And so they aren't going to be as task. And so they aren't going to be as good as fitting it, especially to a good as fitting it, especially to a good as fitting it, especially to a prompt format they may not have seen prompt format they may not have seen prompt format they may not have seen on the first go-around. But fine-tuning on the first go-around. But fine-tuning on the first go-around. But fine-tuning actually is helping it learn the task or actually is helping it learn the task or actually is helping it learn the task or how to express itself. So, a lesson you how to express itself. So, a lesson you how to express itself. So, a lesson you can kind of take away is that your can kind of take away is that your can kind of take away is that your persona that you're looking for is in persona that you're looking for is in persona that you're looking for is in there. We just need to figure out the there. We just need to figure out the there. We just need to figure out the way to summon it or elicit it.

  12. way to summon it or elicit it. way to summon it or elicit it. And that lesson actually takes us to the And that lesson actually takes us to the And that lesson actually takes us to the third technique, which I want to third technique, which I want to third technique, which I want to highlight to show how sophisticated your highlight to show how sophisticated your highlight to show how sophisticated your techniques can get if you're just using techniques can get if you're just using techniques can get if you're just using so-called prompting alone, but using so-called prompting alone, but using so-called prompting alone, but using careful calibration and thinking. careful calibration and thinking. careful calibration and thinking. So, So, So, in this one, this team did something in this one, this team did something in this one, this team did something very clever. They set up a system prompt very clever. They set up a system prompt very clever. They set up a system prompt that was demographics. They showed a that was demographics. They showed a that was demographics. They showed a product concept. And then they asked how product concept. And then they asked how product concept. And then they asked how likely would you be to purchase this likely would you be to purchase this likely would you be to purchase this product? product? product? And they gave it the same scale from one And they gave it the same scale from one And they gave it the same scale from one to five, five being most likely, one to five, five being most likely, one to five, five being most likely, one being the least likely to purchase, like being the least likely to purchase, like being the least likely to purchase, like you'd expect, kind of your basic naive you'd expect, kind of your basic naive you'd expect, kind of your basic naive prompting pattern. prompting pattern. prompting pattern. And then they said, "Well, you know, And then they said, "Well, you know, And then they said, "Well, you know, hearkening back to that paper, although hearkening back to that paper, although hearkening back to that paper, although I don't know if they were inspired by I don't know if they were inspired by I don't know if they were inspired by it, it, it, they said, 'Well, large language models they said, 'Well, large language models they said, 'Well, large language models aren't used to doing surface, but they aren't used to doing surface, but they aren't used to doing surface, but they are more used to expressing themselves are more used to expressing themselves are more used to expressing themselves in text.' in text.' in text.' So, they said as instead of giving us a So, they said as instead of giving us a So, they said as instead of giving us a one to five rating, give us a set of one to five rating, give us a set of one to five rating, give us a set of text. So, the example here is, 'I'm text. So, the example here is, 'I'm text. So, the example here is, 'I'm somewhat interested. If it works well somewhat interested. If it works well somewhat interested. If it works well and isn't too expensive, I might give it and isn't too expensive, I might give it and isn't too expensive, I might give it a try.' a try.' a try.' And then to map that text to the one And then to map that text to the one And then to map that text to the one through five willingness to pay, they through five willingness to pay, they through five willingness to pay, they had humans write out corresponding text had humans write out corresponding text had humans write out corresponding text for what they would expect. So, if it's for what they would expect. So, if it's for what they would expect. So, if it's a one, "Hell no, I'll never buy that."

  13. a one, "Hell no, I'll never buy that." a one, "Hell no, I'll never buy that." Five, "Absolutely, I'll buy 20." Right? Five, "Absolutely, I'll buy 20." Right? Five, "Absolutely, I'll buy 20." Right? They had them write out examples of each They had them write out examples of each They had them write out examples of each one of the different options, and then one of the different options, and then one of the different options, and then they measured the semantic similarity they measured the semantic similarity they measured the semantic similarity between the text that came out of the between the text that came out of the between the text that came out of the model and those human examples. And that model and those human examples. And that model and those human examples. And that gave them a vector over which they can gave them a vector over which they can gave them a vector over which they can basically measure a probability basically measure a probability basically measure a probability distribution of where this text that distribution of where this text that distribution of where this text that came out of the model lands. So, what I came out of the model lands. So, what I came out of the model lands. So, what I like about this is it's actually a like about this is it's actually a like about this is it's actually a distribution. Kind of feels like, you distribution. Kind of feels like, you distribution. Kind of feels like, you know, human. Some days I might say four, know, human. Some days I might say four, know, human. Some days I might say four, some might Some days I might say five in some might Some days I might say five in some might Some days I might say five in this graph, but rarely would I say one, this graph, but rarely would I say one, this graph, but rarely would I say one, two, or three in this example. two, or three in this example. two, or three in this example. And what they were able to show is that And what they were able to show is that And what they were able to show is that they were not able to only reconstruct they were not able to only reconstruct they were not able to only reconstruct accurate values for willingness to pay. accurate values for willingness to pay. accurate values for willingness to pay. They were able to capture the They were able to capture the They were able to capture the distribution. Because one of the distribution. Because one of the distribution. Because one of the important failure modes we haven't important failure modes we haven't important failure modes we haven't talked about is that LLMs, even when talked about is that LLMs, even when talked about is that LLMs, even when they get the persona averages right, they get the persona averages right, they get the persona averages right, they very often lose the details. The they very often lose the details. The they very often lose the details. The variations get muddled together in the variations get muddled together in the variations get muddled together in the middle. middle. middle. This chart at the bottom, This chart at the bottom, This chart at the bottom, basically that horizontal axis is a basically that horizontal axis is a basically that horizontal axis is a measure of the entire shape similarity. measure of the entire shape similarity. measure of the entire shape similarity. And one means perfectly identical, and And one means perfectly identical, and And one means perfectly identical, and zero means not. And what you can see is zero means not. And what you can see is zero means not. And what you can see is the naive way, in the purple or I guess the naive way, in the purple or I guess the naive way, in the purple or I guess pink uh doesn't do as well as the yellow pink uh doesn't do as well as the yellow pink uh doesn't do as well as the yellow which is up near the top of the range.

  14. which is up near the top of the range. which is up near the top of the range. So that means it really did a good job So that means it really did a good job So that means it really did a good job not only understanding what the ultimate not only understanding what the ultimate not only understanding what the ultimate choice was but how well that choice choice was but how well that choice choice was but how well that choice varied. varied. varied. Okay. Let's talk about how to measure Okay. Let's talk about how to measure Okay. Let's talk about how to measure alignment from a synthetic persona. alignment from a synthetic persona. alignment from a synthetic persona. One of the things that our traditional One of the things that our traditional One of the things that our traditional market research clients are sometimes market research clients are sometimes market research clients are sometimes surprised by and disappointed is that surprised by and disappointed is that surprised by and disappointed is that you cannot use statistical you cannot use statistical you cannot use statistical synthetic personas to boost statistical synthetic personas to boost statistical synthetic personas to boost statistical significance. You can take an significance. You can take an significance. You can take an underrepresented population and get more underrepresented population and get more underrepresented population and get more values out of it but you can't say it's values out of it but you can't say it's values out of it but you can't say it's statistically significant. And to statistically significant. And to statistically significant. And to understand this it helps to go back to understand this it helps to go back to understand this it helps to go back to that weather analogy. If I want to know that weather analogy. If I want to know that weather analogy. If I want to know how much it rains today in San Francisco how much it rains today in San Francisco how much it rains today in San Francisco and I used to live here so I know it and I used to live here so I know it and I used to live here so I know it rains a lot, I'd stick a weather gauge. rains a lot, I'd stick a weather gauge. rains a lot, I'd stick a weather gauge. If I want to know with more certainty, If I want to know with more certainty, If I want to know with more certainty, I'd stick a thousand weather gauges and I'd stick a thousand weather gauges and I'd stick a thousand weather gauges and those would increase the accuracy of my those would increase the accuracy of my those would increase the accuracy of my estimate. estimate. estimate. But if I want to know if it's going to But if I want to know if it's going to But if I want to know if it's going to rain tomorrow, if I take a forecast and rain tomorrow, if I take a forecast and rain tomorrow, if I take a forecast and I rerun it a thousand times without I rerun it a thousand times without I rerun it a thousand times without changing the input, that doesn't change changing the input, that doesn't change changing the input, that doesn't change my certainty of that forecast. It my certainty of that forecast. It my certainty of that forecast. It improves my estimate of what the model improves my estimate of what the model improves my estimate of what the model is telling me but it doesn't make the is telling me but it doesn't make the is telling me but it doesn't make the forecast itself more accurate. And forecast itself more accurate. And forecast itself more accurate. And that's what happens when you basically that's what happens when you basically that's what happens when you basically are rerunning a synthetic persona with are rerunning a synthetic persona with are rerunning a synthetic persona with no changes to input. So the lesson is no changes to input. So the lesson is no changes to input. So the lesson is more synthetic samples aren't actually more synthetic samples aren't actually more synthetic samples aren't actually going to improve your statistical going to improve your statistical going to improve your statistical significance for the most part.

  15. significance for the most part. significance for the most part. So what you need to do is you need to do So what you need to do is you need to do So what you need to do is you need to do what you do with weather forecast. You'd what you do with weather forecast. You'd what you do with weather forecast. You'd basically check against what actually basically check against what actually basically check against what actually happened or in our case what humans happened or in our case what humans happened or in our case what humans actually said. And that's where we're actually said. And that's where we're actually said. And that's where we're going to basically be measuring going to basically be measuring going to basically be measuring distributions of data. Unlike classic distributions of data. Unlike classic distributions of data. Unlike classic e-vals where there's clearly a right and e-vals where there's clearly a right and e-vals where there's clearly a right and wrong and you can score how many were wrong and you can score how many were wrong and you can score how many were right and how many wrong, now we need to right and how many wrong, now we need to right and how many wrong, now we need to measure the data as a comparison of measure the data as a comparison of measure the data as a comparison of distributions. And there are many ways distributions. And there are many ways distributions. And there are many ways for distributions to get wrong. They for distributions to get wrong. They for distributions to get wrong. They could be completely wildly off. They can could be completely wildly off. They can could be completely wildly off. They can as we mentioned get the average right as we mentioned get the average right as we mentioned get the average right but the shape of the distribution wrong. but the shape of the distribution wrong. but the shape of the distribution wrong. And so you're going to need multiple And so you're going to need multiple And so you're going to need multiple metrics to capture how well your model metrics to capture how well your model metrics to capture how well your model is reflecting different personas. Um I is reflecting different personas. Um I is reflecting different personas. Um I recommend using a correlation type recommend using a correlation type recommend using a correlation type metric along with one of these shape metric along with one of these shape metric along with one of these shape type metrics which capture what the type metrics which capture what the type metrics which capture what the underlying shape of the distribution is. underlying shape of the distribution is. underlying shape of the distribution is. The other thing you need to do is The other thing you need to do is The other thing you need to do is estimate the fundamental noise in your estimate the fundamental noise in your estimate the fundamental noise in your ground truth data. That experiment I ground truth data. That experiment I ground truth data. That experiment I talked about in the beginning where they talked about in the beginning where they talked about in the beginning where they got 83% accuracy, got 83% accuracy, got 83% accuracy, the key smart thing they did is they the key smart thing they did is they the key smart thing they did is they took those humans and they brought them took those humans and they brought them took those humans and they brought them back 2 weeks later and they redid the back 2 weeks later and they redid the back 2 weeks later and they redid the battery of surveys and personality tests battery of surveys and personality tests battery of surveys and personality tests and they found that the humans on and they found that the humans on and they found that the humans on average were only 80% consistent to average were only 80% consistent to average were only 80% consistent to themselves.

  16. themselves. themselves. So that sets a noise floor as how So that sets a noise floor as how So that sets a noise floor as how accurate our models could ever get accurate our models could ever get accurate our models could ever get because the humans themselves are because the humans themselves are because the humans themselves are fundamentally noisy. And so the 83% is fundamentally noisy. And so the 83% is fundamentally noisy. And so the 83% is actually normalized against that. actually normalized against that. actually normalized against that. If you can do this and bring your humans If you can do this and bring your humans If you can do this and bring your humans back, that's great. Very often you back, that's great. Very often you back, that's great. Very often you can't. So the way you can kind of can't. So the way you can kind of can't. So the way you can kind of artificially do this is take your ground artificially do this is take your ground artificially do this is take your ground truth human data, break it into two truth human data, break it into two truth human data, break it into two chunks, and then pretend one is chunks, and then pretend one is chunks, and then pretend one is synthetic and one is human, and then synthetic and one is human, and then synthetic and one is human, and then measure the correlation and repeat that measure the correlation and repeat that measure the correlation and repeat that hundreds and thousands of times and hundreds and thousands of times and hundreds and thousands of times and average it, and that'll set kind of a average it, and that'll set kind of a average it, and that'll set kind of a noise floor that your ground truth data noise floor that your ground truth data noise floor that your ground truth data where half of it's synthetic, half of it where half of it's synthetic, half of it where half of it's synthetic, half of it real, could be the level of accuracy you real, could be the level of accuracy you real, could be the level of accuracy you could hope to get. could hope to get. could hope to get. So hopefully by now you have an So hopefully by now you have an So hopefully by now you have an appreciation for why I think weather appreciation for why I think weather appreciation for why I think weather forecasts are the best lens to forecasts are the best lens to forecasts are the best lens to understand synthetic personas. They are understand synthetic personas. They are understand synthetic personas. They are not people, they are forecasts, and we not people, they are forecasts, and we not people, they are forecasts, and we should treat them accordingly. Both should treat them accordingly. Both should treat them accordingly. Both systems are bounded, both systems will systems are bounded, both systems will systems are bounded, both systems will be improving over time, and they're most be improving over time, and they're most be improving over time, and they're most trustworthy when they're validated trustworthy when they're validated trustworthy when they're validated against reality. against reality. against reality. Um Um Um Now synthetic personas are very often Now synthetic personas are very often Now synthetic personas are very often cast in the market against human cast in the market against human cast in the market against human research. And I think that's unfortunate research. And I think that's unfortunate research. And I think that's unfortunate because they're actually complementary because they're actually complementary because they're actually complementary to each other. Let me give you two to each other. Let me give you two to each other. Let me give you two reasons why. One is that reasons why. One is that reasons why. One is that we're entering an era where humans are we're entering an era where humans are we're entering an era where humans are no longer the sole economic actor. Every no longer the sole economic actor. Every no longer the sole economic actor. Every action your human customer is taking in action your human customer is taking in action your human customer is taking in terms of awareness, consideration, or a terms of awareness, consideration, or a terms of awareness, consideration, or a purchase decision to buy is being purchase decision to buy is being purchase decision to buy is being increasingly mediated by AI agents. So, increasingly mediated by AI agents. So, increasingly mediated by AI agents. So, a human-only study is actually not the

  17. a human-only study is actually not the a human-only study is actually not the gold truth. What we really need to gold truth. What we really need to gold truth. What we really need to understand is what does the human plus understand is what does the human plus understand is what does the human plus agent ecosystem look like? And then finally, the alternative to a And then finally, the alternative to a synthetic persona is not human research. synthetic persona is not human research. synthetic persona is not human research. In most cases, it's no research or it's In most cases, it's no research or it's In most cases, it's no research or it's somebody's opinion. What really happens somebody's opinion. What really happens somebody's opinion. What really happens is you've done a survey of humans and is you've done a survey of humans and is you've done a survey of humans and you get a question and if it's in the you get a question and if it's in the you get a question and if it's in the survey, you can just answer it. That's survey, you can just answer it. That's survey, you can just answer it. That's very simple to do. But what typically very simple to do. But what typically very simple to do. But what typically happens is it's 2 months later and happens is it's 2 months later and happens is it's 2 months later and you're like, we need to answer this you're like, we need to answer this you're like, we need to answer this question which we didn't ask. Well, then question which we didn't ask. Well, then question which we didn't ask. Well, then somebody needs to be like, "Uh I think somebody needs to be like, "Uh I think somebody needs to be like, "Uh I think it would be this by extrapolation." it would be this by extrapolation." it would be this by extrapolation." Uh expert plus a synthetic persona is Uh expert plus a synthetic persona is Uh expert plus a synthetic persona is going to give you a better result to going to give you a better result to going to give you a better result to that. So, what we like to tell customers that. So, what we like to tell customers that. So, what we like to tell customers is synthetic extends your human data to is synthetic extends your human data to is synthetic extends your human data to more phases of your development process. more phases of your development process. more phases of your development process. It can go more places your existing It can go more places your existing It can go more places your existing research can't. One of the most exciting research can't. One of the most exciting research can't. One of the most exciting directions is to actually run directions is to actually run directions is to actually run simulations. We didn't get time for simulations. We didn't get time for simulations. We didn't get time for this, but it's called generative this, but it's called generative this, but it's called generative agent-based modeling where we can take agent-based modeling where we can take agent-based modeling where we can take each of these personas and simulate with each of these personas and simulate with each of these personas and simulate with the dynamics and how they'll interface the dynamics and how they'll interface the dynamics and how they'll interface interface and interact with each other.

  18. interface and interact with each other. interface and interact with each other. And ultimately, what this will let you And ultimately, what this will let you And ultimately, what this will let you do is turn your human data into a living do is turn your human data into a living do is turn your human data into a living queryable asset. queryable asset. queryable asset. If you're interested in doing that with If you're interested in doing that with If you're interested in doing that with your data, feel free to reach out to us. your data, feel free to reach out to us. your data, feel free to reach out to us. We help market research and insights We help market research and insights We help market research and insights teams generate and use synthetic teams generate and use synthetic teams generate and use synthetic personas in AI. You can find us on the personas in AI. You can find us on the personas in AI. You can find us on the web at insights sciences.ai web at insights sciences.ai web at insights sciences.ai and my contact information is on the and my contact information is on the and my contact information is on the slide. I hope you have a good slide. I hope you have a good slide. I hope you have a good conference. Thank you. conference. Thank you. conference. Thank you. >> [applause]

Summary

The main theme is the rise of synthetic personas, AI-generated individuals used for market research and product testing. These personas are likened to weather forecasting, enabled by advancements in compute and data. The practical takeaway is to understand both the promise and limitations of synthetic personas, as they, like weather, have a point beyond which accuracy diminishes.

View original episode ↗