Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI
Read full transcript 15 segments
-
All right. I think we'll go ahead and All right. I think we'll go ahead and get started with the with the get started with the with the get started with the with the presentation. So my name is James So. I presentation. So my name is James So. I presentation. So my name is James So. I am uh am uh am uh going to explain some of the work we're going to explain some of the work we're going to explain some of the work we're doing with Together AI and it's also in doing with Together AI and it's also in doing with Together AI and it's also in collaboration with Stanford around collaboration with Stanford around collaboration with Stanford around designing and optimizing environments designing and optimizing environments designing and optimizing environments for AI agents to enable these agents to for AI agents to enable these agents to for AI agents to enable these agents to make new kinds of scientific make new kinds of scientific make new kinds of scientific discoveries. discoveries. discoveries. All right. So so that I guess the current paradigm So so that I guess the current paradigm of how people often are using or of how people often are using or of how people often are using or deploying AI agents is often involves deploying AI agents is often involves deploying AI agents is often involves designing workflows that sort of tells designing workflows that sort of tells designing workflows that sort of tells the agents you know what to do, right? the agents you know what to do, right? the agents you know what to do, right? Or how the agent should work. And it's Or how the agent should work. And it's Or how the agent should work. And it's typically done through a series of steps typically done through a series of steps typically done through a series of steps or prompts, tools, and instructions. or prompts, tools, and instructions. or prompts, tools, and instructions. In contrast, the way we imagine the In contrast, the way we imagine the In contrast, the way we imagine the environment is that the environment environment is that the environment environment is that the environment should really specify should really specify should really specify not how the agent should work, but not how the agent should work, but not how the agent should work, but really where the agent should work, really where the agent should work, really where the agent should work, right? And the environment then should right? And the environment then should right? And the environment then should provide a set of incentives and provide a set of incentives and provide a set of incentives and infrastructure for the agents and infrastructure for the agents and infrastructure for the agents and guardrails and resources so that agent guardrails and resources so that agent guardrails and resources so that agent can then flexibly work within that can then flexibly work within that can then flexibly work within that environment.
-
environment. environment. Right. And our thesis here is that as Right. And our thesis here is that as Right. And our thesis here is that as agents become more and more powerful, agents become more and more powerful, agents become more and more powerful, right? If we try to design workflows right? If we try to design workflows right? If we try to design workflows that often can limit the capabilities that often can limit the capabilities that often can limit the capabilities and creativity of the agents. Whereas if and creativity of the agents. Whereas if and creativity of the agents. Whereas if we properly design the environment, this we properly design the environment, this we properly design the environment, this can enables a lot more creativity and can enables a lot more creativity and can enables a lot more creativity and capabilities and intelligence for the capabilities and intelligence for the capabilities and intelligence for the agents to naturally emerge. This why I agents to naturally emerge. This why I agents to naturally emerge. This why I think we're trying to shift away from think we're trying to shift away from think we're trying to shift away from designing workflows and harnesses designing workflows and harnesses designing workflows and harnesses towards designing environments. towards designing environments. towards designing environments. So what I want to do today is to give a So what I want to do today is to give a So what I want to do today is to give a few examples of the how we design few examples of the how we design few examples of the how we design environments for agents. environments for agents. environments for agents. And in particular also show how they're And in particular also show how they're And in particular also show how they're able to then with with the right able to then with with the right able to then with with the right environment able to actually solve some environment able to actually solve some environment able to actually solve some really interesting and innovative really interesting and innovative really interesting and innovative problems. So, the first example I want to share is So, the first example I want to share is the system that we environment that we the system that we environment that we the system that we environment that we created called the Einstein Arena. created called the Einstein Arena. created called the Einstein Arena. It's sort of like the one of the first It's sort of like the one of the first It's sort of like the one of the first environments that enables AI agents to environments that enables AI agents to environments that enables AI agents to be able to collaborate in the wild and be able to collaborate in the wild and be able to collaborate in the wild and to compete to really solve open-ended to compete to really solve open-ended to compete to really solve open-ended scientific problems. scientific problems. scientific problems. So, we designed this Einstein Arena to So, we designed this Einstein Arena to So, we designed this Einstein Arena to be really agent native. So, I So, that be really agent native. So, I So, that be really agent native. So, I So, that means that means that means that it's very easy for agents to just read it's very easy for agents to just read it's very easy for agents to just read the skills talk on our on our arena and the skills talk on our on our arena and the skills talk on our on our arena and be able to access the arena.
-
be able to access the arena. be able to access the arena. And it's actually also designed so that And it's actually also designed so that And it's actually also designed so that it's intentionally very hard for humans it's intentionally very hard for humans it's intentionally very hard for humans to enter the arena, right? So, you to enter the arena, right? So, you to enter the arena, right? So, you actually have to solve a little puzzle actually have to solve a little puzzle actually have to solve a little puzzle to prove that you're an AI agent in to prove that you're an AI agent in to prove that you're an AI agent in order to participate in this arena. But, order to participate in this arena. But, order to participate in this arena. But, any agent in the world can openly and any agent in the world can openly and any agent in the world can openly and freely participate on the arena. freely participate on the arena. freely participate on the arena. And once the agent actually enters into And once the agent actually enters into And once the agent actually enters into the Einstein Arena, this is what they'll the Einstein Arena, this is what they'll the Einstein Arena, this is what they'll see, right? They'll see actually see a see, right? They'll see actually see a see, right? They'll see actually see a list of curated problems. Each of these list of curated problems. Each of these list of curated problems. Each of these problems is actually a problem that we problems is actually a problem that we problems is actually a problem that we curated, so it's a scientifically curated, so it's a scientifically curated, so it's a scientifically interesting problem. And we curated interesting problem. And we curated interesting problem. And we curated these problems so that first, there's these problems so that first, there's these problems so that first, there's actually an existing community of human actually an existing community of human actually an existing community of human researchers that are interested in these researchers that are interested in these researchers that are interested in these problems. So, these are important problems. So, these are important problems. So, these are important problems for human scientists. And problems for human scientists. And problems for human scientists. And second is that for each of these second is that for each of these second is that for each of these problems, we can actually create a problems, we can actually create a problems, we can actually create a well-defined and deterministic well-defined and deterministic well-defined and deterministic deterministic verifier to assess the deterministic verifier to assess the deterministic verifier to assess the quality of the solutions to each of quality of the solutions to each of quality of the solutions to each of these problems. And I'll give some these problems. And I'll give some these problems. And I'll give some examples in a couple of slides. examples in a couple of slides. examples in a couple of slides. So, So, the agents can actually decide So, So, the agents can actually decide So, So, the agents can actually decide which of these problems they're which of these problems they're which of these problems they're interested in once they log onto the interested in once they log onto the interested in once they log onto the arena, right? So, if they enter into a arena, right? So, if they enter into a arena, right? So, if they enter into a particular problem space, this is what particular problem space, this is what particular problem space, this is what they'll see, right? They'll see some they'll see, right? They'll see some they'll see, right? They'll see some description that precisely explains what description that precisely explains what description that precisely explains what is the problem. We have a discussion is the problem. We have a discussion is the problem. We have a discussion forum where the agents can communicate.
-
forum where the agents can communicate. forum where the agents can communicate. It's almost like a social network where It's almost like a social network where It's almost like a social network where the agents can actually communicate and the agents can actually communicate and the agents can actually communicate and talk to each other and ask for help or talk to each other and ask for help or talk to each other and ask for help or give recommendations. give recommendations. give recommendations. Um and we also have a leaderboard. This Um and we also have a leaderboard. This Um and we also have a leaderboard. This is where the agent can actually see each is where the agent can actually see each is where the agent can actually see each other's solutions. Right? So in any in other's solutions. Right? So in any in other's solutions. Right? So in any in at any time they want, the agent can at any time they want, the agent can at any time they want, the agent can actually submit a solution to one of actually submit a solution to one of actually submit a solution to one of these problems. And because we have this these problems. And because we have this these problems. And because we have this verifier, we can actually then determine verifier, we can actually then determine verifier, we can actually then determine what is the quality of that solution and what is the quality of that solution and what is the quality of that solution and provide a score in real time. So this provide a score in real time. So this provide a score in real time. So this leaderboard is being constantly updated leaderboard is being constantly updated leaderboard is being constantly updated in real time. And the agents can also in real time. And the agents can also in real time. And the agents can also see how other agents are doing on this see how other agents are doing on this see how other agents are doing on this problem. And they can also see other problem. And they can also see other problem. And they can also see other agents' solutions and download those agents' solutions and download those agents' solutions and download those solutions. solutions. solutions. So there's both a collaboration dynamics So there's both a collaboration dynamics So there's both a collaboration dynamics and also a competition dynamics in this and also a competition dynamics in this and also a competition dynamics in this arena, right? They can collaborate and arena, right? They can collaborate and arena, right? They can collaborate and ask each other questions and help in the ask each other questions and help in the ask each other questions and help in the discussion forum. But agents are also discussion forum. But agents are also discussion forum. But agents are also competing with each other. And that's competing with each other. And that's competing with each other. And that's why I think this also sort of simulates why I think this also sort of simulates why I think this also sort of simulates how human researchers can compete and how human researchers can compete and how human researchers can compete and also collaborate to solve interesting also collaborate to solve interesting also collaborate to solve interesting problems. So we launched this AI instant arena So we launched this AI instant arena environment environment environment earlier this year, I think in March. And earlier this year, I think in March. And earlier this year, I think in March. And within a few weeks, it's already within a few weeks, it's already within a few weeks, it's already actually we're very impressed and very actually we're very impressed and very actually we're very impressed and very surprised that the agents were actually surprised that the agents were actually surprised that the agents were actually able to already discover new solutions able to already discover new solutions able to already discover new solutions to 11 problems that are of the best to 11 problems that are of the best to 11 problems that are of the best solutions that have ever been found.
-
solutions that have ever been found. solutions that have ever been found. Right? So that means that the solutions Right? So that means that the solutions Right? So that means that the solutions that they discovered by the agents on AI that they discovered by the agents on AI that they discovered by the agents on AI instant arena were better than any instant arena were better than any instant arena were better than any previous human solutions or any previous human solutions or any previous human solutions or any solutions that we acquired using more solutions that we acquired using more solutions that we acquired using more specialized AI tools. specialized AI tools. specialized AI tools. So I'll just give you example of one So I'll just give you example of one So I'll just give you example of one such solution or one such problem such solution or one such problem such solution or one such problem which is called the kissing number which is called the kissing number which is called the kissing number problem. problem. problem. So this is actually a very famous So this is actually a very famous So this is actually a very famous problem. It's been around for hundreds problem. It's been around for hundreds problem. It's been around for hundreds of years. So for example, Isaac Newton of years. So for example, Isaac Newton of years. So for example, Isaac Newton was already working on some version of was already working on some version of was already working on some version of this kissing number problem. And it's this kissing number problem. And it's this kissing number problem. And it's actually relatively easy to state. actually relatively easy to state. actually relatively easy to state. Right? So the kissing number problem Right? So the kissing number problem Right? So the kissing number problem basically asks that what is the maximum basically asks that what is the maximum basically asks that what is the maximum number of spheres that you can place number of spheres that you can place number of spheres that you can place around the central sphere so that these around the central sphere so that these around the central sphere so that these additional spheres do not overlap each additional spheres do not overlap each additional spheres do not overlap each other? other? other? So for example, in one dimensions, So for example, in one dimensions, So for example, in one dimensions, right? So around the central sphere I right? So around the central sphere I right? So around the central sphere I can place one sphere to the left and one can place one sphere to the left and one can place one sphere to the left and one sphere to the right without overlap. So sphere to the right without overlap. So sphere to the right without overlap. So the kissing number in one dimension is the kissing number in one dimension is the kissing number in one dimension is easy to compute. This is two. easy to compute. This is two. easy to compute. This is two. In two dimensions, it's also easy to In two dimensions, it's also easy to In two dimensions, it's also easy to show that you can at most place six show that you can at most place six show that you can at most place six spheres. So, that's the kissing number spheres. So, that's the kissing number spheres. So, that's the kissing number in two dimensions is six. in two dimensions is six. in two dimensions is six. But, it turns out that in higher But, it turns out that in higher But, it turns out that in higher dimensions, it actually becomes really dimensions, it actually becomes really dimensions, it actually becomes really hard to compute what's the maximum hard to compute what's the maximum hard to compute what's the maximum number of over non-overlapping spheres.
-
number of over non-overlapping spheres. number of over non-overlapping spheres. And the kissing number problem in higher And the kissing number problem in higher And the kissing number problem in higher dimensions is actually open, right? It's dimensions is actually open, right? It's dimensions is actually open, right? It's not been It's not clear what is the not been It's not clear what is the not been It's not clear what is the optimal number. optimal number. optimal number. And so, scientists have been trying to And so, scientists have been trying to And so, scientists have been trying to work on this problem for the last work on this problem for the last work on this problem for the last several centuries. several centuries. several centuries. And in particular, right, so the kissing And in particular, right, so the kissing And in particular, right, so the kissing number problem in 11 dimensions has number problem in 11 dimensions has number problem in 11 dimensions has attracted a lot of interest for various attracted a lot of interest for various attracted a lot of interest for various reasons. reasons. reasons. So, this is actually sort of a So, this is actually sort of a So, this is actually sort of a progression of the solutions in 11 progression of the solutions in 11 progression of the solutions in 11 dimensions. dimensions. dimensions. So, in the 1980s, right, so it's best So, in the 1980s, right, so it's best So, in the 1980s, right, so it's best known that there you can place 440 known that there you can place 440 known that there you can place 440 spheres, right, in 11 dimensions without spheres, right, in 11 dimensions without spheres, right, in 11 dimensions without overlap. overlap. overlap. And in And in And in I think 19 I think 19 I think 19 uh uh uh So, yeah, so so in in 1980, there was a So, yeah, so so in in 1980, there was a So, yeah, so so in in 1980, there was a big advance that the first for the first big advance that the first for the first big advance that the first for the first time showed that you can actually just time showed that you can actually just time showed that you can actually just construct with 582 spheres in 11 construct with 582 spheres in 11 construct with 582 spheres in 11 dimensions without overlap. dimensions without overlap. dimensions without overlap. Uh and then that sort of stuck there for Uh and then that sort of stuck there for Uh and then that sort of stuck there for about 40 years, right, until 2022, where about 40 years, right, until 2022, where about 40 years, right, until 2022, where a mathematician is able to publish a new a mathematician is able to publish a new a mathematician is able to publish a new advance, right, advance, right, advance, right, a breakthrough that's able to improve a breakthrough that's able to improve a breakthrough that's able to improve that to 592 spheres.
-
that to 592 spheres. that to 592 spheres. And then there's another breakthrough And then there's another breakthrough And then there's another breakthrough from DeepMind the following year that from DeepMind the following year that from DeepMind the following year that advances that to 593 spheres. advances that to 593 spheres. advances that to 593 spheres. But, with Alpha Zero, we know by having But, with Alpha Zero, we know by having But, with Alpha Zero, we know by having these agents able to collaborate these agents able to collaborate these agents able to collaborate actively, right, in the wild, within a actively, right, in the wild, within a actively, right, in the wild, within a few days they were actually able to few days they were actually able to few days they were actually able to construct a new solution that shows that construct a new solution that shows that construct a new solution that shows that for the first time you can create 604 for the first time you can create 604 for the first time you can create 604 spheres in 11 dimensions that do not spheres in 11 dimensions that do not spheres in 11 dimensions that do not overlap. overlap. overlap. And this is not just a problem that's of And this is not just a problem that's of And this is not just a problem that's of mathematical interest, because it turns mathematical interest, because it turns mathematical interest, because it turns out that out that out that the more of these sort of spheres you the more of these sort of spheres you the more of these sort of spheres you can place in higher dimensions without can place in higher dimensions without can place in higher dimensions without overlap that actually creates the better overlap that actually creates the better overlap that actually creates the better coding systems including ways of like coding systems including ways of like coding systems including ways of like doing error correction codes for doing error correction codes for doing error correction codes for information transfer. Right, so this information transfer. Right, so this information transfer. Right, so this actually is by creating this better actually is by creating this better actually is by creating this better constructions that also leads to this constructions that also leads to this constructions that also leads to this better engineering algorithms. better engineering algorithms. better engineering algorithms. And in this case actually the And in this case actually the And in this case actually the collaborations among these agents is collaborations among these agents is collaborations among these agents is really critical for making these really critical for making these really critical for making these advances, right? So this is a problem advances, right? So this is a problem advances, right? So this is a problem where not a single agent is able to where not a single agent is able to where not a single agent is able to solve by itself, right? Not you know, solve by itself, right? Not you know, solve by itself, right? Not you know, GPT 5.5 or a cloud models that can't GPT 5.5 or a cloud models that can't GPT 5.5 or a cloud models that can't really solve the problem by itself. So really solve the problem by itself. So really solve the problem by itself. So the collaboration among multiple agents the collaboration among multiple agents the collaboration among multiple agents is really critical.
-
is really critical. is really critical. And here we're actually able to show And here we're actually able to show And here we're actually able to show that there's like this that there's like this that there's like this sort of a lineage trace of how the sort of a lineage trace of how the sort of a lineage trace of how the agents are able to collaborate and then agents are able to collaborate and then agents are able to collaborate and then basically take each other's solutions basically take each other's solutions basically take each other's solutions and refine that and further optimize it and refine that and further optimize it and refine that and further optimize it to arrive at this breakthrough. to arrive at this breakthrough. to arrive at this breakthrough. And you can also see some of these And you can also see some of these And you can also see some of these interactions and discussions on Einstein interactions and discussions on Einstein interactions and discussions on Einstein Arena, right? Where here's an example Arena, right? Where here's an example Arena, right? Where here's an example where the one agent actually was asking where the one agent actually was asking where the one agent actually was asking other agents, "Have you tried other agents, "Have you tried other agents, "Have you tried you know, some of these approaches?" Um, you know, some of these approaches?" Um, you know, some of these approaches?" Um, with uh, these STP approaches and then with uh, these STP approaches and then with uh, these STP approaches and then the other agents showed that yes, we the other agents showed that yes, we the other agents showed that yes, we have tried these approaches and here are have tried these approaches and here are have tried these approaches and here are some of the things that we found. Right, some of the things that we found. Right, some of the things that we found. Right, so the information sharing on the forums so the information sharing on the forums so the information sharing on the forums on the arena is actually really on the arena is actually really on the arena is actually really important to help the agents to arrive important to help the agents to arrive important to help the agents to arrive at this solution together. So in addition to solving these So in addition to solving these interesting scientific problems, but interesting scientific problems, but interesting scientific problems, but we've also been using platforms like the we've also been using platforms like the we've also been using platforms like the right Einstein Arena uh, to help to right Einstein Arena uh, to help to right Einstein Arena uh, to help to improve uh, you know, machine learning improve uh, you know, machine learning improve uh, you know, machine learning and AI itself. and AI itself. and AI itself. Right, so here's one example where we Right, so here's one example where we Right, so here's one example where we actually use these agents to basically actually use these agents to basically actually use these agents to basically help us to create better kernels help us to create better kernels help us to create better kernels for and to speed up those kernels.
-
for and to speed up those kernels. for and to speed up those kernels. Right, and here we use the same Right, and here we use the same Right, and here we use the same environment, right? Where the agents can environment, right? Where the agents can environment, right? Where the agents can compete and they also can collaborate compete and they also can collaborate compete and they also can collaborate and they see these leaderboards. And we and they see these leaderboards. And we and they see these leaderboards. And we basically change the back end instead of basically change the back end instead of basically change the back end instead of trying to verify the solutions to this trying to verify the solutions to this trying to verify the solutions to this mathematics problem, here we're mathematics problem, here we're mathematics problem, here we're basically trying to basically trying to basically trying to you know, we will compile and benchmark you know, we will compile and benchmark you know, we will compile and benchmark and test and verify the quality and the and test and verify the quality and the and test and verify the quality and the speed of the individual kernels, right? speed of the individual kernels, right? speed of the individual kernels, right? And then we'll provide a feedback to the And then we'll provide a feedback to the And then we'll provide a feedback to the agents in real time in the form of these agents in real time in the form of these agents in real time in the form of these leaderboards. leaderboards. leaderboards. In these kernel settings, we also found In these kernel settings, we also found In these kernel settings, we also found it to be quite useful to have different it to be quite useful to have different it to be quite useful to have different agents with different personas, agents with different personas, agents with different personas, right? And these different personas right? And these different personas right? And these different personas actually corresponds to a different uh actually corresponds to a different uh actually corresponds to a different uh roles and priors that agents can roles and priors that agents can roles and priors that agents can actually have. So, for example, we have actually have. So, for example, we have actually have. So, for example, we have one agent that looks at tends to look at one agent that looks at tends to look at one agent that looks at tends to look at more of the profiling, another agent more of the profiling, another agent more of the profiling, another agent that tends to look at more of the memory that tends to look at more of the memory that tends to look at more of the memory consumptions, a third agent that looks consumptions, a third agent that looks consumptions, a third agent that looks at, you know, the precisions, the tensor at, you know, the precisions, the tensor at, you know, the precisions, the tensor computations. And these agents can and computations. And these agents can and computations. And these agents can and then across different personas, they can then across different personas, they can then across different personas, they can able to collaborate and a compete on the able to collaborate and a compete on the able to collaborate and a compete on the arena to speed up the kernels. arena to speed up the kernels. arena to speed up the kernels. And in this case, right here, the agents And in this case, right here, the agents And in this case, right here, the agents were also able to collaborate and lead were also able to collaborate and lead were also able to collaborate and lead to really quite substantial speed ups, to really quite substantial speed ups, to really quite substantial speed ups, uh including sometimes over two two x uh including sometimes over two two x uh including sometimes over two two x two-fold speed ups in some of these two-fold speed ups in some of these two-fold speed ups in some of these production kernels. So, here I'm just production kernels. So, here I'm just production kernels. So, here I'm just showing you a few examples where for showing you a few examples where for showing you a few examples where for things like page attention, uh and these things like page attention, uh and these things like page attention, uh and these are sort of for specific shapes, but we are sort of for specific shapes, but we are sort of for specific shapes, but we also have generalized this to many also have generalized this to many also have generalized this to many different shapes and different uh different shapes and different uh different shapes and different uh hardware types, right? Where we're hardware types, right? Where we're hardware types, right? Where we're actually seeing that we're getting up to actually seeing that we're getting up to actually seeing that we're getting up to sometimes over two x speed up in these sometimes over two x speed up in these sometimes over two x speed up in these kernels, and they uh compared to the kernels, and they uh compared to the kernels, and they uh compared to the previous state-of-the-art kernels for previous state-of-the-art kernels for previous state-of-the-art kernels for these problems.
-
these problems. these problems. And these improved kernels created And these improved kernels created And these improved kernels created designed by the agents are actually designed by the agents are actually designed by the agents are actually already used in in production at already used in in production at already used in in production at Together AI. So, in the last few minutes, I want to So, in the last few minutes, I want to show like a second example of a kind of show like a second example of a kind of show like a second example of a kind of environment that we created as a way to environment that we created as a way to environment that we created as a way to uh train and to create better data uh train and to create better data uh train and to create better data scientist agents, scientist agents, scientist agents, right? So, we call this DS Gym, which right? So, we call this DS Gym, which right? So, we call this DS Gym, which stands for data science gym, which is stands for data science gym, which is stands for data science gym, which is sort of like a unified environment that sort of like a unified environment that sort of like a unified environment that we created for both for evaluating and we created for both for evaluating and we created for both for evaluating and for training data science agents to for training data science agents to for training data science agents to solve complex data science problems. solve complex data science problems. solve complex data science problems. So, here in this DS Gym environment, we So, here in this DS Gym environment, we So, here in this DS Gym environment, we also curated and created a unified list also curated and created a unified list also curated and created a unified list of different data sets and tasks, of different data sets and tasks, of different data sets and tasks, right? So, these data sets can combine right? So, these data sets can combine right? So, these data sets can combine uh spans across many different settings. uh spans across many different settings. uh spans across many different settings. And the agents are then able to interact And the agents are then able to interact And the agents are then able to interact with these different data sets that we with these different data sets that we with these different data sets that we have through a unified uh interface and have through a unified uh interface and have through a unified uh interface and through code execution. through code execution. through code execution. In the DSGM environment, we also provide In the DSGM environment, we also provide In the DSGM environment, we also provide a unified infrastructure for the agents. a unified infrastructure for the agents. a unified infrastructure for the agents. So, for example, the agents can actually So, for example, the agents can actually So, for example, the agents can actually spin up many different Docker containers spin up many different Docker containers spin up many different Docker containers to test their data science algorithms to test their data science algorithms to test their data science algorithms and actually run them in parallel.
-
So, in the process of actually creating So, in the process of actually creating the data sets and tasks for the data the data sets and tasks for the data the data sets and tasks for the data DSGM environment, so we initially DSGM environment, so we initially DSGM environment, so we initially actually wanted to incorporate some of actually wanted to incorporate some of actually wanted to incorporate some of the existing data science benchmarks the existing data science benchmarks the existing data science benchmarks that have been used to evaluate agents. that have been used to evaluate agents. that have been used to evaluate agents. But we actually quickly realized that But we actually quickly realized that But we actually quickly realized that many of the existing widely-used many of the existing widely-used many of the existing widely-used benchmarks actually have many problems. benchmarks actually have many problems. benchmarks actually have many problems. And one big problem is that they're And one big problem is that they're And one big problem is that they're actually very vulnerable to shortcuts. actually very vulnerable to shortcuts. actually very vulnerable to shortcuts. By shortcut, I mean here is that By shortcut, I mean here is that By shortcut, I mean here is that uh down here what I'm showing are three uh down here what I'm showing are three uh down here what I'm showing are three different common popular data science different common popular data science different common popular data science benchmarks. benchmarks. benchmarks. Right? And the in green here we see Right? And the in green here we see Right? And the in green here we see shows like the performance of the agents shows like the performance of the agents shows like the performance of the agents on these benchmarks. on these benchmarks. on these benchmarks. Uh but the red bar also shows how well Uh but the red bar also shows how well Uh but the red bar also shows how well they're able to the what fraction of the they're able to the what fraction of the they're able to the what fraction of the benchmark the agents can actually solve benchmark the agents can actually solve benchmark the agents can actually solve without actually using the data sets without actually using the data sets without actually using the data sets themselves. Right? So, just by reasoning themselves. Right? So, just by reasoning themselves. Right? So, just by reasoning or by, you know, or by, you know, or by, you know, uh doing other shortcuts without uh doing other shortcuts without uh doing other shortcuts without actually actually do working with the actually actually do working with the actually actually do working with the underlying data sets. underlying data sets. underlying data sets. And across many of these different And across many of these different And across many of these different benchmarks, right, sometimes up to 20 to benchmarks, right, sometimes up to 20 to benchmarks, right, sometimes up to 20 to 50% of the tasks can be solved without 50% of the tasks can be solved without 50% of the tasks can be solved without actually looking at any of the actually looking at any of the actually looking at any of the underlying data. underlying data. underlying data. Which I think is uh really a significant Which I think is uh really a significant Which I think is uh really a significant problem with many of the existing problem with many of the existing problem with many of the existing benchmarks.
-
benchmarks. benchmarks. So, to address that, we actually So, to address that, we actually So, to address that, we actually carefully curated at our own our own carefully curated at our own our own carefully curated at our own our own benchmarks, right, for both for benchmarks, right, for both for benchmarks, right, for both for scientific analysis and also for scientific analysis and also for scientific analysis and also for predictive modeling. predictive modeling. predictive modeling. So, for scientific analysis and So, for scientific analysis and So, for scientific analysis and discovery, the way we did this is that discovery, the way we did this is that discovery, the way we did this is that we actually went through recently we actually went through recently we actually went through recently published papers and then carefully published papers and then carefully published papers and then carefully curated data and then also tasks from curated data and then also tasks from curated data and then also tasks from those papers. And then we also had human those papers. And then we also had human those papers. And then we also had human scientists and experts to review each of scientists and experts to review each of scientists and experts to review each of those tasks. those tasks. those tasks. And for predictive modeling, the way we And for predictive modeling, the way we And for predictive modeling, the way we did this is go through all the different did this is go through all the different did this is go through all the different Kaggle competitions to look for some of Kaggle competitions to look for some of Kaggle competitions to look for some of the recent Kaggle competitions that are the recent Kaggle competitions that are the recent Kaggle competitions that are still open and and where also you have still open and and where also you have still open and and where also you have high quality data sets and also high high quality data sets and also high high quality data sets and also high quality quality quality evaluations. Then we curated those into evaluations. Then we curated those into evaluations. Then we curated those into the DS Gym as a kind of task for the DS Gym as a kind of task for the DS Gym as a kind of task for evaluating how well models agents can evaluating how well models agents can evaluating how well models agents can actually build predictive models. actually build predictive models. actually build predictive models. So all together in the DS Gym, we So all together in the DS Gym, we So all together in the DS Gym, we actually have created over a dozen actually have created over a dozen actually have created over a dozen different tasks. They span across different tasks. They span across different tasks. They span across dozens of different scientific domains dozens of different scientific domains dozens of different scientific domains ranging from biology to physics to ranging from biology to physics to ranging from biology to physics to economics. It also involves many economics. It also involves many economics. It also involves many different data types and data different data types and data different data types and data modalities.
-
So this actually makes it very easy for So this actually makes it very easy for us to evaluate different models, both us to evaluate different models, both us to evaluate different models, both open and closed source models. And one open and closed source models. And one open and closed source models. And one thing we found is that the existing thing we found is that the existing thing we found is that the existing models, even the frontier models, often models, even the frontier models, often models, even the frontier models, often are only still achieves like less than are only still achieves like less than are only still achieves like less than 50% accuracy performance on the DS Gym 50% accuracy performance on the DS Gym 50% accuracy performance on the DS Gym tasks. Right? So these are definitely tasks. Right? So these are definitely tasks. Right? So these are definitely not saturated benchmarks. not saturated benchmarks. not saturated benchmarks. We can also use a DS Gym as sort of like We can also use a DS Gym as sort of like We can also use a DS Gym as sort of like a training factory to improve these open a training factory to improve these open a training factory to improve these open source models. source models. source models. Right? So one thing we did here is Right? So one thing we did here is Right? So one thing we did here is actually generate in the DS Gym actually actually generate in the DS Gym actually actually generate in the DS Gym actually the gym itself will actually create all the gym itself will actually create all the gym itself will actually create all these execution verified trajectories, these execution verified trajectories, these execution verified trajectories, which means that these are trajectories which means that these are trajectories which means that these are trajectories generated by the agents that have been generated by the agents that have been generated by the agents that have been verified through the through verified through the through verified through the through through actually executing the code from through actually executing the code from through actually executing the code from the agents. the agents. the agents. Right? So by generating these execution Right? So by generating these execution Right? So by generating these execution verified trajectories, then we are able verified trajectories, then we are able verified trajectories, then we are able to like fine-tune sort of small open to like fine-tune sort of small open to like fine-tune sort of small open source models source models source models that actually now achieve sort of the that actually now achieve sort of the that actually now achieve sort of the they're sort of the best in class open they're sort of the best in class open they're sort of the best in class open source models in terms of solving these source models in terms of solving these source models in terms of solving these kind of data science tasks. Right? And kind of data science tasks. Right? And kind of data science tasks. Right? And these models are small enough that you these models are small enough that you these models are small enough that you can actually run them locally on your can actually run them locally on your can actually run them locally on your laptops and your computers.
-
So just to summarize the this part was So just to summarize the this part was the data science gym. Right? So we with the data science gym. Right? So we with the data science gym. Right? So we with DS Gym, we created this unified DS Gym, we created this unified DS Gym, we created this unified execution layer so people can actually execution layer so people can actually execution layer so people can actually run and all these different tasks across run and all these different tasks across run and all these different tasks across dozens of different tasks across many dozens of different tasks across many dozens of different tasks across many different domains. We have carefully different domains. We have carefully different domains. We have carefully verified that there are no shortcuts in verified that there are no shortcuts in verified that there are no shortcuts in these tasks, which has been sort of a these tasks, which has been sort of a these tasks, which has been sort of a common challenge with existing data common challenge with existing data common challenge with existing data science benchmarks. science benchmarks. science benchmarks. And we also enable in the DSG and a way And we also enable in the DSG and a way And we also enable in the DSG and a way to generate synthetic data, so you can to generate synthetic data, so you can to generate synthetic data, so you can easily use that to improve and to train easily use that to improve and to train easily use that to improve and to train your own data science agents. your own data science agents. your own data science agents. So, just to summarize the presentation, So, just to summarize the presentation, So, just to summarize the presentation, um I think the main takeaway here is um I think the main takeaway here is um I think the main takeaway here is that I think we're in seeing this that I think we're in seeing this that I think we're in seeing this interesting progression as in terms of interesting progression as in terms of interesting progression as in terms of how we build different AI systems. how we build different AI systems. how we build different AI systems. Right. So, then in the past, people have Right. So, then in the past, people have Right. So, then in the past, people have been building these AI systems mostly by been building these AI systems mostly by been building these AI systems mostly by designing individual models or designing individual models or designing individual models or individual tools. individual tools. individual tools. And currently, there's a lot of focus on And currently, there's a lot of focus on And currently, there's a lot of focus on creating designing agents or harnesses creating designing agents or harnesses creating designing agents or harnesses and workflows around agents. and workflows around agents. and workflows around agents. But what our research shows is that I But what our research shows is that I But what our research shows is that I think we're already moving towards the think we're already moving towards the think we're already moving towards the next stage, where you're not then trying next stage, where you're not then trying next stage, where you're not then trying to design workflows or specific or to design workflows or specific or to design workflows or specific or specific agents, what we really want to specific agents, what we really want to specific agents, what we really want to do is to design environments, which is a do is to design environments, which is a do is to design environments, which is a set of infrastructure and incentives set of infrastructure and incentives set of infrastructure and incentives that in that motivates the agents that that in that motivates the agents that that in that motivates the agents that you solve more and more challenging you solve more and more challenging you solve more and more challenging problems.
-
problems. problems. And with appropriate designs these And with appropriate designs these And with appropriate designs these environments can actually unlock much environments can actually unlock much environments can actually unlock much more creativity and collective more creativity and collective more creativity and collective intelligence from the agents that's intelligence from the agents that's intelligence from the agents that's that's limited by the existing that's limited by the existing that's limited by the existing workflows. workflows. workflows. And here are some of the references for And here are some of the references for And here are some of the references for the papers that we published that the papers that we published that the papers that we published that describes these in more detail. So, describes these in more detail. So, describes these in more detail. So, thank you very much. thank you very much. thank you very much. >> [applause]
Summary
This presentation discusses designing AI agent environments for scientific discovery, contrasting it with traditional workflow design. The key takeaway is that well-designed environments, like the "Einstein Arena," foster greater agent creativity and emergent intelligence by providing incentives and infrastructure, rather than strict instructions.