Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
Read full transcript 17 segments
-
Uh so this talk is called scaling to Uh so this talk is called scaling to long horizons. My name is Ross. Uh I'm long horizons. My name is Ross. Uh I'm long horizons. My name is Ross. Uh I'm the CEO of GR. We're a London-based the CEO of GR. We're a London-based the CEO of GR. We're a London-based reinforcement learning company. Uh reinforcement learning company. Uh reinforcement learning company. Uh before GR, I was the reasoning lead at before GR, I was the reasoning lead at before GR, I was the reasoning lead at Meta AI working on Llamas, uh Galactica, Meta AI working on Llamas, uh Galactica, Meta AI working on Llamas, uh Galactica, lots of other models back in the day. lots of other models back in the day. lots of other models back in the day. I'm joined by Chengxi, uh co-founder and I'm joined by Chengxi, uh co-founder and I'm joined by Chengxi, uh co-founder and president of GR. Uh and yeah, hit today president of GR. Uh and yeah, hit today president of GR. Uh and yeah, hit today we're going to talk about algorithms, we're going to talk about algorithms, we're going to talk about algorithms, environments, compute, all the things environments, compute, all the things environments, compute, all the things you need to do to get agents scaling uh you need to do to get agents scaling uh you need to do to get agents scaling uh to kind of long with tasks. to kind of long with tasks. to kind of long with tasks. So we're going to have two parts of this So we're going to have two parts of this So we're going to have two parts of this talk today. I'm going to first of all talk today. I'm going to first of all talk today. I'm going to first of all start with a personal perspective about, start with a personal perspective about, start with a personal perspective about, you know, the early days, the golden age you know, the early days, the golden age you know, the early days, the golden age of language modeling in between like of language modeling in between like of language modeling in between like maybe 2020 and 2023. maybe 2020 and 2023. maybe 2020 and 2023. Uh I'll talk about, like I said, all Uh I'll talk about, like I said, all Uh I'll talk about, like I said, all these models and some of our early these models and some of our early these models and some of our early reinforcement learning efforts for LLMs. reinforcement learning efforts for LLMs. reinforcement learning efforts for LLMs. And then Chengxi is going to talk about, And then Chengxi is going to talk about, And then Chengxi is going to talk about, you know, what's ahead, you know, what you know, what's ahead, you know, what you know, what's ahead, you know, what the next frontiers. And yeah, that's the next frontiers. And yeah, that's the next frontiers. And yeah, that's going to be a really talk with a lot of going to be a really talk with a lot of going to be a really talk with a lot of alpha, so I'd encourage you to stick alpha, so I'd encourage you to stick alpha, so I'd encourage you to stick around for that. around for that. around for that. So the journey so far. So my journey So the journey so far. So my journey So the journey so far. So my journey started here. Um so this was the Papers started here. Um so this was the Papers started here. Um so this was the Papers With Code team. I'm sure many of you With Code team. I'm sure many of you With Code team. I'm sure many of you used Papers With Code back in the day.
-
used Papers With Code back in the day. used Papers With Code back in the day. So we were a London-based startup 2019. So we were a London-based startup 2019. So we were a London-based startup 2019. Uh we were acquired by Meta later that Uh we were acquired by Meta later that Uh we were acquired by Meta later that year. And then we had a crazy transition year. And then we had a crazy transition year. And then we had a crazy transition within Meta to do research. Um so we within Meta to do research. Um so we within Meta to do research. Um so we did, like I said, Galactica. Then after did, like I said, Galactica. Then after did, like I said, Galactica. Then after ChatGPT came out, we started the ChatGPT came out, we started the ChatGPT came out, we started the post-training for Llama 2, Llama 3. So post-training for Llama 2, Llama 3. So post-training for Llama 2, Llama 3. So all the great work you uh saw there was all the great work you uh saw there was all the great work you uh saw there was folks in this room. And lots of other folks in this room. And lots of other folks in this room. And lots of other interesting stuff that never got interesting stuff that never got interesting stuff that never got published as well. Uh reasoning Llama published as well. Uh reasoning Llama published as well. Uh reasoning Llama and lots of other things. Um and lots of other things. Um and lots of other things. Um so yeah, this this small team I'd like so yeah, this this small team I'd like so yeah, this this small team I'd like to think like the open weight kind of to think like the open weight kind of to think like the open weight kind of revolution started in this room. And you revolution started in this room. And you revolution started in this room. And you know, know, know, it really like hit home this idea to me it really like hit home this idea to me it really like hit home this idea to me that kind of small focused teams, like that kind of small focused teams, like that kind of small focused teams, like even in like the age of scaling, can do even in like the age of scaling, can do even in like the age of scaling, can do amazing things if people are aligned. amazing things if people are aligned. amazing things if people are aligned. Now for me, uh things got particularly Now for me, uh things got particularly Now for me, uh things got particularly crazy in 2022. Uh so let me tell you a crazy in 2022. Uh so let me tell you a crazy in 2022. Uh so let me tell you a story. Um the media perception is that story. Um the media perception is that story. Um the media perception is that ChatGPT came out of nowhere, you know, ChatGPT came out of nowhere, you know, ChatGPT came out of nowhere, you know, shocked the world, and that's how kind shocked the world, and that's how kind shocked the world, and that's how kind of the modern AI wave started. But, you of the modern AI wave started. But, you of the modern AI wave started. But, you know, I have a different personal know, I have a different personal know, I have a different personal perspective on this because 2 weeks perspective on this because 2 weeks perspective on this because 2 weeks before ChatGPT came along, there was before ChatGPT came along, there was before ChatGPT came along, there was another language model called Galactica.
-
another language model called Galactica. another language model called Galactica. So let's talk about Galactica. Galactica So let's talk about Galactica. Galactica So let's talk about Galactica. Galactica and, you know, ChatGPT, you know, they and, you know, ChatGPT, you know, they and, you know, ChatGPT, you know, they were both, you know, in some respects were both, you know, in some respects were both, you know, in some respects quite similar. They're both based on quite similar. They're both based on quite similar. They're both based on pretty good base models. Galactica pretty good base models. Galactica pretty good base models. Galactica itself was a base model, and then itself was a base model, and then itself was a base model, and then ChatGPT was based on GPT-3.5. ChatGPT was based on GPT-3.5. ChatGPT was based on GPT-3.5. But there was a clear difference in But there was a clear difference in But there was a clear difference in outcomes. So Galactica at the time outcomes. So Galactica at the time outcomes. So Galactica at the time shipped with this slight base model shipped with this slight base model shipped with this slight base model demo. And as you guys know now, like demo. And as you guys know now, like demo. And as you guys know now, like base models, they come with a lot of base models, they come with a lot of base models, they come with a lot of quirks. You know, they hallucinate. You quirks. You know, they hallucinate. You quirks. You know, they hallucinate. You prompt them to do, you know, silly prompt them to do, you know, silly prompt them to do, you know, silly things, they would do silly things. things, they would do silly things. things, they would do silly things. Whereas ChatGPT wasn't just a base Whereas ChatGPT wasn't just a base Whereas ChatGPT wasn't just a base model, but had this like crucial model, but had this like crucial model, but had this like crucial reinforcement learning from human reinforcement learning from human reinforcement learning from human feedback pipeline. And this was the key feedback pipeline. And this was the key feedback pipeline. And this was the key thing that made LLMs like really thing that made LLMs like really thing that made LLMs like really products for the first time. So I like products for the first time. So I like products for the first time. So I like to think in a weird kind of way, this is to think in a weird kind of way, this is to think in a weird kind of way, this is like the first like kind of natural like the first like kind of natural like the first like kind of natural experiment showing you that kind of RL experiment showing you that kind of RL experiment showing you that kind of RL like provides value, right? Uh like provides value, right? Uh like provides value, right? Uh and to my misfortune, it was like a very and to my misfortune, it was like a very and to my misfortune, it was like a very personal like kind of a natural personal like kind of a natural personal like kind of a natural experiment and you know, Galactica blew experiment and you know, Galactica blew experiment and you know, Galactica blew up. Uh but that's like a good like kind up. Uh but that's like a good like kind up. Uh but that's like a good like kind of lesson there. A good base model is of lesson there. A good base model is of lesson there. A good base model is not enough. So I took that lesson quite not enough. So I took that lesson quite not enough. So I took that lesson quite early on. early on. early on. So like I said, RLHF made LLMs products.
-
So like I said, RLHF made LLMs products. So like I said, RLHF made LLMs products. They were the thing that kind of made They were the thing that kind of made They were the thing that kind of made LLMs cross the Rubicon into something LLMs cross the Rubicon into something LLMs cross the Rubicon into something that wasn't just a toy, but used by now that wasn't just a toy, but used by now that wasn't just a toy, but used by now billions of people. But you didn't have billions of people. But you didn't have billions of people. But you didn't have to wait until ChatGPT to see this. Like to wait until ChatGPT to see this. Like to wait until ChatGPT to see this. Like even at the time, like InstructGPT in even at the time, like InstructGPT in even at the time, like InstructGPT in 2022 had these pretty stunning results. 2022 had these pretty stunning results. 2022 had these pretty stunning results. Like a 1 billion parameter model with Like a 1 billion parameter model with Like a 1 billion parameter model with RLHF was outperforming 175 billion RLHF was outperforming 175 billion RLHF was outperforming 175 billion models. So two orders of magnitude fewer models. So two orders of magnitude fewer models. So two orders of magnitude fewer parameters, but getting better results. parameters, but getting better results. parameters, but getting better results. So that was astonishing. So if you were So that was astonishing. So if you were So that was astonishing. So if you were paying attention closely, you know, paying attention closely, you know, paying attention closely, you know, maybe we should have been as well, but maybe we should have been as well, but maybe we should have been as well, but we were focused on a, you know, bloody we were focused on a, you know, bloody we were focused on a, you know, bloody base model, which is hard work in 2022. base model, which is hard work in 2022. base model, which is hard work in 2022. But that shows you how you know But that shows you how you know But that shows you how you know important even basic RL is. important even basic RL is. important even basic RL is. And the Galactica demo itself, I mean, And the Galactica demo itself, I mean, And the Galactica demo itself, I mean, it set off a storm. So, ancient history it set off a storm. So, ancient history it set off a storm. So, ancient history now, but we put out a demo. We let now, but we put out a demo. We let now, but we put out a demo. We let people play around with it um because people play around with it um because people play around with it um because it's kind of cool. And at the time it's kind of cool. And at the time it's kind of cool. And at the time people got scared. So, it was like, you people got scared. So, it was like, you people got scared. So, it was like, you know, like I said, you prompt it on like know, like I said, you prompt it on like know, like I said, you prompt it on like a research paper on Dyson spheres or a a research paper on Dyson spheres or a a research paper on Dyson spheres or a report on the benefits of eating crushed report on the benefits of eating crushed report on the benefits of eating crushed glass and people are like, "Oh my god." glass and people are like, "Oh my god." glass and people are like, "Oh my god." Um so, that was the state of things in Um so, that was the state of things in Um so, that was the state of things in 2022. And yeah, I'll be honest, the meta 2022. And yeah, I'll be honest, the meta 2022. And yeah, I'll be honest, the meta association didn't you know help us association didn't you know help us association didn't you know help us either. Um and yeah, the tragic story in either. Um and yeah, the tragic story in either. Um and yeah, the tragic story in a way was a lot of the, you know, novel a way was a lot of the, you know, novel a way was a lot of the, you know, novel work was maybe overshadowed.
-
work was maybe overshadowed. work was maybe overshadowed. But you know, the paradox of this whole But you know, the paradox of this whole But you know, the paradox of this whole thing is that at the time Galactica was thing is that at the time Galactica was thing is that at the time Galactica was actually a bloody good model. Um like it actually a bloody good model. Um like it actually a bloody good model. Um like it outperformed Palm, Chinchilla, GPT-3.5 outperformed Palm, Chinchilla, GPT-3.5 outperformed Palm, Chinchilla, GPT-3.5 with a lot less compute in scientific with a lot less compute in scientific with a lot less compute in scientific domains. It was state-of-the-art. So domains. It was state-of-the-art. So domains. It was state-of-the-art. So again, that also shows you how powerful again, that also shows you how powerful again, that also shows you how powerful RL is. You can have a sota base model, RL is. You can have a sota base model, RL is. You can have a sota base model, but that is not enough. but that is not enough. but that is not enough. Um so, here you see on like a math who's Um so, here you see on like a math who's Um so, here you see on like a math who's kind of beating Chinchilla. kind of beating Chinchilla. kind of beating Chinchilla. Uh kind of latex equations, you know, Uh kind of latex equations, you know, Uh kind of latex equations, you know, science, you know, it was getting around science, you know, it was getting around science, you know, it was getting around 68% compared to GPT-3.5, 49%. So, 68% compared to GPT-3.5, 49%. So, 68% compared to GPT-3.5, 49%. So, crushing there. crushing there. crushing there. And chain of thought as well. So, you And chain of thought as well. So, you And chain of thought as well. So, you know, a Palm at the time which was a know, a Palm at the time which was a know, a Palm at the time which was a Google Brain model, 540 billion. You Google Brain model, 540 billion. You Google Brain model, 540 billion. You know, that's 30 billion in Galactica was know, that's 30 billion in Galactica was know, that's 30 billion in Galactica was getting like 36 versus 19%. So, double getting like 36 versus 19%. So, double getting like 36 versus 19%. So, double the performance, order of magnitude less the performance, order of magnitude less the performance, order of magnitude less results. So, that again reinforces base results. So, that again reinforces base results. So, that again reinforces base good base models not enough. good base models not enough. good base models not enough. But it introduced some key ideas which I But it introduced some key ideas which I But it introduced some key ideas which I think are very important. I mean, think are very important. I mean, think are very important. I mean, Galactica was the first LM to really Galactica was the first LM to really Galactica was the first LM to really crack data efficiency. 105 billion uh crack data efficiency. 105 billion uh crack data efficiency. 105 billion uh token corpus compared to you know, token corpus compared to you know, token corpus compared to you know, trillion uh tokens in Chinchilla. And it trillion uh tokens in Chinchilla. And it trillion uh tokens in Chinchilla. And it was a really contrarian at the time cuz was a really contrarian at the time cuz was a really contrarian at the time cuz you know, at the time everyone was like, you know, at the time everyone was like, you know, at the time everyone was like, "Okay, we just need more tokens." And "Okay, we just need more tokens." And "Okay, we just need more tokens." And Galactica said, "No. High quality, you Galactica said, "No. High quality, you Galactica said, "No. High quality, you know, curated data sets really matter."
-
know, curated data sets really matter." know, curated data sets really matter." And that was a real driver of those And that was a real driver of those And that was a real driver of those results you just saw. results you just saw. results you just saw. And it was also like the first major LM And it was also like the first major LM And it was also like the first major LM to really crack multi-epoch training. It to really crack multi-epoch training. It to really crack multi-epoch training. It sounds ridiculous now, but at the time sounds ridiculous now, but at the time sounds ridiculous now, but at the time the consensus was you don't do more than the consensus was you don't do more than the consensus was you don't do more than an epoch. Um but this kind of rule of an epoch. Um but this kind of rule of an epoch. Um but this kind of rule of thumb, you may have heard of it like thumb, you may have heard of it like thumb, you may have heard of it like four uh epochs repeated data, that was four uh epochs repeated data, that was four uh epochs repeated data, that was formalized later. but Galactica was the formalized later. but Galactica was the formalized later. but Galactica was the first like real empirical result for first like real empirical result for first like real empirical result for that. that. that. Now, perhaps more importantly, there's Now, perhaps more importantly, there's Now, perhaps more importantly, there's this idea of thinking tokens. And some this idea of thinking tokens. And some this idea of thinking tokens. And some of you might remember this, but it was of you might remember this, but it was of you might remember this, but it was like really quite buried within the like really quite buried within the like really quite buried within the paper. So, around this time there were paper. So, around this time there were paper. So, around this time there were like different ideas for reasoning. like different ideas for reasoning. like different ideas for reasoning. There was chain of thought, which was There was chain of thought, which was There was chain of thought, which was one idea. It's where you prompt for like one idea. It's where you prompt for like one idea. It's where you prompt for like kind of like the steps. There were kind of like the steps. There were kind of like the steps. There were scratch pads where you just like put scratch pads where you just like put scratch pads where you just like put like very like numerical kind of like very like numerical kind of like very like numerical kind of intermediate steps. But, Galactica was intermediate steps. But, Galactica was intermediate steps. But, Galactica was really this first idea which said, "No, really this first idea which said, "No, really this first idea which said, "No, this is an internal working memory this is an internal working memory this is an internal working memory process. This is an internal thinking. process. This is an internal thinking. process. This is an internal thinking. You should be inside these tags, and you You should be inside these tags, and you You should be inside these tags, and you should spend the inference computer should spend the inference computer should spend the inference computer before you get to an answer, right?" So, before you get to an answer, right?" So, before you get to an answer, right?" So, these are all like quite like pressing these are all like quite like pressing these are all like quite like pressing ideas. ideas. ideas. But, you know, Galactica came and went, But, you know, Galactica came and went, But, you know, Galactica came and went, blew up, and then, you know, we were blew up, and then, you know, we were blew up, and then, you know, we were kind of tasked as a team to kind of spin kind of tasked as a team to kind of spin kind of tasked as a team to kind of spin up the post-training effort for Llama. up the post-training effort for Llama. up the post-training effort for Llama. But, I had like a personal obsession, But, I had like a personal obsession, But, I had like a personal obsession, which was like reasoning. And I had like which was like reasoning. And I had like which was like reasoning. And I had like a really simple idea at the time, which a really simple idea at the time, which a really simple idea at the time, which is, "What if we applied kind of is, "What if we applied kind of is, "What if we applied kind of reinforcement learning pressure to this reinforcement learning pressure to this reinforcement learning pressure to this like kind of thinking tags as work?"
-
like kind of thinking tags as work?" like kind of thinking tags as work?" Like, what if we just optimize the thing Like, what if we just optimize the thing Like, what if we just optimize the thing in between the thinking? And if that in between the thinking? And if that in between the thinking? And if that sounds familiar, then this is like kind sounds familiar, then this is like kind sounds familiar, then this is like kind of what Deep Seek 01 ended up doing 2 of what Deep Seek 01 ended up doing 2 of what Deep Seek 01 ended up doing 2 years later. But, there's a key years later. But, there's a key years later. But, there's a key difference. So, at the time we only had difference. So, at the time we only had difference. So, at the time we only had Llama 2 base models, terrible Llama 2 base models, terrible Llama 2 base models, terrible mathematics corpus, terrible results on mathematics corpus, terrible results on mathematics corpus, terrible results on math. And, you know, the context window, math. And, you know, the context window, math. And, you know, the context window, you know, we're all like, you know, very you know, we're all like, you know, very you know, we're all like, you know, very context-rich now, 1 million tokens. You context-rich now, 1 million tokens. You context-rich now, 1 million tokens. You know, the back in the day it was 4,000. know, the back in the day it was 4,000. know, the back in the day it was 4,000. It wasn't too fun. It wasn't too fun. It wasn't too fun. But, we still had a recipe at the time, But, we still had a recipe at the time, But, we still had a recipe at the time, and this is unpublished, but it was a and this is unpublished, but it was a and this is unpublished, but it was a really good for the meta. So, RSP was really good for the meta. So, RSP was really good for the meta. So, RSP was this. this. this. Number one, continue pre-training of Number one, continue pre-training of Number one, continue pre-training of Llama 2 towards mathematics and science Llama 2 towards mathematics and science Llama 2 towards mathematics and science data. So, that's the first thing. Llama data. So, that's the first thing. Llama data. So, that's the first thing. Llama 2, math corpus, so let's fix that. 2, math corpus, so let's fix that. 2, math corpus, so let's fix that. Number two, PPO with verifiable rewards. Number two, PPO with verifiable rewards. Number two, PPO with verifiable rewards. But, notice this isn't GRPO, right? So, But, notice this isn't GRPO, right? So, But, notice this isn't GRPO, right? So, we had a time and a strong outcome we had a time and a strong outcome we had a time and a strong outcome reward model to initialize the value reward model to initialize the value reward model to initialize the value model. That was a key thing, lots of model. That was a key thing, lots of model. That was a key thing, lots of data on value models at the time. And data on value models at the time. And data on value models at the time. And internally at the time, this kind of internally at the time, this kind of internally at the time, this kind of recipe led to state-of-the-art results recipe led to state-of-the-art results recipe led to state-of-the-art results on math and reasoning. So, we were kind on math and reasoning. So, we were kind on math and reasoning. So, we were kind of like, "Wow, this is like really shows of like, "Wow, this is like really shows of like, "Wow, this is like really shows the power of having the right the power of having the right the power of having the right objective."
-
objective." objective." But, the really fascinating thing is But, the really fascinating thing is But, the really fascinating thing is like we had great results, but we didn't like we had great results, but we didn't like we had great results, but we didn't have like inference time scaling. We have like inference time scaling. We have like inference time scaling. We didn't have this reflective behavior didn't have this reflective behavior didn't have this reflective behavior that became like the hallmark of R1 and that became like the hallmark of R1 and that became like the hallmark of R1 and O1. You know, back weights, you know, O1. You know, back weights, you know, O1. You know, back weights, you know, back tracking, all this kind of stuff. back tracking, all this kind of stuff. back tracking, all this kind of stuff. So, it begs like the question like, So, it begs like the question like, So, it begs like the question like, "Why? Well, why didn't we have that "Why? Well, why didn't we have that "Why? Well, why didn't we have that moment?" moment?" moment?" And we got an answer around like 2 years And we got an answer around like 2 years And we got an answer around like 2 years later. So, there's a couple of like later. So, there's a couple of like later. So, there's a couple of like things going on here, but essentially things going on here, but essentially things going on here, but essentially better base models were the thing that better base models were the thing that better base models were the thing that really got RL cooking. And when DeepMind really got RL cooking. And when DeepMind really got RL cooking. And when DeepMind came out, I was kind of shocked at the came out, I was kind of shocked at the came out, I was kind of shocked at the time. I was like, "Holy we just time. I was like, "Holy we just time. I was like, "Holy we just like tried the same thing. We didn't like tried the same thing. We didn't like tried the same thing. We didn't have this. What's going on here?" And in have this. What's going on here?" And in have this. What's going on here?" And in a weird kind of way, the real lesson was a weird kind of way, the real lesson was a weird kind of way, the real lesson was it was just like the bitter lesson, like it was just like the bitter lesson, like it was just like the bitter lesson, like the most purest form of bitter lesson the most purest form of bitter lesson the most purest form of bitter lesson possible. Like, better base models, more possible. Like, better base models, more possible. Like, better base models, more RL computes, bigger context windows, and RL computes, bigger context windows, and RL computes, bigger context windows, and that's all you need for this kind of that's all you need for this kind of that's all you need for this kind of emergent behavior. It also like says emergent behavior. It also like says emergent behavior. It also like says something like quite important about the something like quite important about the something like quite important about the sociology of like research because the sociology of like research because the sociology of like research because the fact that OpenAI had this model, you fact that OpenAI had this model, you fact that OpenAI had this model, you know, GPT-4 level model before anyone know, GPT-4 level model before anyone know, GPT-4 level model before anyone else, it allowed them to see further, else, it allowed them to see further, else, it allowed them to see further, right? So, that's a really interesting right? So, that's a really interesting right? So, that's a really interesting point. Like, the the age of scaling point. Like, the the age of scaling point. Like, the the age of scaling means that if you have certain means that if you have certain means that if you have certain prerequisites in place, you become prerequisites in place, you become prerequisites in place, you become smarter, you see further, you see more smarter, you see further, you see more smarter, you see further, you see more ideas. So, really interesting point.
-
ideas. So, really interesting point. ideas. So, really interesting point. So, this was me in 2024. I was a very So, this was me in 2024. I was a very So, this was me in 2024. I was a very sad panda, uh defeated by ChatGPT and sad panda, uh defeated by ChatGPT and sad panda, uh defeated by ChatGPT and O1. O1. O1. But, I wasn't deterred. Um so, I wanted But, I wasn't deterred. Um so, I wanted But, I wasn't deterred. Um so, I wanted to seek the next wave, and I still was to seek the next wave, and I still was to seek the next wave, and I still was like convinced that kind of reasoning like convinced that kind of reasoning like convinced that kind of reasoning hadn't been solved. So, we started GR uh hadn't been solved. So, we started GR uh hadn't been solved. So, we started GR uh to take on truly like big tasks. And to take on truly like big tasks. And to take on truly like big tasks. And with that in mind, I'm going to hand with that in mind, I'm going to hand with that in mind, I'm going to hand over to Chengxi. He's going to talk over to Chengxi. He's going to talk over to Chengxi. He's going to talk about what we're kind of thinking about about what we're kind of thinking about about what we're kind of thinking about now. now. now. >> Ross. Ross. Ross. >> Hi, everyone. I'm Chengxi Taylor, >> Hi, everyone. I'm Chengxi Taylor, co-founder and president of General co-founder and president of General co-founder and president of General Intelligence Inc. Intelligence Inc. Intelligence Inc. I'm going to share what it takes to I'm going to share what it takes to I'm going to share what it takes to scale to long horizon. First, I want to make it clear. Long First, I want to make it clear. Long horizon task is not just an engineering horizon task is not just an engineering horizon task is not just an engineering problem. It is a mindset. problem. It is a mindset. problem. It is a mindset. If we want to solve humanity's biggest If we want to solve humanity's biggest If we want to solve humanity's biggest problems, such as cure cancer, solve problems, such as cure cancer, solve problems, such as cure cancer, solve millennium's prize problem, or go into millennium's prize problem, or go into millennium's prize problem, or go into Mars, Mars, Mars, we have to be patient. It take time. And we have to be patient. It take time. And we have to be patient. It take time. And if we want AI to move us towards that if we want AI to move us towards that if we want AI to move us towards that level impact, we have to think about level impact, we have to think about level impact, we have to think about long horizon.
-
But here's the But here's the first problem. We have a scarce context first problem. We have a scarce context first problem. We have a scarce context window. window. window. If you take Fermat's Last Theorem as If you take Fermat's Last Theorem as If you take Fermat's Last Theorem as example, example, example, what it take for the mathematician was what it take for the mathematician was what it take for the mathematician was over 10 years time of reading paper, over 10 years time of reading paper, over 10 years time of reading paper, writing down thoughts in the scratch writing down thoughts in the scratch writing down thoughts in the scratch pad, or taking a walk to generate the pad, or taking a walk to generate the pad, or taking a walk to generate the creative ideas. creative ideas. creative ideas. If we convert to token, that's probably If we convert to token, that's probably If we convert to token, that's probably tens of a billions of even hundreds tens of a billions of even hundreds tens of a billions of even hundreds billions. billions. billions. But where we are now, just a 1 million But where we are now, just a 1 million But where we are now, just a 1 million token token token context window. context window. context window. So one solution is use compaction. So So one solution is use compaction. So So one solution is use compaction. So what it essentially does is generate the what it essentially does is generate the what it essentially does is generate the token until the end of the context token until the end of the context token until the end of the context window, summarize, and then on top of window, summarize, and then on top of window, summarize, and then on top of that generate more tokens. that generate more tokens. that generate more tokens. And the beauty of applying RL in the And the beauty of applying RL in the And the beauty of applying RL in the situation is kind of like kill two birds situation is kind of like kill two birds situation is kind of like kill two birds with one stone. You apply RL to the with one stone. You apply RL to the with one stone. You apply RL to the compaction and also the task. compaction and also the task. compaction and also the task. But here's the problem. With long But here's the problem. With long But here's the problem. With long horizon, there are three issues. The horizon, there are three issues. The horizon, there are three issues. The first is the gradient variance scales first is the gradient variance scales first is the gradient variance scales with the length. And the second is a with the length. And the second is a with the length. And the second is a sparse reward.
-
sparse reward. sparse reward. And you have this credit assignment And you have this credit assignment And you have this credit assignment problem. And finally, there's also problem. And finally, there's also problem. And finally, there's also variable length of the trajectory that variable length of the trajectory that variable length of the trajectory that adds to the problem of optimization. adds to the problem of optimization. adds to the problem of optimization. So to solve this issue, we can apply So to solve this issue, we can apply So to solve this issue, we can apply critics, which is the value model. critics, which is the value model. critics, which is the value model. And value model can reduce the variance And value model can reduce the variance And value model can reduce the variance and also have a couple advantages, such and also have a couple advantages, such and also have a couple advantages, such as on the trajectory level that fits as on the trajectory level that fits as on the trajectory level that fits compaction very well, and also encourage compaction very well, and also encourage compaction very well, and also encourage the batch diversity. the batch diversity. the batch diversity. And also, I'll talk later on And also, I'll talk later on And also, I'll talk later on bootstrapping. Basically, get signal bootstrapping. Basically, get signal bootstrapping. Basically, get signal before the end of the episode. before the end of the episode. before the end of the episode. But, the downside for this is that um But, the downside for this is that um But, the downside for this is that um it's more complicated than GRPO, and it's more complicated than GRPO, and it's more complicated than GRPO, and basically, you have to train another basically, you have to train another basically, you have to train another value model alongside with the policy value model alongside with the policy value model alongside with the policy model. model. model. And there's some tools to help with the And there's some tools to help with the And there's some tools to help with the context limitations, such as a file context limitations, such as a file context limitations, such as a file system tools, which essentially like a system tools, which essentially like a system tools, which essentially like a scratchpad for AI to write to the uh scratchpad for AI to write to the uh scratchpad for AI to write to the uh reasoning thought. reasoning thought. reasoning thought. And self-search tools, which allows And self-search tools, which allows And self-search tools, which allows agent to search over the previous agent to search over the previous agent to search over the previous trajectory.
-
trajectory. trajectory. And then you have archive tools. And then you have archive tools. And then you have archive tools. In the case like auto research, you can In the case like auto research, you can In the case like auto research, you can build upon your previous result. But, we build upon your previous result. But, we build upon your previous result. But, we have to be careful. In other scenario, have to be careful. In other scenario, have to be careful. In other scenario, you don't want an AI to cheat by just to you don't want an AI to cheat by just to you don't want an AI to cheat by just to grab the previous answer without grab the previous answer without grab the previous answer without thinking. thinking. thinking. So, So, So, how good are the current model on this how good are the current model on this how good are the current model on this long horizon task? long horizon task? long horizon task? In the general reasoning, we construct In the general reasoning, we construct In the general reasoning, we construct the benchmark called a Kelly bench, the benchmark called a Kelly bench, the benchmark called a Kelly bench, where we're actually featured in the where we're actually featured in the where we're actually featured in the front page of the Financial Times. Kind front page of the Financial Times. Kind front page of the Financial Times. Kind of a caught us off guard how much the of a caught us off guard how much the of a caught us off guard how much the mainstream have interest in this. mainstream have interest in this. mainstream have interest in this. So, basically, what we did is that we So, basically, what we did is that we So, basically, what we did is that we allowed the agents to build machine allowed the agents to build machine allowed the agents to build machine learning models to learning models to learning models to um trade in the football matches over um trade in the football matches over um trade in the football matches over 1-year horizon. In this case, it's a 1-year horizon. In this case, it's a 1-year horizon. In this case, it's a Premier League, if you're interested in Premier League, if you're interested in Premier League, if you're interested in football. football. football. And we're so fascinated by this because And we're so fascinated by this because And we're so fascinated by this because there's a real money to be made, and if there's a real money to be made, and if there's a real money to be made, and if it was a successful, couldn't make a it was a successful, couldn't make a it was a successful, couldn't make a billions. There's a whole industry on billions. There's a whole industry on billions. There's a whole industry on sports betting. And unlike things like a sports betting. And unlike things like a sports betting. And unlike things like a cargo competition, this has a real-world cargo competition, this has a real-world cargo competition, this has a real-world implication.
-
implication. implication. But, But, But, here's the result. here's the result. here's the result. As you can see, As you can see, As you can see, we gave all the frontier models a 100K we gave all the frontier models a 100K we gave all the frontier models a 100K to start. All of them lost her to start. All of them lost her to start. All of them lost her Sad. And that captured the public's Sad. And that captured the public's Sad. And that captured the public's imagination. Oh, AI is not as great as imagination. Oh, AI is not as great as imagination. Oh, AI is not as great as they thought. they thought. they thought. And why are models so bad at long And why are models so bad at long And why are models so bad at long horizon? horizon? horizon? First, I believe now the AI industry is First, I believe now the AI industry is First, I believe now the AI industry is a little bit too biased towards coding a little bit too biased towards coding a little bit too biased towards coding and procedure task. What I mean is that and procedure task. What I mean is that and procedure task. What I mean is that the current task is a most formulated the current task is a most formulated the current task is a most formulated like like like do this and fix that. Normally that do this and fix that. Normally that do this and fix that. Normally that limits the solution like one or two. limits the solution like one or two. limits the solution like one or two. There isn't just too much space for There isn't just too much space for There isn't just too much space for creativity. creativity. creativity. And second of all, not enough focus on And second of all, not enough focus on And second of all, not enough focus on open-ended task. open-ended task. open-ended task. We live in the real world with a lot of We live in the real world with a lot of We live in the real world with a lot of a complexity, uncertainty, and that's a complexity, uncertainty, and that's a complexity, uncertainty, and that's not fully captured by the current not fully captured by the current not fully captured by the current benchmark. benchmark. benchmark. And also, there isn't enough simulation And also, there isn't enough simulation And also, there isn't enough simulation of the real world. We live in the world of the real world. We live in the world of the real world. We live in the world that there are other players. Like in that there are other players. Like in that there are other players. Like in today's conference room, there are other today's conference room, there are other today's conference room, there are other real people who have a different real people who have a different real people who have a different thought, a different games than you have thought, a different games than you have thought, a different games than you have in your mind. That's the complexity in your mind. That's the complexity in your mind. That's the complexity that's not fully captured.
-
And another thing I want to talk about And another thing I want to talk about is the long horizon impact on compute. is the long horizon impact on compute. is the long horizon impact on compute. We know that GPUs are scarce and We know that GPUs are scarce and We know that GPUs are scarce and precious resources. And in this case of precious resources. And in this case of precious resources. And in this case of a long horizon reasoning, you have to be a long horizon reasoning, you have to be a long horizon reasoning, you have to be careful about how to optimize your use careful about how to optimize your use careful about how to optimize your use between training and inference. between training and inference. between training and inference. And pipeline RL is a quite popular And pipeline RL is a quite popular And pipeline RL is a quite popular technique nowadays. So basically, it's a technique nowadays. So basically, it's a technique nowadays. So basically, it's a trade-off between off-policy and the GPU trade-off between off-policy and the GPU trade-off between off-policy and the GPU utilization. utilization. utilization. So traditionally, you let inference run So traditionally, you let inference run So traditionally, you let inference run towards the end and then you start to towards the end and then you start to towards the end and then you start to train the model. train the model. train the model. But in the case of long horizon, you But in the case of long horizon, you But in the case of long horizon, you have to wait until the inference finish. have to wait until the inference finish. have to wait until the inference finish. What the pipeline RL does is that you What the pipeline RL does is that you What the pipeline RL does is that you let the sequence to be generated and you let the sequence to be generated and you let the sequence to be generated and you start to train the model while there's a start to train the model while there's a start to train the model while there's a still more sequences being generated. still more sequences being generated. still more sequences being generated. And you see this created off-policy. But And you see this created off-policy. But And you see this created off-policy. But from the our experience, normally from the our experience, normally from the our experience, normally off-policy up to eight steps is okay. off-policy up to eight steps is okay. off-policy up to eight steps is okay. So, essentially, we made a trade-off So, essentially, we made a trade-off So, essentially, we made a trade-off between the off-policy and the GPU between the off-policy and the GPU between the off-policy and the GPU utilization.
-
utilization. utilization. But, here comes the issue. As we the But, here comes the issue. As we the But, here comes the issue. As we the long horizon indicates, sometimes the long horizon indicates, sometimes the long horizon indicates, sometimes the inference would take weeks or even more. inference would take weeks or even more. inference would take weeks or even more. In that case, inevitably, it will goes In that case, inevitably, it will goes In that case, inevitably, it will goes beyond the constraint of the eight steps beyond the constraint of the eight steps beyond the constraint of the eight steps of our policy. So, your GPU have just to of our policy. So, your GPU have just to of our policy. So, your GPU have just to sit there idle and wait for it to sit there idle and wait for it to sit there idle and wait for it to finish. finish. finish. And if you don't want to wait, as I And if you don't want to wait, as I And if you don't want to wait, as I mentioned before, mentioned before, mentioned before, applying the value model allows you to applying the value model allows you to applying the value model allows you to bootstrap. What it means is that before bootstrap. What it means is that before bootstrap. What it means is that before the end of the episode, you generate the end of the episode, you generate the end of the episode, you generate expectation. It's like a dopamine in expectation. It's like a dopamine in expectation. It's like a dopamine in human brain. And that allows you to human brain. And that allows you to human brain. And that allows you to train the model. But, here's another train the model. But, here's another train the model. But, here's another trade-off. While you utilize the GPU trade-off. While you utilize the GPU trade-off. While you utilize the GPU fully, you introduce the value model fully, you introduce the value model fully, you introduce the value model bias. So, there's always a bit of bias. So, there's always a bit of bias. So, there's always a bit of trade-off in those solutions. trade-off in those solutions. trade-off in those solutions. And I want to also mention that in the And I want to also mention that in the And I want to also mention that in the long horizon, infrastructure is long horizon, infrastructure is long horizon, infrastructure is important, especially for the important, especially for the important, especially for the environment. And Open Review was a environment. And Open Review was a environment. And Open Review was a product is a platform by General product is a platform by General product is a platform by General Reasoning. If you're interested, you can Reasoning. If you're interested, you can Reasoning. If you're interested, you can check it out. openreview.ai.
-
check it out. openreview.ai. check it out. openreview.ai. So, it's a place where host over 350 So, it's a place where host over 350 So, it's a place where host over 350 environments and with a single API environments and with a single API environments and with a single API endpoint. And we use this for our endpoint. And we use this for our endpoint. And we use this for our internal RL and also some frontier labs internal RL and also some frontier labs internal RL and also some frontier labs and new labs are using this. So, So, to summarize both Ross and my speech, it to summarize both Ross and my speech, it to summarize both Ross and my speech, it has been a long journey as long horizon has been a long journey as long horizon has been a long journey as long horizon indicate. indicate. indicate. We as a team have seen the paradigms in We as a team have seen the paradigms in We as a team have seen the paradigms in AI reasoning on pre-training and agents AI reasoning on pre-training and agents AI reasoning on pre-training and agents in the past few years. But, looking in the past few years. But, looking in the past few years. But, looking ahead, what makes us really excited is ahead, what makes us really excited is ahead, what makes us really excited is the long horizon. the long horizon. the long horizon. And it requires us to think, have a new And it requires us to think, have a new And it requires us to think, have a new thinking on the algorithm, environments, thinking on the algorithm, environments, thinking on the algorithm, environments, and compute. There are a lot of and compute. There are a lot of and compute. There are a lot of challenges and trade-offs, but we find challenges and trade-offs, but we find challenges and trade-offs, but we find it's really exciting to take on this it's really exciting to take on this it's really exciting to take on this journey because, as I mentioned in the journey because, as I mentioned in the journey because, as I mentioned in the very beginning, long horizon is not just very beginning, long horizon is not just very beginning, long horizon is not just engineering problem. It is a mindset, engineering problem. It is a mindset, engineering problem. It is a mindset, and if we really are ambitious to solve and if we really are ambitious to solve and if we really are ambitious to solve humanity's biggest problems, this is the humanity's biggest problems, this is the humanity's biggest problems, this is the journey for everyone.
-
journey for everyone. journey for everyone. And that's also the mission for general And that's also the mission for general And that's also the mission for general reasoning. So, if you are interested, reasoning. So, if you are interested, reasoning. So, if you are interested, follow us. General Reasoning, we're follow us. General Reasoning, we're follow us. General Reasoning, we're London-based AI research company. Thank London-based AI research company. Thank London-based AI research company. Thank you.
Summary
This tech talk, "Scaling to Long Horizons," discusses the challenges and strategies for developing large-scale AI agents. Key subjects include reinforcement learning, large language models like Llama and Galactica, and the importance of algorithms, environments, and compute. The practical takeaway is that small, focused teams can achieve remarkable results in AI development through alignment, even in an era of scaling.